Underwater operation and maintenance digital twinning and dual-mode parallel operation method based on real-time data
By using underwater binocular cameras and high-precision hand-eye calibration technology, combined with hierarchical multimodal feature constraint registration, an underwater operation and maintenance digital twin system was constructed. This system enables three-dimensional spatial pose estimation and virtual-real synchronous control for underwater operations, solving the problems of visual limitations and insufficient real-time feedback in underwater operation and maintenance, and improving the safety and accuracy of operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHEJIANG UNIV
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies cannot achieve macro-level global situational awareness, micro-level non-destructive vision, and real-time interactive control in underwater operations. This makes it difficult for operators to accurately judge the details, distance, and physical state of targets, resulting in a lack of presence and spatial depth. Furthermore, existing methods cannot provide real-time feedback to support efficient and precise underwater operations.
The underwater operation and maintenance digital twin method based on real-time data and dual-mode parallel operation is adopted. The observation point cloud is collected by underwater binocular camera, and the three-dimensional spatial pose estimation of the target object is realized by combining high-precision hand-eye calibration and hierarchical multimodal feature constraint registration. Through virtual-real synchronous control and pose mapping, an immersive digital twin space is constructed, which supports the operation process of seamless switching between simulation and reality.
It significantly improves the safety and accuracy of underwater operation and maintenance. Through a dual-mode parallel operation method, it provides a sense of visual presence and spatial depth perception, reduces operational risks and the probability of equipment collisions, and realizes visualization and precise control of the entire underwater operation process.
Smart Images

Figure CN121870735A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an underwater parallel operation method, and to the field of underwater operation and maintenance operations and digital twin virtual-real collaboration technology, specifically to an underwater operation and maintenance digital twin and dual-mode parallel operation method based on real-time data. Background Technology
[0002] In underwater operations, operators rely on traditional visual feedback, which involves displaying a single underwater video stream. Due to unfavorable imaging factors such as uneven underwater lighting, scattering from suspended objects, and color distortion, the image quality is poor and contrast is low. Operators struggle to accurately judge the details, distance, and physical state of targets, resulting in a severe lack of presence and spatial depth. While some visual enhancement methods, such as image restoration and enhancement algorithms, can improve image quality, they are essentially improvements on a two-dimensional plane and cannot fundamentally provide three-dimensional spatial perception. Although multi-camera switching can expand the field of view, it still relies on operators manually switching between multiple two-dimensional images and mentally piecing them together, leading to a high cognitive load and a high risk of misjudgment. In recent years, some research has attempted to enhance immersion using offline 3D reconstruction and VR technology, but the reconstruction process is time-consuming, real-time mapping is impossible, and the generated static models cannot reflect dynamic changes in the operational status, making them unsuitable for real-time decision-making. Therefore, existing technologies cannot construct a digital twin virtual-real collaborative solution that integrates macro-level global situational awareness, micro-level non-destructive vision, and real-time interactive control to overcome the limitations of two-dimensional vision and the lag of offline models, and achieve efficient, accurate, and safe underwater operation and maintenance. Summary of the Invention
[0003] To address the problems existing in the background technology, this invention provides a method for underwater operation and maintenance digital twin based on real-time data and dual-mode parallel operation. This invention constructs an immersive digital twin space integrating operator learning and operation through a real-time data stream and a dual-mode parallel approach. It recreates a 1:1 underwater operation and maintenance scenario for operators, providing a sense of visual presence and spatial depth perception, thereby improving the safety, accuracy, and decision-making efficiency of underwater operation and maintenance operations.
[0004] The technical solution adopted in this invention is: The underwater operation and maintenance digital twin and dual-mode parallel operation method based on real-time data of the present invention includes: Step 1: Use an underwater robotic arm with an underwater binocular camera mounted at the end to collect observation point clouds of the target object in the underwater scene in real time. Then, relying on high-precision hand-eye calibration parameters, use a hierarchical multimodal feature constraint registration method and a cross-coordinate system pose mapping method to obtain the three-dimensional spatial pose of the target object.
[0005] The second step is to acquire the angle data of each joint of the underwater robotic arm in real time, and combine the three-dimensional spatial pose of the target object and the underwater scene to establish a digital twin scene including the digital twin robotic arm. A virtual-real synchronous control and pose mapping mechanism is designed to establish a data channel between the real underwater robotic arm and the digital twin robotic arm, so as to realize the accurate reproduction of the motion trajectory of the real robotic arm by the digital twin robotic arm.
[0006] Step 3: Based on the data channel between the real underwater robotic arm and the digital twin robotic arm, the underwater robotic arm is controlled to perform precise operations on the target object through two parallel working methods that can be switched between simulation and real-world control, thus achieving dual-mode parallel operation.
[0007] In the first step, the hierarchical multimodal feature constraint registration method includes an initial alignment layer, a feature constraint optimization layer containing a multimodal feature consistency metric function, and a precise convergence layer. The observation point cloud of the target object and the reference model point cloud are processed by the initial alignment layer to obtain the alignment transformation matrix of the target object. After the alignment transformation matrix is passed through the feature constraint optimization layer and the precise convergence layer, the three-dimensional spatial pose of the target object in the camera coordinate system of the underwater binocular camera is obtained according to the multimodal feature consistency metric function.
[0008] The target object's observation point cloud is a point cloud acquired by an underwater binocular camera in the camera coordinate system, and the target object's reference model point cloud is a point cloud obtained based on the target object's 3D model in the model coordinate system. The target object's observation point cloud and the reference model point cloud are processed through an initial alignment layer to first obtain the feature similarity between each point in the target object's observation point cloud and the reference model point cloud. Then, point pairs with similarity higher than a preset similarity threshold are selected to form a matching point set. C Based on the matching point set C Construct the alignment transformation matrix of the target object as follows: in, T These are the pose transformation parameters; p i and q j These are the matching point sets. C Observation points in i and reference model points j ; It is the square of the 2-norm.
[0009] In the feature constraint optimization layer, the alignment transformation matrix is... Model feature points are obtained after projection. And serve as a search center, Thus, a radius of [missing information] is constructed. The spherical neighborhood, radius as follows: in, r 0 represents the initial search radius; For the number of iterations, This is the attenuation coefficient.
[0010] With radius The spherical neighborhood serves as the adaptive spatial constraint boundary for the multimodal feature consistency measurement function, strictly controlling the effective range of feature consistency. Outliers are eliminated through iterative screening, thereby providing high-confidence input data for subsequent multimodal feature consistency measurement functions.
[0011] The aforementioned multimodal feature consistency measurement function Specifically as follows: in, T These are the pose transformation parameters; For adaptive weight parameters; and These are the point distance term and the feature preservation term, respectively. p i and q j These are the matching point sets. C Observation points in i and reference model points j ; F ( ) is the feature extraction function.
[0012] Furthermore, a multimodal feature consistency metric function is set through a precise convergence layer, and an adaptive step size update mechanism is used to adjust the pose transformation parameters. T By performing iterative solutions, the three-dimensional spatial pose of the target object in the camera coordinate system of the underwater stereo camera is finally obtained. .
[0013] In the first step, the target object is positioned in three-dimensional space within the camera coordinate system of the underwater binocular camera. The three-dimensional spatial pose of the target object in the base coordinate system of the underwater robotic arm is obtained after processing using a cross-coordinate system pose mapping method. ,as follows: in, It is a rotation matrix; and These are the offset vectors along the y and z axes between the bases of the underwater binocular camera and the underwater robotic arm, respectively.
[0014] In the third step, the operator drives the digital twin robotic arm to perform operations on the target object in the digital twin scenario, thereby obtaining the real-time joint angles and end-effector poses of the digital twin robotic arm as operation commands. The operation commands are then sent to the underwater robotic arm in real time for control, forming a simulation-real-scene operation process.
[0015] This invention presents a digital twin and dual-mode parallel operation method for underwater operations based on real-time data, utilizing binocular cameras, an underwater electric robotic arm, VR equipment, and remote control devices. It calculates the target's pose in real time using binocular vision and a hierarchical feature registration algorithm, and achieves synchronous control of the virtual and real robotic arms based on joint data and pose mapping. A high-fidelity underwater twin scene integrating real data is constructed in Unity3D, supporting physical collision detection and interactive operation. Simultaneously, it provides both simulation and real-world working modes, supporting seamless switching and closed-loop optimization between risk-free pre-simulation and real-world operations. Ultimately, it achieves visualization, simulation, and precise control of the entire underwater operation process, significantly improving the accuracy, safety, and efficiency of operations.
[0016] The beneficial effects of this invention are: This invention constructs a synchronous mapping between the real underwater environment and the virtual scene by integrating high-precision pose estimation from binocular vision recognition with real-time data feedback from the robotic arm joints. Operators, using VR devices, can immerse themselves in a first-person perspective to observe the real-time relative pose between the robotic arm's end effector and the underwater target. This effectively overcomes challenges such as low visibility, visual occlusion, and light scattering in traditional underwater operations, providing reliable spatial perception and decision support for precise operations and significantly reducing operational risks and the probability of equipment collisions in complex underwater environments.
[0017] This invention implements a one-click switching and parallel operation method for dual working modes. In simulation mode, risk-free task rehearsals and operational training can be conducted, significantly saving equipment wear and time costs. In real-world mode, the system uses virtual-real synchronization technology to map the real equipment status to the virtual environment in real time, supporting precise underwater operation control. The two modes achieve seamless switching and data synchronization through a unified human-machine interface, forming a closed-loop operation process of "simulation rehearsal → real-world execution → data feedback → optimization." This ensures operational continuity and provides a complete integrated solution, significantly improving the reliability and adaptability of underwater operation and maintenance.
[0018] This invention aims to solve the problems of low target positioning accuracy and insufficient pre-operation simulation and real-time feedback capabilities in underwater environments. It can realize visualization, simulation and precise control of the entire underwater operation process, significantly improving the accuracy, safety and efficiency of the operation. Attached Figure Description
[0019] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a general block diagram of the present invention; Figure 3 This is a flowchart of the real-world mode of the present invention; Figure 4 This is a flowchart of the simulation mode of the present invention. Detailed Implementation
[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. The specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0021] like Figure 1 As shown, the underwater operation and maintenance digital twin and dual-mode parallel operation method based on real-time data of the present invention is as follows: Step 1: Use an underwater robotic arm with an underwater binocular camera mounted at the end to collect observation point clouds of the target object in the underwater scene in real time. Then, relying on high-precision hand-eye calibration parameters, use a hierarchical multimodal feature constraint registration method and a cross-coordinate system pose mapping method to obtain the three-dimensional spatial pose of the target object.
[0022] The hierarchical multimodal feature-constrained registration method includes an initial alignment layer, a feature constraint optimization layer containing a multimodal feature consistency metric function, and a precise convergence layer. After processing the observation point cloud of the target object and the reference model point cloud through the initial alignment layer, the alignment transformation matrix of the target object is obtained. This matrix is then passed through the feature constraint optimization layer and the precise convergence layer. Finally, the 3D spatial pose of the target object in the camera coordinate system of the underwater stereo camera is obtained based on the multimodal feature consistency metric function. The hierarchical multimodal feature constraint registration method calculates the transformation relationship from the model coordinate system to the camera coordinate system, thereby calculating the pose of the target object relative to the camera coordinate system. That is, the model coordinate system To the camera coordinate system The transformation.
[0023] The observation point cloud of the target object is a point cloud in the camera coordinate system acquired by an underwater binocular camera. The reference model point cloud of the target object is a point cloud in the model coordinate system obtained based on the 3D model of the target object. Both are obtained by converting the target object's CAD model into standard 3D data. and reference model point cloud Through initial alignment layer processing, a coarse, global pose estimation is quickly achieved, providing a good starting point for subsequent fine registration. First, the feature similarity between each point in the observation point cloud of the target object and the reference model point cloud is obtained. Then, point pairs with similarity higher than a preset similarity threshold are selected to form a matching point set. C Based on the matching point set C Construct the alignment transformation matrix of the target object as follows: in, T These are the pose transformation parameters. , R For rotation matrix, t It is a translation vector. It is a three-dimensional special Euclidean group; p i and q j These are the matching point sets. C Observation points in i and reference model points j ; It is the square of the 2-norm.
[0024] Pose transformation parameters T Based on minimizing the generalized feature distance function The transformation for each sample is calculated as follows: in, For the set of feature extraction operators; For the first Feature extraction operators For the first Adaptive weighting coefficients for feature extraction operators (such as geometric structure spectrum, surface curvature distribution, etc.); For the first The metric function of the feature space corresponding to the feature extraction operator.
[0025] After defining the function to minimize the generalized feature distance, initial alignment begins by calculating the multi-scale geometric features of the point cloud. as follows: in, The surface change rate is a characteristic. The distribution characteristics are the included angle of the normal direction. The feature is the local curvature entropy feature, and the superscript T denotes the transpose of the vector. Then, a feature similarity matrix is constructed. as follows: in, The bandwidth parameter of the Gaussian kernel function; and These are the matching point sets. C Observation points in i and reference model points j Multiscale geometric features; For matching point sets C Observation points in i Features and reference model points j Spatial distance constraints for features.
[0026] Based on feature similarity matrix A high-confidence matching point set is constructed by selecting point pairs with high similarity scores and satisfying geometric constraints. This transforms the similarity measure in the feature space into a geometric registration problem in Euclidean space. Then, the transformation matrix of the initial alignment layer is solved using the feature correspondence.
[0027] In the feature-constrained optimization layer, the alignment transformation matrix is... Model feature points are obtained after projection. And serve as a search center, Thus, a radius of [missing information] is constructed. The spherical neighborhood, radius as follows: in, r 0 represents the initial search radius; For the number of iterations, This is the attenuation coefficient.
[0028] With radius The spherical neighborhood serves as the adaptive spatial constraint boundary for the multimodal feature consistency measurement function, strictly controlling the effective range of feature consistency. Outliers are eliminated through iterative screening, thereby providing high-confidence input data for subsequent multimodal feature consistency measurement functions.
[0029] Multimodal feature consistency measurement function Specifically as follows: in, T These are the pose transformation parameters; For adaptive weight parameters; and These are the point distance term and the feature preservation term, respectively. p i and qj These are the matching point sets. C Observation points in i and reference model points j ; F ( ) represents the feature extraction function; feature preservation term While ensuring spatial alignment, the high-level feature descriptions also remain consistent, enhancing the algorithm's robustness to noise and partial occlusion.
[0030] Furthermore, a multimodal feature consistency metric function is set through a precise convergence layer, and an adaptive step size update mechanism is used to adjust the pose transformation parameters. T By performing iterative solutions, the three-dimensional spatial pose of the target object in the camera coordinate system of the underwater stereo camera is finally obtained. , , and These are the optimized rotation matrix and translation vector, respectively.
[0031] The feature-constrained optimization layer utilizes multimodal features and their consistency constraints to perform fine-grained pose solving and optimization. The precision convergence layer aims to eliminate minute errors, ensuring that the solution converges stably and accurately to the optimal value, and employs an adaptive step-size update strategy to update the pose transformation parameters. T ,in, This is an adaptive step size factor. This is represented by a Lie algebra. The precise convergence layer uses a stopping criterion based on the rate of change of error, setting two convergence conditions, and stopping iteration only when both are met: in, Representing the t During the next iteration value, Representing the t -1 iteration value; and These are the first and second convergence thresholds, respectively.
[0032] After the iteration is complete, the pose matrix is obtained as follows: This matrix accurately describes the spatial pose of the target object in the binocular camera coordinate system {O}.
[0033] The three-dimensional spatial pose of the target object in the camera coordinate system of the underwater binocular camera. The three-dimensional spatial pose of the target object in the base coordinate system of the underwater robotic arm is obtained after processing using a cross-coordinate system pose mapping method. ,as follows: in, It is a rotation matrix; and These are the offset vectors along the y and z axes between the bases of the underwater binocular camera and the underwater robotic arm, respectively.
[0034] Transforming the coordinates between the stereo camera and the underwater robotic arm requires a rotation matrix due to the different coordinate system definitions. Transform the point from the binocular camera coordinate system {O} to the robotic arm base coordinate system {B}. In the base coordinate system {B}: x-axis forward, y-axis left, z-axis upward. In the binocular camera coordinate system {O}: x-axis right, y-axis downward, z-axis forward along the optical axis. Then its rotation matrix... for: The columns of the rotation matrix represent the directions of the x, y, and z axes of {O} within {B}. The relative distance between the binocular camera and the robotic arm base in the actual system is measured. Since there is an offset between the binocular camera and the robotic arm base in the actual system, this offset vector is determined after hand-eye calibration. Therefore, the offset vector in {B} is represented as Thus, the target pose in {O} is obtained. Transform the relations in {B}.
[0035] When acquiring the pose of an underwater target, the system uses a binocular camera to acquire its depth and attitude information in real time, reconstructing the precise spatial position of the actual target. Since the depth and pose error of binocular readings is at the millimeter level, a hierarchical multimodal feature constraint registration framework is proposed. The "hierarchical multimodal feature constraint" in this framework emphasizes joint constraint and optimization of the target's spatial pose through features from multiple layers and various sensor data. The method achieves high-precision pose estimation against underwater environmental interference by constructing a multimodal feature consistency metric function and an adaptive step-size optimization mechanism. The multimodal feature consistency metric function is a loss function used to measure the degree of matching or consistency between feature data from different sensors. The adaptive step-size optimization mechanism is an optimization strategy that dynamically adjusts the step size for parameter updates based on the current registration state (e.g., error magnitude, convergence status). The overall process consists of three layers: an initial alignment layer, a feature constraint optimization layer, and a precise convergence layer.
[0036] The second step is to acquire the angle data of each joint of the underwater robotic arm in real time, and combine the three-dimensional spatial pose of the target object and the underwater scene to establish a digital twin scene including the digital twin robotic arm. A virtual-real synchronous control and pose mapping mechanism is designed to establish a data channel between the real underwater robotic arm and the digital twin robotic arm, so as to realize the accurate reproduction of the motion trajectory of the real robotic arm by the digital twin robotic arm.
[0037] Based on real-time data streams of the target object's 3D spatial pose and the joint angles of the real robotic arm, a digital twin scene consistent with the real physical environment is constructed and dynamically updated in the Unity3D engine. This dynamically rendered scene includes a virtual underwater lighting environment, a virtual water body with simulated physical properties, a digital twin robotic arm driven by real data, and a virtual target object with pose mapping, forming a simulation environment with physical realism and interactive immersion.
[0038] A simulated underwater experimental scenario is constructed within a digital twin environment. A virtual robotic arm model, driven by joint data from a real robotic arm, is deployed, and a point cloud model of the target object is overlaid. The specific position of this point cloud model is determined by the pose established in the first step. Operators control the real robotic arm via teleoperation devices. The joint angle data of the real robotic arm is collected in real-time through the ROS system and transmitted to the Unity3D client via TCP communication. Upon receiving this data, the Unity3D client directly drives the corresponding joints of the digital twin robotic arm. Note that a calibration is required during the initial system setup. First, a specific, easily identifiable calibration feature point is selected in the real world—a three-dimensional sphere fixed on a workbench. Then, the real robotic arm is controlled so that its end effector points precisely at this feature point from different angles and positions, and the end effector pose at each moment is recorded. This is a 4×4 homogeneous transformation matrix, derived from the rotation matrix. Translation vector Composition, that is The pose of the digital twin robotic arm model is the end effector pose of the virtual robotic arm model in Unity3D. A fixed transformation matrix needs to be found. The points are transformed from the real robotic arm coordinate system to the digital twin robotic arm coordinate system, that is: in, This represents the actual end effector pose of the robotic arm.
[0039] Then, by collecting multiple sets of data, a system of linear equations can be obtained. For the th... Data collected this time: By solving the above system of equations, we can obtain This process is similar to the robot hand-eye calibration problem, and the least squares method is used here to find the optimal solution. After completing the virtual-real mapping, the next step is to provide real-time actuation for the digital twin robotic arm and the target object.
[0040] The system receives the joint angle of the robotic arm using the TCP protocol. Because the chirality of a real robotic arm differs from that of a virtual robotic arm in Unity3D, a joint space mapping function needs to be constructed to... Perform a mapping transformation, and then use it to directly drive the virtual robotic arm model in Unity3D.
[0041] The construction of the digital twin scenario is as follows: An interactive 3D simulation environment is built based on a real underwater scene model. Keyboard and VR controllers are used to control virtual space movement, viewpoint switching, and the robotic arm's grasping operations. The system introduces a physics engine-based collision detection mechanism. Colliders matching the shape of the robotic arm gripper and the underwater target are added, and corresponding physical materials are configured. The target's collider is set as a kinematic rigid body, and its gravity properties are enabled, allowing it to respond to external forces when not being grasped. To achieve automatic grasping, a collision detection script is bound to these two colliders. This script continuously monitors the contact state between the gripper and the target's colliders: once a collision is detected and the grasping conditions are met, a control signal is immediately triggered, driving the robotic arm to tighten the gripper to grasp the target. At this time, the target's kinematic properties are dynamically adjusted to ensure it can stably follow the robotic arm's movement. When the robotic arm moves to the designated position, the operator controls the robotic arm to release the gripper via button commands. The grasping state is released, and the target's gravity properties are activated, thus completing a full grasping-placement cycle. Simulates the physical contact, interference, and force feedback effects during the grasping process, providing a highly realistic virtual environment for operational training and task rehearsals.
[0042] Step 3: Based on the data channel between the real underwater robotic arm and the digital twin robotic arm, the underwater robotic arm is controlled to perform precise operations on the target object through two parallel working methods that can be switched between simulation and real-world control, thus achieving dual-mode parallel operation.
[0043] In a digital twin scenario, the operator drives a digital twin robotic arm to perform operations on a target object, thereby obtaining the real-time joint angles and end-effector pose of the digital twin robotic arm as operating commands. These commands are then sent in real-time to the underwater robotic arm for control, forming a simulation-real-scene operation process. In simulation mode, the operator can conduct risk-free task planning, collision detection, and operational drills based on the digital twin scenario. In real-scene mode, the system sends operating commands to the real robotic arm in real-time, achieving precise operations on underwater targets, thus forming a "simulation-real-scene" operation process.
[0044] The digital twin system of this invention not only constructs a high-fidelity underwater operation and maintenance scenario, but also integrates a two-dimensional digital twin interface to present multi-sensor data information. It integrates communication status monitoring, mode switching control, multi-channel video monitoring, equipment signal connection, operational status, robotic arm joint data, underwater operations, and other information as the core information hub of the system. The communication connection module allows users to input the host's IP address (Internet Protocol) and port number via keyboard to establish TCP (Transmission Control Protocol) communication with the ROS (Robot Operating System) system, which serves as the control core in the real environment. The mode switching module enables one-click switching between simulation mode and real-world mode. The multi-channel video monitoring module displays real-time first-person perspective video from the virtual 3D (Three-Dimensional) scene and real third-person perspective camera video streams transmitted via RTSP (Real Time Streaming Protocol) network, providing multi-dimensional visual monitoring. The signal connection module detects the connection status of TCP communication and video signals in real time. The operational status module accurately marks the current task state of the system. The robotic arm joint data module maps and displays the angles of each joint of the real robotic arm in real time, while simultaneously driving the virtual robotic arm to achieve motion following.
[0045] In mode switching control, this system innovatively provides a dual working mode that can run in parallel and supports hot switching: real-scene control and simulation training. This constructs an integrated closed-loop operation system of "training-practice". In simulation training mode, operators can conduct risk-free full-process operation rehearsals based on high-fidelity digital twin scenarios, including visualized task planning, real-time physical collision detection, and operational skills training. The system automatically records key indicators such as operation time and path accuracy for performance evaluation. When switching to real-scene control mode, operation commands are sent to the real robotic arm controller with millisecond-level latency, driving it to complete precise operations on underwater targets. The two modes achieve seamless one-click switching through a unified human-machine interface. All operation commands and status data are kept in bidirectional synchronization, forming an intelligent closed-loop operation process of "simulation rehearsal → real-scene execution → data feedback → optimization".
[0046] like Figure 2As shown, the system architecture proposed in this invention mainly consists of an information perception and processing module and a digital twin system, which communicate via the TCP protocol. In the information perception and processing module, target segmentation and binocular depth reading are first performed based on video signals acquired by an underwater binocular camera to obtain the geometric information of the target object; simultaneously, the Phantom device is used to calculate and issue commands for the underwater robotic arm's remote operation. The processed data is transmitted to the digital twin system via TCP communication. This system specifically includes four core units: a data transmission unit, an environment perception and reconstruction unit, a virtual robotic arm unit, and a human-machine interaction unit. The data transmission unit is responsible for real-time communication with the real environment at the slave end; the environment perception and reconstruction unit uses depth information and binocular video signals to recreate the real underwater operation scene; the virtual robotic arm unit is responsible for driving the virtual robotic arm model to maintain synchronized movement with the real robotic arm; and the human-machine interaction unit provides the operator with immersive visual feedback and an operational interface. In a specific embodiment of the invention, taking the task of grasping a target object by an underwater five-DOF electric robotic arm as an example, the digital twin system in the VR headset provides the operator with an immersive visual experience. The operator can then use a VR controller or a Plantom teleoperation device to operate the robotic arm's grasping task. The system offers two different operating modes: simulation mode and real-world mode. In simulation mode, the operator can control the movement and viewing angle within the headset's field of view using a keyboard or VR controller. The simulation rehearses the robotic arm's grasping and placing of the target object, assisting in experimental rehearsals. In real-world mode, the communication module connects the underwater electric robotic arm and a binocular camera. The digital twin robotic arm replicates the joint movements of the real robotic arm in real time, synchronously performing the grasping and placing of the target object, assisting the operator in accurately judging the robotic arm's spatial posture and reducing operational risks in complex underwater environments.
[0047] like Figure 3 As shown, the real-world operation process mainly consists of three core modules: data acquisition, data processing, and digital twin presentation. First, in the data acquisition module, the operator issues control commands via a teleoperation device to drive the real robotic arm and collects joint motion data of the robotic arm in real time. Simultaneously, an underwater binocular camera acquires image information of the underwater target, generating a target point cloud and calculating the initial pose. Subsequently, the acquired real-time angle data and point cloud data are transmitted to the data processing module via TCP protocol. This module has two core functions: first, it uses a hierarchical multimodal feature constraint algorithm to register the target point cloud, calculate the optimized high-precision pose T, and perform coordinate transformation; second, it establishes a virtual-real synchronous mapping mechanism to handle the data correspondence between the real world and the digital world. Finally, all processed state data is input into the digital twin system. The system uses a base coordinate transformation matrix to map the states of the real robotic arm and the target object to virtual space, driving the virtual robotic arm and the target object model to achieve high-fidelity synchronous mapping and real-time rendering in the Unity3D scene.
[0048] like Figure 4 As shown, the simulation mode is mainly used for pre-playing and training, and its logic begins at the human-computer interaction layer. The operator sets the desired position of the robotic arm's end effector using a keyboard or VR controller. The system uses inverse kinematics to solve for the target angles of each joint and processes the ROV motion control commands in parallel. When performing the mechanical grasping action in the virtual environment, the system introduces a physics engine to perform collision detection and dynamic simulation (including gravity, friction, and collision attributes), accurately calculates the interaction state between the virtual robotic arm and the underwater environment and pipelines, and presents the grasping feedback to the operator in real time, thereby verifying the feasibility of the operation strategy in a risk-free environment.
[0049] In practical implementation, this invention sets up an underwater experimental scenario. A 3D-printed model of an underwater robotic arm and an underwater target object is placed in an experimental pool. Outside the pool are the robotic arm's control board (Jetson NX), the Phantom teleoperation device, and the VIVECOMMOS device (VR headset and controllers). The VR headset displays a digital twin scene with "real-world-simulation" dual-mode switching for underwater operations and maintenance. The VR controllers and Phantom teleoperation device are used to control the joint angles of the robotic arm.
[0050] Before the experiment begins, all experimental equipment must be checked and prepared to ensure it is functioning properly. The specific steps are as follows: 1) Debugging the digital twin system: Connect the VIVE COMMOS device to the computer, put on the helmet and create a room inside, run the exe file of the digital twin system in Unity3D, and you can see the digital twin system in the helmet. Use the controller to control the movement of the ROV and switch the viewpoint to check if all functions are normal.
[0051] Inspection of the underwater electric manipulator: Start the electric manipulator and check the range of motion and flexibility of each joint to ensure it is in good working order. At the same time, check the connections of its drive system and signal acquisition system to ensure they are normal.
[0052] Communication connection test: Start the master PC, run the remote operation control software, ensure that the digital twin system can receive the signals transmitted by the electric robotic arm, check the connection status of the communication network, and ensure stable communication between the master and slave ends.
[0053] To fully verify the performance of this digital twin, the following experimental tasks are designed: In real-world mode, the operator wears a VIVE COMMOS device, adjusts the viewing angle using the handle, and uses the Plantom teleoperation device to control the robotic arm to grab the target object and place it on the pipe.
[0054] In simulation mode (no physical scene is present and no real robotic arm is connected), the operator wears a VR headset and uses a controller to move the virtual robotic arm, simulating the robotic arm grabbing the target object and moving it to a designated position.
[0055] Task 1: Directly operate the real-world mode Task 2: First, use the simulation mode to familiarize yourself with the remote operation process, and then operate the live mode.
[0056] The results of this experiment show that the "simulation first, then practical operation" process for Task 2 demonstrates significant advantages over the Task 1 process, which involves direct real-world operation, in terms of work efficiency, operational accuracy, and cost control. Specifically, these advantages are reflected in: Improved task completion efficiency: The average task completion time for Task 2 is reduced by approximately 40% compared to Task 1. After an average of 3-5 complete rehearsals in simulation mode, operators become more familiar with the process paths and operational nodes during real-world operations, resulting in significantly enhanced action continuity and reduced ineffective movements, thus enabling efficient completion of grasping and placement.
[0057] Improved operational accuracy and success rate: The first attempt success rate for Task 2 reached 90%, far exceeding the 60% for Task 1. The operator pre-optimized the grasping strategy (such as approach angle and clamping force control point) in the virtual environment, avoiding the number of adjustments required due to improper strategies in real-world operations. The positioning accuracy of the robotic arm's end effector was improved by approximately 25% according to calculations.
[0058] This experiment adopted an optimized "simulation first, then real-world operation" process, meaning the simulation mode was executed first, followed by the real-world mode. This arrangement significantly improved the success rate and operational smoothness of the final operation. Operators first practiced the entire process of grasping, placing, and retrieving the target object in a completely virtual environment using VR controllers without risk. This greatly improved their operational proficiency, verified and optimized operational strategies, and simultaneously validated the functionality of the system's core algorithms. Based on this thorough rehearsal, operators then executed the real-world task. Their existing muscle memory and validated optimal strategies could be directly applied to the real environment, making their operations more precise and decisive. This not only significantly reduced the risk of task failure due to operational errors or inappropriate strategies but also effectively saved the expensive operating costs of real underwater equipment, ultimately forming a highly efficient and reliable closed loop of "virtual rehearsal guiding real-world operation."
[0059] The above content is merely a technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.
Claims
1. A method for underwater operation and maintenance based on real-time data using digital twins and dual-mode parallel operation, characterized in that, include: Step 1: Use an underwater robotic arm with an underwater binocular camera mounted at the end to collect observation point cloud of the target object in the underwater scene in real time, and then use a hierarchical multimodal feature constraint registration method and a cross-coordinate system pose mapping method to obtain the three-dimensional spatial pose of the target object; Step 2: Real-time acquisition of angle data of each joint of the underwater robotic arm, and combined with the three-dimensional spatial pose of the target object and the underwater scene, to establish a digital twin scene including the digital twin robotic arm, and to establish a data channel between the real underwater robotic arm and the digital twin robotic arm. Step 3: Based on the data channel between the real underwater robotic arm and the digital twin robotic arm, the underwater robotic arm is controlled to perform precise operations on the target object through two parallel working methods that can be switched between simulation and real-scene control, thus realizing dual-mode parallel operation.
2. The underwater operation and maintenance digital twin and dual-mode parallel operation method based on real-time data as described in claim 1, characterized in that: In the first step, the hierarchical multimodal feature constraint registration method includes an initial alignment layer, a feature constraint optimization layer containing a multimodal feature consistency metric function, and a precise convergence layer. The observation point cloud of the target object and the reference model point cloud are processed by the initial alignment layer to obtain the alignment transformation matrix of the target object. After the alignment transformation matrix is passed through the feature constraint optimization layer and the precise convergence layer, the three-dimensional spatial pose of the target object in the camera coordinate system of the underwater binocular camera is obtained according to the multimodal feature consistency metric function.
3. The real-time data based underwater operational maintenance digital twin and dual mode parallel operation method of claim 2, wherein: The observation point cloud of the target object is a point cloud acquired by an underwater binocular camera in the camera coordinate system, and the reference model point cloud of the target object is a point cloud obtained based on the three-dimensional model of the target object in the model coordinate system. The observation point cloud of the target object is processed through an initial alignment layer to first obtain the feature similarity between each point in the observation point cloud of the target object and the reference model point cloud. Then, point pairs with similarity higher than a preset similarity threshold are selected to form a matching point set. C Based on the matching point set C Construct the alignment transformation matrix of the target object as follows: in, T These are the pose transformation parameters; p i and q j These are the matching point sets. C Observation points in i and reference model points j ; It is the square of the 2-norm.
4. The real-time data based underwater operation and maintenance digital twin and double-mode parallel operation method according to claim 2, characterized in that: In the feature constraint optimization layer, the alignment transformation matrix is... Model feature points are obtained after projection. And serve as a search center, Thus, a radius of [missing information] is constructed. The spherical neighborhood, radius as follows: in, r 0 represents the initial search radius; For the number of iterations, The attenuation coefficient; Spherical neighborhoods of radius are used as adaptive spatial constraint boundaries for the multimodal feature consistency measure function.
5. The real-time data based underwater operation and maintenance digital twin and double-mode parallel operation method according to claim 2, characterized in that: The multi-modal feature consistency measure function In detail as follows: in, T These are the pose transformation parameters; For adaptive weight parameters; and These are the point distance term and the feature preservation term, respectively. p i and q j These are the matching point sets. C Observation points in i and reference model points j ; F ( ) represents the feature extraction function; And through the precise convergence layer setting multimodal feature consistency measure function adopts adaptive step update mechanism to the pose transformation parameters T Iterative solution is carried out, and finally the three-dimensional space pose of the target object in the camera coordinate system of the underwater binocular camera is obtained .
6. The underwater operation and maintenance digital twin and dual-mode parallel operation method based on real-time data as described in claim 5, characterized in that: In the first step, the target object is positioned in three-dimensional space within the camera coordinate system of the underwater binocular camera. The three-dimensional spatial pose of the target object in the base coordinate system of the underwater robotic arm is obtained after processing using a cross-coordinate system pose mapping method. ,as follows: in, It is a rotation matrix; and These are the offset vectors along the y and z axes between the bases of the underwater binocular camera and the underwater robotic arm, respectively.
7. The real-time data based underwater operation and maintenance digital twin and double-mode parallel operation method according to claim 1, characterized in that: In the third step, the digital twin robotic arm is driven to perform operations on the target object in the digital twin scenario, thereby obtaining the real-time joint angles and end-effector poses of the digital twin robotic arm as operation commands. The operation commands are then sent to the underwater robotic arm in real time for control, forming a simulation-real-scene operation process.