Binocular visual servoing method for autonomous operation of underwater vehicle-dual-arm manipulator

Through binocular visual servoing method and Kalman filtering technology, combined with the manipulator end joint controller, the problem of discontinuous target positioning of underwater vehicle-dual-arm manipulator in complex environments is solved, and autonomous and intelligent underwater operations are realized.

CN120182800BActive Publication Date: 2025-09-30HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510563148.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-09-30
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing underwater vehicle-dual-arm manipulators cannot effectively and continuously locate targets in complex underwater environments, affecting the success rate and autonomy of operations.

Method used

The binocular visual servo method is used to obtain two-dimensional images and depth images, combined with the target detector and Kalman filtering method to achieve unbiased estimation of the target's three-dimensional coordinates. The manipulator end joint controller is designed to adapt to the target posture and improve the grasping success rate.

Benefits of technology

Realizing continuous positioning and autonomous grasping of targets in complex underwater environments improves the autonomy and intelligence of underwater operations and enhances the success rate and efficiency of operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182800B_ABST
    Figure CN120182800B_ABST
Patent Text Reader

Abstract

This application belongs to the field of underwater operations and specifically discloses a binocular visual servoing method for autonomous operation of an underwater vehicle-two-arm manipulator. The method includes: obtaining a two-dimensional image and a depth image in front of the underwater vehicle-two-arm manipulator; inputting the two-dimensional image into a lightweight target detector to obtain a target recognition frame; designing a method for obtaining target three-dimensional coordinate observation values ​​based on the recognition frame shape to reduce the impact of parallax holes; solving the device motion matrix based on the three-dimensional coordinates of key points in the images at adjacent moments; using the target three-dimensional coordinates as the state vector, designing a nonlinear Kalman filtering method based on the motion matrix to achieve continuous estimation of the target position; using the target three-dimensional coordinates as a reference, driving the vehicle-manipulator to move, and using the two-dimensional image to drive the rotation of the manipulator's end joints so that after the manipulator reaches the target, its actuator can effectively clamp the target. Through this application, the autonomy and intelligence of underwater operations can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of underwater operations, and more specifically, relates to a binocular visual servoing method for autonomous operation of an underwater vehicle-dual-arm manipulator. Background Art

[0002] With the growing demand for marine resource development, advanced underwater operation technologies are playing an increasingly important role. Underwater vehicle-manipulator systems are widely used in various marine operations, such as seabed mining, aquaculture, deep-sea oil and gas system operations and maintenance, and underwater rescue. As operational scenarios become increasingly complex, the demand for underwater vehicle-manipulator systems to achieve environmental awareness and operational autonomy is also increasing.

[0003] However, in practical applications, various noise interferences in complex underwater environments affect target detection based on visible light cameras and binocular stereo matching, posing a significant challenge to continuous target localization. Furthermore, the pose of underwater targets is highly random, and failure to adjust the grasping angle accordingly will significantly impact the success rate of the operation. Therefore, to improve the autonomy and efficiency of underwater operations, advanced visual servoing methods are needed. For example, industrial vision technology can be used to achieve real-time monitoring and precise measurement of the underwater environment, providing accurate environmental information for the operation of underwater vehicles and manipulators. Simultaneously, industrial control software can process and analyze this visual data, generating control commands to achieve precise control of the underwater vehicle-arm manipulator, thereby improving the automation level and efficiency of underwater operations. This combination not only improves the safety and reliability of underwater operations but also expands their application areas, ensuring the adaptability of underwater vehicle-manipulator systems to complex tasks. Summary of the Invention

[0004] In view of the defects of the existing technology, the purpose of this application is to provide a binocular visual servoing method for autonomous operation of an underwater vehicle-dual-arm manipulator, aiming to solve the problem that the existing underwater vehicle-dual-arm manipulator is unable to continuously locate the target in a complex underwater environment.

[0005] To achieve the above objectives, in a first aspect, the present application provides a binocular visual servoing method for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, and the method comprises:

[0006] Acquire a two-dimensional image and a depth image of the target in front of the underwater vehicle-dual-arm manipulator; input the two-dimensional image into a target detector to obtain a target recognition frame; then, determine a three-dimensional coordinate search area on the depth image based on the target recognition frame, remove pixels with invalid depth values ​​in the search area based on the depth image, and retain pixels with intermediate depth values ​​by sorting them according to depth value, thereby determining the three-dimensional coordinate observation value of the target;

[0007] Determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at the previous and next moments;

[0008] The three-dimensional coordinates and three-dimensional velocity of the target are used as the state vector, the three-dimensional coordinate observation value of the target is used as the observation vector, and the three-dimensional coordinate prediction value of the target is obtained in combination with the motion parameters. The three-dimensional coordinate observation value and the prediction value are fused based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target;

[0009] The device is driven to move to a preset position with an unbiased estimate of the target's three-dimensional coordinates as a reference; then, the actuator of at least one manipulator is controlled to move toward the target based on the unbiased estimate of the target's three-dimensional coordinates, and the end joint of the corresponding actuator is driven to rotate using a two-dimensional image of the target, so that after the manipulator reaches the target, its actuator effectively clamps the target.

[0010] It should be noted that in some application scenarios, the above-mentioned target detector can also obtain the target category, and then, based on the target category, it can be determined whether to control the actuators of one or both manipulators to move toward the target. This application does not specifically limit this, and those skilled in the art can design it as needed.

[0011] For example, the depth image is obtained through stereo matching using a binocular camera.

[0012] It can be understood that the present application inputs a two-dimensional image into a lightweight target detector to obtain a target recognition frame; designs a method for obtaining target three-dimensional coordinate observation values ​​based on the shape of the recognition frame to reduce the impact of parallax holes; solves the device motion matrix based on the three-dimensional coordinates of key points in the images at adjacent moments; uses the target three-dimensional coordinates as the state vector, and designs a nonlinear Kalman filtering method based on the motion matrix to achieve unbiased estimation of the target position.

[0013] Specifically, after acquiring the target frame, the present application combines the depth image acquired by binocular vision to obtain the three-dimensional coordinate observation value of the target. Compared with the prior art that directly selects the three-dimensional coordinates of the center point of the target frame when acquiring the target frame, the present application combines the depth information of the pixels around the center of the target frame, eliminates invalid values ​​and outliers, and then jointly calculates the three-dimensional coordinates of the target. The calculated coordinate values ​​are more accurate. In addition, the Kalman filter method combined with the hull motion estimation makes the positioning of the target more continuous. It effectively solves the problems of discontinuous and uneven target positioning caused by parallax holes, detector misses, hull shaking, etc. In addition, when controlling the manipulator, the present application also adjusts the angle of the manipulator end joint in combination with the target image. Compared with the prior art that does not consider the posture relationship between the manipulator end joint and the target, it can greatly improve the probability of effectively clamping the target and ensure the success rate of underwater operations. Furthermore, the present application can obtain the target type based on the target detector, and then control the operation of a single manipulator or dual manipulators in combination with the target type, with high flexibility.

[0014] The visual servoing method designed in this application can enable underwater vehicles and manipulators to complete a variety of autonomous operation tasks, and has the ability to adapt to the randomness of the target posture. It can also be expanded to dual-arm operation tasks, which can effectively improve operation efficiency and success rate.

[0015] In one embodiment, the object detector includes: a feature extractor, an encoder, and a decoder;

[0016] The feature extractor is used to extract the features of the target in the two-dimensional image;

[0017] The encoder is used to adjust the features extracted by the feature extractor to the same scale and splice them in the channel dimension to obtain features. ; Through convolution of features Perform dimensionality reduction to obtain features ; Use fast Fourier transform to transform features Convert to frequency domain, corresponding to spectrum ; Adjust the spectrum through learnable filters The information in the image is converted back to the time domain using the inverse fast Fourier transform to obtain new features. The new features are sequentially input into the first multi-layer perceptron, the deformable convolutional network, and the second multi-layer perceptron for processing to obtain the features enhanced by the encoder.

[0018] The decoder is used to obtain a target recognition frame through depth-wise separable convolution.

[0019] In one embodiment, determining the three-dimensional coordinate observation value of the target includes:

[0020] The center point of the target recognition frame is used as the center of the search area on the depth image, and the width and height of the search area are determined according to the size of the target recognition frame;

[0021] Match the depth image with the two-dimensional image based on pixel coordinates to obtain the depth value of each pixel in the search area;

[0022] Sort all pixels in the search area according to their depth values;

[0023] Eliminate invalid pixels in the search area based on a preset sorting threshold to obtain a new pixel set;

[0024] Keep the pixel points in the middle of the new pixel set;

[0025] The three-dimensional coordinates of all retained pixels are averaged to obtain the three-dimensional coordinate observation value of the target.

[0026] In one embodiment, retaining pixels with middle depth values ​​in the new pixel set includes:

[0027] Calculate the upper quartile of the sorted set of new pixels and lower quartile :

[0028]

[0029] in, n is the total number of pixels included in the new pixel set;

[0030] if and If is an integer, take and The corresponding position depth value is used as the threshold and ;

[0031] if and If it is not an integer, round it up and down to get , , , , and take the four depth values ​​of the corresponding positions , , , ,

[0032] The depth value satisfy The pixels are retained.

[0033] In one embodiment, the motion parameters of the device are obtained by combining the three-dimensional coordinates of key points at two preceding and following moments, including:

[0034] The first i Error function of key points Set to:

[0035]

[0036] in, Indicates the current moment i The three-dimensional coordinates of the key points, Indicates the last moment i The three-dimensional coordinates of the key points, R Represents the device's three-dimensional motion rotation matrix, Represents the device translation vector; and forming the motion parameters;

[0037] With the goal of minimizing the sum of squares of the error functions of all key points, solve the and .

[0038] In one embodiment, obtaining the predicted three-dimensional coordinates of the target in combination with the motion parameters includes:

[0039] According to the target k State vector at time -1 , through the state transfer matrix Get the prediction result of the state vector at time k :

[0040]

[0041] Among them, the state vector ; ; t Indicates the time interval between the previous and next moments;

[0042] The goal is k The three-dimensional coordinate prediction value at the time ; Motion parameters and Combined k -1 and k The three-dimensional coordinates of the key points at that moment determine the device's three-dimensional motion rotation matrix and device translation vector.

[0043] In one embodiment, the unbiased estimate of the target's three-dimensional coordinates is for:

[0044]

[0045] in, is the predicted value of the three-dimensional coordinate, is the three-dimensional coordinate observation value; H is the observation transformation matrix, ; is a filter gain matrix; and / or

[0046]

[0047]

[0048]

[0049] in, for k -1 moment covariance matrix; is the state transfer matrix; and All represent noise; For combination k -1 and k The 3D motion rotation matrix of the device determined by the 3D coordinates of the key points at the moment; For prediction k The covariance matrix of the moments; After considering the motion parameters k Moment covariance matrix.

[0050] In one embodiment, using a two-dimensional image of a target to drive the end joint of the manipulator to rotate includes:

[0051] The two-dimensional image is input into a network model, and the rotation angle of the end joint is output; the network model is trained based on a sample group consisting of a determined target two-dimensional image and the rotation angle of the end joint, or the network model is learned by a mapping relationship between a preset target two-dimensional image and the rotation angle of the end joint; and / or

[0052] The network model includes: a convolutional neural network, a global average pooling layer, and a fully connected layer; the convolutional neural network includes channel-by-channel convolution or stationary point convolution to extract high-dimensional features from a two-dimensional image, the global average pooling layer is used to fuse the high-dimensional features in the spatial dimension, and the fully connected layer includes an output node for predicting the rotation angle of the end joint based on the fused high-dimensional features;

[0053] The loss function of the network model for:

[0054]

[0055] in, N represents the number of samples; Indicates the desired rotation angle; represents the rotation angle predicted by the network model, Represents a two-dimensional image.

[0056] In a second aspect, the present application provides a binocular visual servo system for autonomous operation of an underwater vehicle and a dual-arm manipulator. The underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations. The system includes:

[0057] An image acquisition unit, used to acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator;

[0058] The target observation unit is configured to input the two-dimensional image into a target detector to obtain a target recognition frame; then determine a three-dimensional coordinate search area on the depth image based on the target recognition frame, remove pixels with invalid depth values ​​in the search area in combination with the depth image, and retain pixels with intermediate depth values ​​in order of depth value, thereby determining the three-dimensional coordinate observation value of the target;

[0059] a target prediction unit for determining the three-dimensional coordinates of a plurality of key points on a two-dimensional image with reference to the depth image; obtaining motion parameters of the device by combining the three-dimensional coordinates of the key points at two preceding and succeeding moments; and obtaining a predicted three-dimensional coordinate value of the target by combining the motion parameters with the three-dimensional coordinates and three-dimensional velocity of the target as a state vector and the three-dimensional coordinate observation value of the target as an observation vector;

[0060] An unbiased estimation unit is used to fuse the three-dimensional coordinate observation value and the predicted value based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target;

[0061] A servo control unit is used to drive the device to a preset position with an unbiased estimate of the target's three-dimensional coordinates as a reference; then control the actuator of at least one manipulator to move toward the target based on the unbiased estimate of the target's three-dimensional coordinates, and use a two-dimensional image of the target to drive the end joint of the corresponding actuator to rotate, so that after the manipulator reaches the target, its actuator can effectively clamp the target.

[0062] In a third aspect, the present application provides an electronic device comprising: at least one memory for storing programs; and at least one processor for executing the programs stored in the memory. When the programs stored in the memory are executed, the processor is used to execute the method described in the first aspect or any one of the embodiments of the first aspect.

[0063] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program. When the computer program runs on a processor, the processor executes the method described in the first aspect or any one of the embodiments of the first aspect.

[0064] In a fifth aspect, the present application provides a computer program product. When the computer program product runs on a processor, it enables the processor to execute the method described in the first aspect or any embodiment of the first aspect.

[0065] In general, the above technical solutions conceived by this application have the following beneficial effects compared with the existing technologies:

[0066] The present application provides a binocular visual servoing method for autonomous operation of an underwater vehicle and a dual-arm manipulator, designs a lightweight deep neural network target detector to achieve rapid identification of the operation target; designs a continuous estimation method for the target three-dimensional coordinates. Compared with the existing target positioning method, the continuous positioning method disclosed in the present application will not lose the target position in cases of parallax holes, missed detection, etc., and at the same time considers the influence of the nonlinear motion of the hull on positioning to improve robustness; then, the present application designs a decoupled visual servoing method, divides the operation execution process of the underwater vehicle and manipulator into dynamic positioning of the hull based on the target position and target grasping of the manipulator based on the target position and image, and designs a neural network-based manipulator end joint controller to adjust the grasping angle, which can effectively solve the problem of autonomous operation under unknown target posture, effectively improve the autonomy and intelligence of underwater operations, and improve the success rate of operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] Figure 1 Flowchart of the binocular visual servoing method for autonomous operation of an underwater vehicle-dual-arm manipulator provided in an embodiment of the present application;

[0068] Figure 2 A schematic diagram of an underwater vehicle-dual-arm manipulator device provided in an embodiment of the present application;

[0069] Figure 3 A block diagram of decoupled visual servoing provided in an embodiment of the present application;

[0070] Figure 4 A block diagram of a lightweight deep neural network target detector provided in an embodiment of the present application;

[0071] Figure 5 A block diagram of a neural network-based robotic arm end joint controller according to an embodiment of the present application;

[0072] Figure 6 A three-dimensional coordinate curve diagram of the target during the autonomous valve switching operation provided in the embodiment of the present application;

[0073] Figure 7 A graph showing the position changes of the joints and end positions of the robotic arm during the autonomous valve opening and closing operation provided in an embodiment of the present application;

[0074] Figure 8A flowchart showing the completion of a first-person perspective task during an autonomous valve switching operation provided in an embodiment of the present application;

[0075] Figure 9 A three-dimensional coordinate curve diagram of the grabbing points during the autonomous pipeline handling operation provided in the embodiment of the present application;

[0076] Figure 10 This is a graph showing the position changes of the joints and end points of the dual robotic arms during the autonomous pipe handling operation provided by an embodiment of the present application;

[0077] Figure 11 A flowchart showing the completion of a first-person perspective task during an autonomous pipeline handling operation provided in an embodiment of the present application;

[0078] Figure 12 This is a diagram of the architecture of a binocular visual servo system for autonomous operation of an underwater vehicle and a dual-arm manipulator provided in an embodiment of the present application;

[0079] Figure 13 This is a diagram of the electronic device architecture provided in an embodiment of the present application. DETAILED DESCRIPTION

[0080] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0081] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.

[0082] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0083] In the description of the embodiments of the present application, unless otherwise specified, “multiple” means two or more than two. For example, multiple key points means two or more than two key points.

[0084] The embodiments of the present application are described below in conjunction with the drawings in the embodiments of the present application.

[0085] Figure 1Flowchart of the binocular visual servoing method for autonomous operation of an underwater vehicle-dual-arm manipulator provided in an embodiment of the present application; Figure 1 As shown, the following steps are included:

[0086] Step S101: Acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator; input the two-dimensional image into a target detector to obtain a target recognition frame; then, determine a three-dimensional coordinate search area on the depth image based on the target recognition frame, remove pixels with invalid depth values ​​in the search area based on the depth image, and retain pixels with intermediate depth values ​​in order of depth value, thereby determining the three-dimensional coordinate observation value of the target;

[0087] Step S102, determining the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; combining the three-dimensional coordinates of the key points at two previous and subsequent moments to obtain motion parameters of the device;

[0088] Step S103, using the target's three-dimensional coordinates and three-dimensional velocity as a state vector, and the target's three-dimensional coordinate observation value as an observation vector, combined with the motion parameters to obtain a three-dimensional coordinate prediction value of the target, and fusing the three-dimensional coordinate observation value and the prediction value based on the Kalman filter method to obtain an unbiased estimate of the target's three-dimensional coordinate;

[0089] In step S104, the device is driven to move to a preset position with reference to the unbiased estimate of the three-dimensional coordinates of the target; then, the actuator of at least one manipulator is controlled to move toward the target according to the unbiased estimate of the three-dimensional coordinates of the target, and the end joint of the corresponding actuator is driven to rotate using the two-dimensional image of the target, so that after the manipulator reaches the target, its actuator effectively clamps the target.

[0090] The purpose of this application is to design a visual servoing method to enable autonomous operation of an underwater vehicle-manipulator system. This method uses a visible light camera to continuously locate targets in low-light conditions and complex water conditions. Furthermore, this visual servoing method is adaptable to targets of varying types and postures, ensuring adaptive adjustment of the grasping angle, thereby achieving autonomy and intelligence in the underwater vehicle-manipulator system.

[0091] In a more specific embodiment, the above method may include the following steps:

[0092] Obtain images of the underwater vehicle and manipulator system at the previous moment and the current moment obtained through the forward-looking camera.

[0093] The structure of the underwater vehicle-manipulator system is shown in Figure 2 shown.

[0094] Design a lightweight underwater target detector based on deep neural network to quickly obtain the type and pixel coordinates of each target contained in the current image;

[0095] The depth map of the field of view is generated from the binocular image using the internal and external parameters of the camera. Based on the target shape and pixel coordinates, the 3D information search range of the target on the depth map is adaptively calculated. Invalid values ​​and outliers are removed, and the remaining valid values ​​are averaged to obtain the target 3D coordinate observation value at the current moment.

[0096] Detect key points on the image at the previous moment and use optical flow estimation to obtain the pixel coordinates of the corresponding points on the current image. Based on this, obtain the corresponding 3D coordinates of the key points in the depth maps of the previous and next frames, and calculate the 3D motion matrix of the system at two consecutive moments based on least squares;

[0097] The target's three-dimensional coordinates at the current moment are preliminarily estimated using a linear motion model, and the predicted three-dimensional coordinates of the target at the current moment are further calculated based on the motion matrix and this value. The observed and predicted three-dimensional coordinates of the target are fused using the Kalman gain to achieve an unbiased estimate of the target's three-dimensional coordinates in the onboard coordinate system at the current moment.

[0098] Based on the current three-dimensional coordinates of the target and the expected three-dimensional coordinates of the target, the dynamic positioning error of the underwater vehicle and the manipulator system is obtained, and it is controlled to navigate to the expected position; the current three-dimensional coordinates of the target are set as the expected position of the manipulator end, and the expected angles of each joint of the manipulator are calculated based on the kinematic model and the difference between the current angles is used as the control quantity to control the movement of the manipulator; a manipulator end joint controller based on a deep neural network is designed, which takes the target image as input, outputs the rotation angle, and controls the grasping angle of the end effector to adapt to the randomness of the target posture.

[0099] Specifically, the decoupled visual servoing method for the autonomous operation task of the underwater vehicle-manipulator system disclosed in this application is described in detail in the following sections. Figure 3 As shown, including the following:

[0100] (1) A lightweight deep neural network was designed to quickly identify targets in images; (2) A method for obtaining the three-dimensional coordinate values ​​of targets based on binocular stereo vision was designed; (3) A Kalman filter method based on camera nonlinear motion compensation was proposed to achieve continuous target positioning; (4) A decoupling scheme was designed to convert target position and image into control instructions.

[0101] The structure of the designed lightweight underwater target detector is as follows: Figure 4As shown, it includes: a lightweight feature extractor for extracting high-level features from the image, an encoder combining global frequency domain filtering and deformable convolution for enhancing the robustness of the model, and a decoder based on depthwise separable convolution for quickly outputting the target type and its pixel coordinates from the features output by the encoder.

[0102] Specifically, in order to enhance the model's expressiveness and adaptability to targets of different sizes, the present embodiment designs an encoder based on frequency domain filtering and deformable convolution. The features extracted by the feature extractor are adjusted to the same scale and spliced ​​in the channel dimension to obtain the feature By convolution Perform dimensionality reduction to obtain . Using Fast Fourier Transform Convert to frequency domain and get its spectrum . Through a learnable filter To adjust the information in the spectrum:

[0103]

[0104] in is an element-by-element multiplication. Then, it is converted back to the time domain using the inverse fast Fourier transform to obtain the new .

[0105] The newly acquired features are passed through a multi-layer perceptron to further enhance the nonlinear expression ability of the model.

[0106] The features processed by frequency domain filtering are then processed by a deformable convolution to improve the representation of objects of different sizes. Similarly, the processed features are passed through a multi-layer perceptron to improve the model's nonlinear representation capabilities.

[0107] The features processed by the encoder are fed into a decoder consisting of depth-wise separable convolutions to predict the type of target and the center pixel coordinates and size of the recognition box pixel by pixel.

[0108] Furthermore, the specific process of the proposed method for obtaining the three-dimensional coordinate value of the target based on binocular stereo vision is as follows:

[0109] The pixel coordinates of the target center are calculated based on the target recognition frame obtained by the detector as the center of the search area, and then the width and height of the search area are calculated based on the size of the target recognition frame. The corresponding three-dimensional coordinates of each pixel in the area on the depth map are extracted to form a set . Use a larger threshold to remove invalid values, that is, if Then delete it directly and get a new collection .

[0110] Calculate the upper quartile of a set and lower quartile :

[0111]

[0112] Where n is the size of the set. and If it is an integer, the value of the corresponding position is taken as the threshold and At this point, only the three-dimensional information that meets the following conditions needs to be retained:

[0113]

[0114] if and If it is not an integer, it is rounded up and down, and the four values ​​at the corresponding positions are taken. , , , , the threshold calculation formula at this time is:

[0115]

[0116] Get a new collection The mean of the effective values ​​is taken as the three-dimensional coordinate observation value of the target at the current moment:

[0117]

[0118] The process of the proposed Kalman filter method based on camera nonlinear motion compensation is as follows:

[0119] First, we need to estimate the nonlinear motion matrix of the boat between two consecutive frames:

[0120] Detect Shi-Tomasi corner points on the image at the previous moment and obtain the three-dimensional coordinates of these key points at the corresponding positions on the depth map . Use optical flow estimation to find the corresponding key points on the current image and obtain the three-dimensional coordinates of these key points at the corresponding positions on the current depth map. . Define the following error function:

[0121]

[0122] According to the above error, the following least squares problem is defined:

[0123]

[0124] By solving the above problem, the three-dimensional motion rotation matrix of the underwater vehicle and manipulator system between two frames can be obtained. and the translation vector .

[0125] Secondly, define the state vector of the target three-dimensional coordinates and the observation vector They are:

[0126]

[0127] Define the state transition matrix and the observation transformation matrix for:

[0128]

[0129] Finally, the unbiased estimation process of the target three-dimensional coordinates is as follows: k -1 moment status , obtained through the state transition matrix k Prediction results of the state vector at the moment :

[0130]

[0131] in, = ; The state variable at the previous moment is determined by the unbiased estimate at the previous moment and the movement speed of the device.

[0132] According to the covariance matrix of the target at the previous moment , predict the covariance matrix at the current moment through the state transfer matrix :

[0133]

[0134] The rotation matrix in the nonlinear motion estimation result of the hull and translation vectors At the same time, the prediction result of the current state vector is added middle:

[0135]

[0136] The rotation matrix in the nonlinear motion estimation result of the hull Add to the prediction result of the current moment covariance matrix middle:

[0137]

[0138] Through the predicted covariance matrix , observation matrix Calculate the filter gain matrix :

[0139]

[0140] By filtering the gain matrix Fusion k Prediction results at the moment and observation results Get an unbiased estimate of the current target's three-dimensional coordinates :

[0141]

[0142] Update the defense difference matrix :

[0143]

[0144] The decoupling scheme for converting target position and image into control instructions divides the autonomous operation of the underwater vehicle-manipulator system into two parts:

[0145] In the first part, the hull is set to the dynamic positioning task and the three-dimensional coordinates of the target in the hull coordinate system at the current moment are calculated. With the expected coordinates Interpolation As a control error, it is input into the underlying controller to drive the underwater vehicle to approach the target and keep it stationary. Here, approaching means that the distance between the device and the target is less than a preset distance threshold, such as 0.5m.

[0146] The second part is to convert the three-dimensional coordinates of the target in the camera coordinate system Converted to the base coordinate system of the manipulator , the expected value of the end position of the robot Set to The relationship between the target image and the end joint rotation angle is encoded as a learnable mapping A deep neural network To learn:

[0147]

[0148] in, Indicates the rotation angle, represents the target image, Represents the network weight.

[0149] Furthermore, the above-mentioned manipulator end joint controller based on deep neural network has a specific form, such as Figure 5As shown, the network consists of a convolutional neural network, a global average pooling layer, and a fully connected layer. The convolutional neural network uses channel-by-channel or point-by-point convolution to extract high-dimensional features from the target image. Global average pooling is used to fuse high-dimensional features in the spatial dimension. The fully connected layer contains an output node that predicts the rotation angle based on the fused features.

[0150] This application uses the following examples to verify the above technical solution:

[0151] The basic parameters of the underwater vehicle-manipulator system used in this embodiment are shown in Table 1.

[0152] Table 1 Basic parameters of underwater vehicle-manipulator system

[0153]

[0154] This embodiment uses a water tank test to illustrate the effectiveness and advancement of the method proposed in this application. In the test, the underwater vehicle-manipulator will complete the autonomous opening and closing tasks of various valves and the autonomous pipe handling tasks of the dual arms to demonstrate the effectiveness of the control method proposed in this application. The following describes the relevant parameter settings. The initial distance between the underwater vehicle-manipulator and the operating target is a random number between [1m, 2m]. When the hull is performing the dynamic positioning task, the expected position of the target in the hull coordinate system is [0, 0, 0.5m], with an error of 20cm. The error between the end position of the manipulator and the expected position must be less than 5cm.

[0155] The test results are as follows Figures 6 to 11 shown. Figure 6 This is the curve of the target position in the camera coordinate system as different valve switching tasks are executed. Figure 7 The curves for the rotation angles of each joint of the robotic arm and the end position in the base coordinate system are shown in Figure 2. The experimental results show that the discontinuity in target positioning is significantly overcome when using this method. The robotic arm's joints rotate according to the target's position and image, bringing the end position close to the valve. Figure 8 It is the first-person perspective image during the task execution. Figure 8 It can be concluded that when the method of the present application is adopted, the underwater vehicle-manipulator can autonomously complete the switching operations of different valves in random postures.

[0156] Figure 9 This is the position change curve of the target grabbing point of the pipeline during the operation process. Figure 10The curves for the rotation angles of the dual manipulator joints and the end position in the base coordinate system are shown in Figure 2. The test results show that when using this method, the hull successfully approaches the target and performs dynamic positioning. The joints of the dual manipulators rotate based on the target's position and image, bringing the end position close to the pipeline. Figure 11 Displays first-person perspective images during task execution. Figure 11 It can be concluded that when the method provided in the embodiment of the present application is adopted, the underwater vehicle-manipulator can use both hands to complete the grasping of the pipeline.

[0157] In summary, this application discloses a binocular visual servoing method for autonomous operation of an underwater vehicle and a dual-arm manipulator. Under the influence of poor underwater lighting conditions, complex interference, and the inherent swaying of the hull, the underwater vehicle-manipulator system faces significant challenges in autonomously completing operations using its own sensing equipment. To ensure rapid target detection, a lightweight deep neural network target detector was designed. To avoid losing the target position due to parallax holes and missed detections, a continuous three-dimensional target coordinate estimation method was designed. The impact of the nonlinear motion of the hull on positioning was also considered to improve robustness. Subsequently, this application designed a scheme for converting the target position and image into control instructions. This scheme divides the operation execution process of the underwater vehicle and manipulator into dynamic positioning of the hull based on the target position and target grasping by the manipulator based on the target position and image. A neural network-based manipulator end-joint controller was designed to adjust the grasping angle. This effectively solves the problem of autonomous operation under unknown target posture and improves the success rate of operations.

[0158] Figure 12 The architecture diagram of the binocular visual servo system for autonomous operation of the underwater vehicle-dual-arm manipulator provided in the embodiment of the present application; Figure 12 Shown, including:

[0159] The image acquisition unit 1210 is used to acquire a two-dimensional image and a depth image of the target in front of the underwater vehicle-dual-arm manipulator;

[0160] The target observation unit 1220 is configured to input the two-dimensional image into a target detector to obtain a target recognition frame; then, based on the target recognition frame, determine a three-dimensional coordinate search area on the depth image, remove pixels with invalid depth values ​​in the search area in combination with the depth image, and retain pixels with intermediate depth values ​​by sorting the depth values, thereby determining the three-dimensional coordinate observation value of the target;

[0161] The target prediction unit 1230 is configured to determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; obtain motion parameters of the device by combining the three-dimensional coordinates of the key points at two preceding and succeeding moments; and obtain a predicted three-dimensional coordinate value of the target by combining the motion parameters with the three-dimensional coordinates and three-dimensional velocity of the target as a state vector and the three-dimensional coordinate observation value of the target as an observation vector.

[0162] An unbiased estimation unit 1240 is configured to fuse the three-dimensional coordinate observations and predictions based on a Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinates of the target;

[0163] The servo control unit 1250 is used to drive the device to a preset position with reference to the unbiased estimate of the three-dimensional coordinates of the target; then control the actuator of at least one manipulator to move toward the target according to the unbiased estimate of the three-dimensional coordinates of the target, and use the two-dimensional image of the target to drive the end joint of the corresponding actuator to rotate, so that after the manipulator reaches the target, its actuator can effectively clamp the target.

[0164] It should be understood that the above-mentioned system is used to execute the method in the above-mentioned embodiment. The implementation principle and technical effect of the corresponding program unit in the system are similar to those described in the above-mentioned method. The working process of the system can refer to the corresponding process in the above-mentioned method and will not be repeated here.

[0165] Based on the method in the above embodiment, the embodiment of the present application provides an electronic device, such as Figure 13 As shown, the electronic device may include: a processor 1310, a communication interface 1320, a memory 1330, and a communication bus 1340, wherein the processor 1310, the communication interface 1320, and the memory 1330 communicate with each other via the communication bus 1340. The processor 1310 may call the logic instructions in the memory 1330 to execute the method in the above embodiment.

[0166] In addition, the logic instructions in the aforementioned memory 1330 can be implemented in the form of a software functional unit and, when sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application.

[0167] Based on the method in the above embodiment, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program runs on a processor, the processor executes the method in the above embodiment.

[0168] Based on the method in the above embodiment, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor executes the method in the above embodiment.

[0169] It is understood that the processor in the embodiments of the present application may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA), other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0170] The method steps in the embodiments of the present application can be implemented by hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC.

[0171] The above embodiments can be implemented in whole or in part using software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions. When loaded and executed on a computer, the computer program instructions fully or partially produce the processes or functions described in the embodiments of this application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted via the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., solid-state drive (SSD)).

[0172] It will be understood that the various numerical numbers involved in the embodiments of the present application are merely distinctions for the convenience of description and are not intended to limit the scope of the embodiments of the present application.

[0173] It is easy for those skilled in the art to understand that the above is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A binocular visual servoing method for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, characterized in that: Methods include: Acquire a two-dimensional image and depth image of the target in front of the underwater vehicle-dual-arm manipulator; Inputting the two-dimensional image into a target detector to obtain a target recognition frame; Then, the three-dimensional coordinate search area on the depth image is determined based on the target recognition frame, and the pixels with invalid depth values ​​in the search area are removed in combination with the depth image, and the pixels with middle depth values ​​are retained according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; Determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; and obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at the previous and next moments. The three-dimensional coordinates and three-dimensional velocity of the target are used as the state vector, the three-dimensional coordinate observation value of the target is used as the observation vector, and the three-dimensional coordinate prediction value of the target is obtained in combination with the motion parameters. The three-dimensional coordinate observation value and the prediction value are fused based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target; Using an unbiased estimate of the target's three-dimensional coordinates as a reference, the device is driven to move to a preset position; then, based on the unbiased estimate of the target's three-dimensional coordinates, at least one actuator of the manipulator is controlled to move toward the target, and a two-dimensional image of the target is used to drive the end joint of the corresponding actuator to rotate, so that after the manipulator reaches the target, its actuator effectively grasps the target; Determine the three-dimensional coordinate observation values ​​of the target, including: The center point of the target recognition frame is used as the center of the search area on the depth image, and the width and height of the search area are determined according to the size of the target recognition frame; Match the depth image with the two-dimensional image based on pixel coordinates to obtain the depth value of each pixel in the search area; Sort all pixels in the search area according to their depth values; Eliminate invalid pixels in the search area based on a preset sorting threshold to obtain a new pixel set; Keep the pixel points in the middle of the new pixel set; The three-dimensional coordinates of all retained pixels are averaged to obtain the three-dimensional coordinate observation value of the target; Keep the pixel with the middle depth value in the new pixel set, including: Calculate the upper quartile of the sorted set of new pixels and lower quartile : in, n is the total number of pixels included in the new pixel set; if and If is an integer, take and The corresponding position depth value is used as the threshold and ; if and If it is not an integer, round it up and down to get , , , , and take the four depth values ​​of the corresponding positions , , , , The depth value satisfy The pixels are retained.

2. The method according to claim 1, characterized in that The target detector includes: a feature extractor, an encoder and a decoder; The feature extractor is used to extract the features of the target in the two-dimensional image; The encoder is used to adjust the features extracted by the feature extractor to the same scale and splice them in the channel dimension to obtain features. ; Through convolution of features Perform dimensionality reduction to obtain features ; Use fast Fourier transform to transform features Convert to frequency domain, corresponding to spectrum ; Adjust the spectrum through learnable filters The information in the image is converted back to the time domain using the inverse fast Fourier transform to obtain new features. The new features are sequentially input into the first multi-layer perceptron, the deformable convolutional network, and the second multi-layer perceptron for processing to obtain the features enhanced by the encoder. The decoder is used to obtain a target recognition frame through depth-wise separable convolution.

3. The method according to claim 1, characterized in that Combine the three-dimensional coordinates of the key points at the previous and next moments to obtain the motion parameters of the device, including: The first i Error function of key points Set to: in, Indicates the current moment i The three-dimensional coordinates of the key points, Indicates the last moment i The three-dimensional coordinates of the key points, R Represents the device's three-dimensional motion rotation matrix, Represents the device translation vector; With the goal of minimizing the sum of squares of the error functions of all key points, solve the and .

4. The method according to claim 1, wherein Acquiring a predicted three-dimensional coordinate value of the target in combination with the motion parameters includes: According to the target k State vector at time -1 , through the state transfer matrix Get the prediction result of the state vector at time k : Among them, the state vector ; ; t Indicates the time interval between the previous and next moments; The goal is k The three-dimensional coordinate prediction value at the time ; Motion parameters and Combined k -1 and k The three-dimensional coordinates of the key points at that moment determine the device's three-dimensional motion rotation matrix and device translation vector.

5. The method according to claim 1, wherein Unbiased estimate of the target's three-dimensional coordinates for: in, is the predicted value of the three-dimensional coordinate, is the three-dimensional coordinate observation value; H is the observation transformation matrix, ; is a filter gain matrix; and / or in, for k -1 moment covariance matrix; is the state transfer matrix; and All represent noise; For combination k -1 and k The 3D motion rotation matrix of the device determined by the 3D coordinates of the key points at the moment; For prediction k The covariance matrix of the moments; After considering the motion parameters k Moment covariance matrix.

6. The method according to claim 1, characterized in that The two-dimensional image of the target is used to drive the end joint rotation of the manipulator, including: The two-dimensional image is input into a network model, and the rotation angle of the end joint is output; the network model is trained based on a sample group consisting of a determined target two-dimensional image and the rotation angle of the end joint, or the network model is learned by a mapping relationship between a preset target two-dimensional image and the rotation angle of the end joint; and / or The network model includes: a convolutional neural network, a global average pooling layer, and a fully connected layer; the convolutional neural network includes channel-by-channel convolution or point-by-point convolution to extract high-dimensional features from a two-dimensional image, the global average pooling layer is used to fuse the high-dimensional features in the spatial dimension, and the fully connected layer includes an output node for predicting the rotation angle of the end joint based on the fused high-dimensional features; The loss function of the network model for: in, N represents the number of samples; Indicates the desired rotation angle; represents the rotation angle predicted by the network model, Represents a two-dimensional image.

7. A binocular visual servo system for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, characterized in that: The system includes: An image acquisition unit, used to acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator; The target observation unit is configured to input the two-dimensional image into a target detector to obtain a target recognition frame; then determine a three-dimensional coordinate search area on the depth image based on the target recognition frame, remove pixels with invalid depth values ​​in the search area in combination with the depth image, and retain pixels with intermediate depth values ​​in order of depth value, thereby determining the three-dimensional coordinate observation value of the target; The target prediction unit is used to determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; and to obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two moments before and after; and obtaining a predicted three-dimensional coordinate value of the target by combining the three-dimensional coordinates and three-dimensional velocity of the target as a state vector and the three-dimensional coordinate observation value of the target as an observation vector, in combination with the motion parameters; An unbiased estimation unit is used to fuse the three-dimensional coordinate observation value and the predicted value based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target; a servo control unit configured to drive the device to a preset position using an unbiased estimate of the target's three-dimensional coordinates as a reference; thereafter controlling at least one actuator of the manipulator to move toward the target based on the unbiased estimate of the target's three-dimensional coordinates; and using a two-dimensional image of the target to drive the rotation of the end joint of the corresponding actuator so that the actuator effectively grips the target after the manipulator reaches the target; Determine the three-dimensional coordinate observation values ​​of the target, including: The center point of the target recognition frame is used as the center of the search area on the depth image, and the width and height of the search area are determined according to the size of the target recognition frame; Match the depth image with the two-dimensional image based on pixel coordinates to obtain the depth value of each pixel in the search area; Sort all pixels in the search area according to their depth values; Eliminate invalid pixels in the search area based on a preset sorting threshold to obtain a new pixel set; Keep the pixel points in the middle of the new pixel set; The three-dimensional coordinates of all retained pixels are averaged to obtain the three-dimensional coordinate observation value of the target; Keep the pixel with the middle depth value in the new pixel set, including: Calculate the upper quartile of the sorted set of new pixels and lower quartile : in, n is the total number of pixels included in the new pixel set; if and If is an integer, take and The corresponding position depth value is used as the threshold and ; if and If it is not an integer, round it up and down to get , , , , and take the four depth values ​​of the corresponding positions , , , , The depth value satisfy The pixels are retained.

8. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 6.