Binocular vision servo method for autonomous operation of underwater vehicle-double-arm manipulator

Through binocular visual servo method and Kalman filtering technology, the target continuous positioning and autonomous operation of the underwater vehicle-manipulator system in complex environments is achieved, which solves the problems of target positioning discontinuity and posture randomness, and improves the operation success rate and level of autonomy.

CN120182800AActive Publication Date: 2025-06-20HUAZHONG UNIV OF SCI & TECH

Patent Information

Application Number
CN202510563148.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-20
Estimated Expiration
2045-04-30

AI Technical Summary

Technical Problem

Existing underwater vehicles-two-arm robots are difficult to achieve continuous positioning of targets in complex underwater environments, especially when noise interference and target attitudes are very random.

Method used

The binocular vision servo method is adopted to obtain two-dimensional images and depth images, combine the target detector and Kalman filtering method to achieve unbiased estimation of the three-dimensional coordinates, and adjust the grab angle of the robot according to the estimated value to achieve effective clamping.

Benefits of technology

It improves the autonomy and intelligence of underwater operations, enhances the continuity and smoothness of target positioning, improves the success rate of operations, and adapts to different types and postures of goals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182800A_ABST
    Figure CN120182800A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of underwater operation, and particularly discloses a binocular vision servo method for autonomous operation of an underwater vehicle-double-arm manipulator, and the method comprises the steps: obtaining a two-dimensional image and a depth image in front of the underwater vehicle-double-arm manipulator; inputting the two-dimensional image into a lightweight target detector to obtain a target recognition frame; designing a target three-dimensional coordinate observation value acquisition method based on the shape of an identification frame so as to reduce the influence of a parallax hole; solving an equipment motion matrix based on the three-dimensional coordinates of the key points in the images at the adjacent moments; a nonlinear Kalman filtering method based on a motion matrix is designed by taking the three-dimensional coordinates of the target as state vectors, and continuous estimation of the target position is realized; a target three-dimensional coordinate is used as a reference, an aircraft-manipulator is driven to move, and a two-dimensional image is used for driving a tail end joint of the manipulator to rotate, so that an actuator of the manipulator can effectively clamp the target after the manipulator reaches the target. According to the invention, the autonomy and intelligence of underwater operation can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of underwater operations. More specifically, it relates to a binocular vision servo method for autonomous operation of an underwater vehicle - dual - arm manipulator. Background Art

[0002] Currently, with the increasing demand for ocean resource development, advanced underwater operation technologies are playing an increasingly important role. The underwater vehicle - manipulator system has been widely used in various ocean operation scenarios, such as seabed mining, mariculture, deep - sea oil and gas system operation and maintenance, and underwater rescue. With the increasing complexity of operation scenarios, the requirements for the environmental perception ability and operation autonomy of the underwater vehicle - manipulator system are also gradually increasing.

[0003] However, in practical applications, various noise interferences in the complex underwater environment affect target detection based on visible - light cameras and binocular stereo matching, posing a great challenge to the continuous positioning of targets. In addition, the poses of underwater targets have a large randomness. If the grasping angle is not adjusted according to the target's pose, it will greatly affect the operation success rate. Therefore, in order to improve the autonomy and efficiency of operations, it is necessary to study advanced vision servo methods. For example, industrial vision technology can be used to achieve real - time monitoring and precise measurement of the underwater environment, providing accurate environmental information for the operation of underwater vehicles and manipulators. At the same time, industrial control software can process and analyze these vision data to generate control instructions, realizing precise control of the underwater vehicle - arm manipulator, thereby improving the automation level and operation efficiency of underwater operations. This combination can not only improve the safety and reliability of underwater operations but also expand the application fields of underwater operations to meet the adaptability of the underwater vehicle - manipulator system to complex operation tasks. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the purpose of this application is to provide a binocular vision servo method for autonomous operation of an underwater vehicle - dual - arm manipulator, aiming to solve the problem that the existing underwater vehicle - dual - arm manipulator cannot continuously locate targets in a complex underwater environment.

[0005] To achieve the above - mentioned purpose, in the first aspect, this application provides a binocular vision servo method for autonomous operation of an underwater vehicle - dual - arm manipulator. The underwater vehicle and the dual - arm manipulator form a device for performing underwater operations. The method includes: Obtain the two-dimensional image and depth image of the target in front of the underwater vehicle - dual-arm manipulator; input the two-dimensional image into the target detector to obtain the target recognition box; then, based on the target recognition box, determine the three-dimensional coordinate search area on the depth image, combine the depth image to eliminate the pixels with invalid depth values in the search area, and retain the pixels with the middle depth value according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; Determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; combine the three-dimensional coordinates of the key points at the previous and next moments to obtain the motion parameters of the device; Take the three-dimensional coordinates and three-dimensional velocity of the target as the state vector, take the three-dimensional coordinate observation value of the target as the observation vector, combine the motion parameters to obtain the three-dimensional coordinate prediction value of the target, and fuse the three-dimensional coordinate observation value and the prediction value based on the Kalman filtering method to obtain the unbiased estimated value of the three-dimensional coordinates of the target; With the unbiased estimated value of the three-dimensional coordinates of the target as a reference, drive the device to move to the preset position; then, according to the unbiased estimated value of the three-dimensional coordinates of the target, control the actuators of at least one manipulator to move towards the target, and use the two-dimensional image of the target to drive the end joints of the corresponding actuator to rotate, so that the actuator of the manipulator effectively clamps the target after reaching the target.

[0006] It should be noted that in some application scenarios, the above target detector can also obtain the target category. Furthermore, subsequently, it can be determined whether to control the actuators of one manipulator or two manipulators to move towards the target according to the target category. This application does not make special limitations on this, and those skilled in the art can design according to needs.

[0007] Exemplarily, the above depth image is obtained by the binocular camera after stereo matching.

[0008] It can be understood that in this application, the two-dimensional image is input into the lightweight target detector to obtain the target recognition box; a method for obtaining the three-dimensional coordinate observation value of the target based on the shape of the recognition box is designed to reduce the influence of the parallax hole; the motion matrix of the device is solved based on the three-dimensional coordinates of the key points in the images at adjacent previous and next moments; with the three-dimensional coordinates of the target as the state vector, a nonlinear Kalman filtering method based on the motion matrix is designed to achieve the unbiased estimation of the target position.

[0009] Specifically, after obtaining the target bounding box, the present application obtains the three-dimensional coordinate observations of the target by combining the depth image obtained by binocular vision. Compared with directly selecting the three-dimensional coordinates at the center point of the target bounding box in the prior art, the present application combines the depth information of the pixels around the center of the target bounding box, eliminates the invalid values and outliers among them, and jointly calculates the three-dimensional coordinates of the target, and the calculated coordinate values are more accurate. In addition, the Kalman filtering method combined with the hull motion estimation makes the positioning of the target more continuous, effectively solving the problems of discontinuous and uneven target positioning caused by parallax holes, detector missed detections, hull shaking, etc. In addition, when controlling the manipulator, the present application also adjusts the angle of the end joint of the manipulator in combination with the target image. Compared with the prior art that does not consider the pose relationship between the end joint of the manipulator and the target, the probability of effectively clamping the target can be greatly improved, ensuring the success rate of underwater operations. Further, the present application can obtain the target type according to the target detector, and then control the single manipulator or the double manipulator operation in combination with the target type, with high flexibility.

[0010] The visual servo method designed in the present application enables the underwater vehicle and the manipulator to complete a variety of autonomous operation tasks, has the ability to adapt to the randomness of the target pose, and can be extended to dual-arm operation tasks, effectively improving the operation efficiency and success rate.

[0011] In one embodiment, the target detector includes: a feature extractor, an encoder, and a decoder; The feature extractor is used to extract the features of the target in the two-dimensional image; The encoder is used to adjust the features extracted by the feature extractor to the same scale and splice them in the channel dimension to obtain features ; perform dimensionality reduction on the features through convolution to obtain features ; use the fast Fourier transform to transform the features to the frequency domain, corresponding to the spectrum ; adjust the information in the spectrum through a learnable filter, and then use the inverse fast Fourier transform to transform it back to the time domain to obtain new features; input the new features into the first multi-layer perceptron, deformable convolutional network, and the second multi-layer perceptron for processing in sequence to obtain the enhanced features of the encoder; The decoder is used to obtain the target recognition bounding box through depthwise separable convolution.

[0012] In one embodiment, determining the three-dimensional coordinate observations of the target includes: Taking the center point inside the target recognition bounding box as the center of the search area on the depth image, and determining the width and height of the search area according to the size of the target recognition bounding box; Match the depth image and the two-dimensional image according to the pixel coordinates to obtain the depth values of each pixel point in the search area; Sort all the pixel points in the search area according to the depth values; Based on a preset sorting threshold, eliminate the invalid pixel points in the search area to obtain a new set of pixel points; Retain the pixel points with the middle sorting in the new set of pixel points; Average the three-dimensional coordinates of all the retained pixel points to obtain the three-dimensional coordinate observation value of the target..

[0013] In one embodiment, retaining the pixel points with the middle depth value in the new set of pixel points includes: Calculate the upper quartile of the sorting in the new set of pixel points and the lower quartile :

[0014] wherein, n is the total number of pixel points included in the new set of pixel points; If and are integers, then take and the depth values at the corresponding positions as the thresholds and ; If and are not integers, then round them up and down respectively to obtain , , , , and take the four depth values at the corresponding positions , , , ,

[0015] Retain the pixel points when the depth value satisfies .

[0016] In one embodiment, combining the three-dimensional coordinates of the key points at the previous and current moments to obtain the motion parameters of the device includes: Set the error function i of the th key point to:

[0017] wherein, represents the three-dimensional coordinates of the i th key point at the current moment, Indicates the three-dimensional coordinates of the i th key point at the previous moment, R represents the three-dimensional motion rotation matrix of the device, represents the translation vector of the device; and constitute the motion parameters; With the goal of minimizing the sum of the squares of the error functions of all key points, solve the and .

[0018] In one embodiment, obtaining the predicted value of the three-dimensional coordinates of the target in combination with the motion parameters includes: According to the state vector k of the target at -1 moment, through the state transition matrix obtain the prediction result of the state vector at the kth moment:

[0019] where the state vector ; ; t represents the time interval between the previous and current moments; The predicted value k of the three-dimensional coordinates of the target at moment; the motion parameters and are respectively the three-dimensional motion rotation matrix and the translation vector of the device determined by combining the three-dimensional coordinates of the key points at k -1 and k moments.

[0020] In one embodiment, the unbiased estimated value of the three-dimensional coordinates of the target is:

[0021] where is the predicted value of the three-dimensional coordinates, is the observed value of the three-dimensional coordinates; H is the observation conversion matrix, ; is the filtering gain matrix; and / or

[0022]

[0023]

[0024] where is the covariance matrix at k -1 moment; is the state transfer matrix; and All represent noise; For combination k -1 and k The three-dimensional motion rotation matrix of the device determined by the three-dimensional coordinates of the key points at the moment; For prediction k The covariance matrix of the moments; After considering the motion parameters k Moment covariance matrix.

[0025] In one embodiment, using a two-dimensional image of a target to drive the end joint of the manipulator to rotate includes: The two-dimensional image is input into a network model, and the rotation angle of the end joint is output; the network model is trained based on a sample group consisting of a determined target two-dimensional image and the rotation angle of the end joint, or the network model is learned by a mapping relationship between a preset target two-dimensional image and the rotation angle of the end joint; and / or The network model includes: a convolutional neural network, a global average pooling layer and a fully connected layer; the convolutional neural network includes channel-by-channel convolution or stationary point convolution to extract high-dimensional features in a two-dimensional image, the global average pooling layer is used to fuse the high-dimensional features in the spatial dimension, and the fully connected layer includes an output node for predicting the rotation angle of the end joint according to the fused high-dimensional features; The loss function of the network model is for:

[0026] in, N represents the number of samples; Indicates the desired rotation angle; represents the rotation angle predicted by the network model, Represents a two-dimensional image.

[0027] In a second aspect, the present application provides a binocular visual servo system for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, and the system comprises: An image acquisition unit, used to acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator; The target observation unit is used to input the two-dimensional image into the target detector to obtain a target recognition frame; then determine the three-dimensional coordinate search area on the depth image according to the target recognition frame, remove the pixel points with invalid depth values ​​in the search area in combination with the depth image, and retain the pixel points with middle depth values ​​according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; A target prediction unit, configured to determine the three-dimensional coordinates of multiple key points on a two-dimensional image with reference to a depth image; obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two consecutive moments before and after; and use the three-dimensional coordinates and three-dimensional velocity of the target as a state vector, use the observed value of the three-dimensional coordinates of the target as an observation vector, and obtain a predicted value of the three-dimensional coordinates of the target in combination with the motion parameters. An unbiased estimation unit, configured to fuse the observed value and the predicted value of the three-dimensional coordinates based on the Kalman filtering method to obtain an unbiased estimated value of the three-dimensional coordinates of the target. A servo control unit, configured to drive the device to move to a preset position with reference to the unbiased estimated value of the three-dimensional coordinates of the target; then control the actuators of at least one manipulator to move towards the target according to the unbiased estimated value of the three-dimensional coordinates of the target, and drive the end joints of the corresponding actuators to rotate by using the two-dimensional image of the target, so that the actuators of the manipulator can effectively clamp the target after reaching the target.

[0028] In a third aspect, the present application provides an electronic device, including: at least one memory for storing a program; at least one processor for executing the program stored in the memory, and when the program stored in the memory is executed, the processor is used to execute the method described in the first aspect or any one of the embodiments of the first aspect.

[0029] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, and when the computer program runs on a processor, the processor is caused to execute the method described in the first aspect or any one of the embodiments of the first aspect.

[0030] In a fifth aspect, the present application provides a computer program product, and when the computer program product runs on a processor, the processor is caused to execute the method described in the first aspect or any one of the embodiments of the first aspect.

[0031] Generally speaking, compared with the prior art, the above technical solution conceived by the present application has the following beneficial effects: The present application provides a binocular vision servo method for the autonomous operation of an underwater vehicle - dual - arm manipulator. A lightweight deep neural network object detector is designed to achieve rapid recognition of operation targets. A method for continuously estimating the three - dimensional coordinates of targets is designed. Compared with existing target positioning methods, the continuous positioning method disclosed in the present application will not lose the target position in cases such as parallax holes and missed detections, and at the same time takes into account the impact of the non - linear motion of the hull on positioning, improving robustness. Subsequently, the present application designs a decoupled vision servo method, which divides the operation execution process of the underwater vehicle and the manipulator into the hull dynamic positioning based on the target position and the manipulator target grasping based on the target position and the image. Moreover, a manipulator end - joint controller based on a neural network is designed to adjust the grasping angle, which can effectively solve the problem of autonomous operation under unknown target postures, effectively improve the autonomy and intelligence of underwater operations, and increase the operation success rate. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flowchart of the binocular vision servo method for the autonomous operation of an underwater vehicle - dual - arm manipulator provided by an embodiment of the present application; Figure 2 It is a schematic diagram of the equipment composition of an underwater vehicle - dual - arm manipulator provided by an embodiment of the present application; Figure 3 It is a decoupled vision servo block diagram provided by an embodiment of the present application; Figure 4 It is a block diagram of the lightweight deep neural network object detector provided by an embodiment of the present application; Figure 5 It is a block diagram of the manipulator end - joint controller based on a neural network provided by an embodiment of the present application; Figure 6 It is a three - dimensional coordinate curve graph of the target during the autonomous valve opening and closing operation provided by an embodiment of the present application; Figure 7 It is a curve graph of the position changes of each joint and the end of the manipulator during the autonomous valve opening and closing operation provided by an embodiment of the present application; Figure 8 It is a flowchart of the task completion situation from the first - person view during the autonomous valve opening and closing operation provided by an embodiment of the present application; Figure 9 It is a three - dimensional coordinate curve graph of the grasping points during the autonomous pipeline handling operation provided by an embodiment of the present application; Figure 10 It is a curve graph of the position changes of each joint and the end of the dual - manipulator during the autonomous pipeline handling operation provided by an embodiment of the present application; Figure 11 It is a flowchart of the task completion situation from the first - person view during the autonomous pipeline handling operation provided by an embodiment of the present application; Figure 12This is the architecture diagram of the binocular vision servo system for the autonomous operation of an underwater vehicle - dual-arm manipulator provided by the embodiments of the present application; Figure 13 This is the architecture diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0033] In order to make the objectives, technical solutions and advantages of the present application more clear and understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0034] The term "and / or" in this article is an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.

[0035] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific manner.

[0036] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of key points refers to two or more key points, etc.

[0037] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application.

[0038] Figure 1 This is the flowchart of the binocular vision servo method for the autonomous operation of an underwater vehicle - dual-arm manipulator provided by the embodiments of the present application; as Figure 1 shown, it includes the following steps: Step S101, obtain a two-dimensional image and a depth image of a target in front of the underwater vehicle - dual-arm manipulator; input the two-dimensional image into a target detector to obtain a target recognition frame; then, based on the target recognition frame, determine a three-dimensional coordinate search area on the depth image, combine the depth image to eliminate pixel points with invalid depth values in the search area, and retain pixel points with the middle depth value according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; Step S102, determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two consecutive moments; Step S103, taking the three-dimensional coordinates and three-dimensional velocity of the target as the state vector, taking the three-dimensional coordinate observation value of the target as the observation vector, combining the motion parameters to obtain the three-dimensional coordinate prediction value of the target, and fusing the three-dimensional coordinate observation value and the prediction value based on the Kalman filter method to obtain the unbiased estimate value of the three-dimensional coordinate of the target; Step S104, driving the device to move to a preset position with reference to the unbiased estimate of the target's three-dimensional coordinates; then controlling at least one actuator of the manipulator to move toward the target according to the unbiased estimate of the target's three-dimensional coordinates, and using the two-dimensional image of the target to drive the end joint of the corresponding actuator to rotate, so that after the manipulator reaches the target, its actuator effectively clamps the target.

[0039] The purpose of this application is to design a visual servoing method to realize the autonomous operation execution of the underwater vehicle-manipulator system, and to continuously locate the target through a visible light camera under insufficient lighting conditions and complex water conditions. In addition, the visual servoing method has the ability to adapt to targets of different types and postures, ensure the adaptive adjustment of the grasping angle, and realize the autonomy and intelligence of the underwater vehicle-manipulator system.

[0040] In a more specific embodiment, the above method may specifically include the following steps: The images of the underwater vehicle and the manipulator system obtained by the forward-looking camera at the previous moment and the current moment are obtained.

[0041] The structure of the underwater vehicle-manipulator system can be found in Figure 2 shown.

[0042] Design a lightweight underwater target detector based on deep neural network to quickly obtain the types and pixel coordinates of each target contained in the current image; The field of view depth map is generated from the binocular image through the internal and external parameters of the camera. Based on the target shape and pixel coordinates, the three-dimensional information search range of the target on the depth map is adaptively calculated, and invalid values ​​and outliers are removed. The remaining valid values ​​are averaged to obtain the target three-dimensional coordinate observation value at the current moment. Detect key points on the image at the previous moment and use optical flow estimation to obtain the pixel coordinates of the corresponding points on the current image. Based on this, obtain the corresponding three-dimensional coordinates of the key points on the depth maps of the previous and next frames, and calculate the three-dimensional motion matrix of the system at two consecutive moments based on least squares; The three-dimensional coordinates of the target at the current moment are preliminarily estimated through the linear motion model, and the predicted value of the three-dimensional coordinates of the target at the current moment is further calculated based on the motion matrix and this value. The observed value and the predicted value of the three-dimensional coordinates of the target are fused through the Kalman gain to achieve an unbiased estimate of the three-dimensional coordinates of the target at the current moment in the onboard coordinate system; Based on the current target three-dimensional coordinates and the expected target three-dimensional coordinates, the dynamic positioning error of the underwater vehicle and manipulator system is obtained, and it is controlled to navigate to the expected position; the current target three-dimensional coordinates are set as the expected position of the end of the manipulator, and based on the kinematic model, the expected angles of each joint of the manipulator are calculated and the difference from the current angles is used as the control quantity to control the movement of the manipulator; a manipulator end joint controller based on a deep neural network is designed, with the target image as the input, the rotation angle is output, and the grasping angle of the end effector is controlled to adapt to the randomness of the target posture.

[0043] Specifically, the decoupled visual servo method for the autonomous operation task of the underwater vehicle-manipulator system disclosed in this application is shown in Figure 3 and includes the following contents: (1) A lightweight deep neural network is designed to quickly identify the target in the image; (2) A method for obtaining the three-dimensional coordinate values of the target based on binocular stereo vision is designed; (3) A Kalman filtering method based on camera non-linear motion compensation is proposed to achieve continuous target positioning; (4) A decoupling scheme for converting the target position and image into control commands is designed.

[0044] The structure of the designed lightweight underwater target detector is shown in Figure 4 and includes: a lightweight feature extractor for extracting high-level features from the image, an encoder combining global frequency domain filtering and deformable convolution for enhancing the robustness of the model, and a decoder based on depthwise separable convolution for quickly outputting the target category and its pixel coordinates from the features output by the encoder.

[0045] Specifically, in order to enhance the expression ability of the model and its adaptability to targets of different sizes, an encoder based on frequency domain filtering and deformable convolution is designed in the embodiment of this application. The features extracted by the feature extractor are adjusted to the same scale and concatenated in the channel dimension to obtain the feature . Through convolution, is dimensionally reduced to obtain . Using the fast Fourier transform, is converted to the frequency domain to obtain its spectrum . The information in the spectrum is adjusted by a learnable filter :

[0046] where is element-wise multiplication. Then, using the inverse fast Fourier transform, it is converted back to the time domain to obtain the new .

[0047] The newly obtained features pass through a multi-layer perceptron to further enhance the non-linear expression ability of the model.

[0048] The features processed by frequency-domain filtering are processed by a deformable convolution to enhance the representation ability for targets of various sizes. The features processed in the same way are then passed through a multi-layer perceptron to enhance the non-linear representation ability of the model.

[0049] The features processed by the above encoder are fed into a decoder composed of depthwise separable convolutions to predict the category of the target pixel-by-pixel and the central pixel coordinates and sizes of the recognition bounding box.

[0050] Furthermore, the specific process of the proposed method for obtaining the three-dimensional coordinate values of targets based on binocular stereo vision is as follows: Calculate the pixel coordinates of the target center based on the target recognition bounding box obtained by the detector as the center of the search area, and then calculate the width and height of the search area according to the size of the target recognition bounding box. Extract the corresponding three-dimensional coordinates at the corresponding positions of each pixel in the area to form a set . Remove the invalid values therein through a relatively large threshold, that is, if then directly delete it to obtain a new set .

[0051] Calculate the upper quartile of the set and the lower quartile :

[0052] where n is the size of the set. If and are integers, then take the values at the corresponding positions as the thresholds and respectively. At this time, only the three-dimensional information that satisfies the following conditions needs to be retained:

[0053] If and are not integers, then round them up and down respectively, and take the four values at the corresponding positions , , , . At this time, the threshold calculation formula is:

[0054] Take the mean value of the valid values in the new set as the three-dimensional coordinate observation value of the target at the current moment:

[0055] The process of the proposed Kalman filtering method based on camera non-linear motion compensation is as follows: First, it is necessary to estimate the non - linear motion matrix of the hull between two consecutive frames: Detect Shi - Tomasi corner points on the image at the previous moment, and obtain the three - dimensional coordinates of these key points at the corresponding positions on the depth map Find the corresponding key points on the current image through optical flow estimation, and obtain the three - dimensional coordinates of these key points at the corresponding positions on the current depth map Define the following error function:

[0056] Define the following least - squares problem according to the above error:

[0057] By solving the above problem, the three - dimensional motion rotation matrix of the underwater vehicle and manipulator system between two frames can be obtained and the translation vector .

[0058] Secondly, define the state vector of the target three - dimensional coordinates and the observation vector as follows:

[0059] Define the state transition matrix and the observation transformation matrix as:

[0060] Finally, the unbiased estimation process of the target three - dimensional coordinates is as follows: According to the state of the target at k -1 moment , obtain the prediction result of the state vector at k moment :

[0061] where, = ; The state variable at the previous moment is determined by the unbiased estimation value at the previous moment and the motion speed of the device.

[0062] According to the covariance matrix of the target at the previous moment , predict the covariance matrix at the current moment through the state transition matrix:

[0063] The rotation matrix and the translation vector Simultaneously add it to the prediction result of the state vector at the current moment as follows:

[0064] Add the rotation matrix in the non - linear motion estimation result of the hull to the prediction result of the covariance matrix at the current moment as follows:

[0065] Calculate the filtering gain matrix through the predicted covariance matrix , the observation matrix :

[0066] Fuse the prediction result at time k and the observation result through the filtering gain matrix to obtain an unbiased estimate of the three - dimensional coordinates of the current target :

[0067] Update the covariance matrix :

[0068] The decoupling scheme for converting the target position and the image into control commands divides the autonomous operation of the underwater vehicle - manipulator system into two parts: In the first part, set the hull to a dynamic positioning task, and calculate the three - dimensional coordinates of the target in the hull coordinate system at the current moment and the interpolation with the desired coordinates as the control error, and input it into the underlying controller to drive the underwater vehicle to approach the target and stay still. Here, approaching means that the distance between the device and the target is less than a preset distance threshold, such as 0.5m.

[0069] In the second part, convert the three - dimensional coordinates of the target in the camera coordinate system to obtain in the base coordinate system of the manipulator. Set the expected value of the end - effector position of the manipulator as . Encode the relationship between the target image and the end - effector joint rotation angle into a learnable mapping and hand it over to a deep neural network for learning: ​​

[0070] Among them, represents the rotation angle, represents the target image, represents the network weights.

[0071] Furthermore, for the above-mentioned end-effector joint controller of the manipulator based on a deep neural network, its specific form, such as Figure 5 shown, includes: a convolutional neural network, a global average pooling layer, and a fully connected layer. The convolutional neural network contains channel-wise convolution or point-wise convolution to extract high-dimensional features in the job target image. The global average pooling is used to fuse the high-dimensional features in the spatial dimension. The fully connected layer contains an output node for predicting the rotation angle based on the fused features.

[0072] The present application uses the following embodiments to verify the above technical solutions: The basic parameters of the underwater vehicle-manipulator system used in this embodiment are shown in Table 1.

[0073] Table 1 Basic parameters of the underwater vehicle-manipulator system

[0074] This embodiment uses a pool test to illustrate the effectiveness and advancement of the method proposed in the present application. In the test, the underwater vehicle-manipulator will complete various valve autonomous opening and closing tasks and the autonomous pipeline handling tasks of the two arms to demonstrate the effectiveness of the control method proposed in the present application. Next, the relevant parameter settings are introduced. The initial distance between the underwater vehicle-manipulator and the job target is a random number between [1m, 2m]. When the vehicle body is performing the dynamic positioning task, the expected position of the target in the vehicle body coordinate system is [0, 0, 0.5m], and the error is 20cm. The error between the end position of the manipulator and the expected position needs to be less than 5cm.

[0075] The test results are as Figures 6 to 11 shown. Figure 6 is the change curve of the position of the target in the camera coordinate system with the execution of different valve opening and closing tasks. Figure 7 is the change curve of the rotation angles of the joints of the manipulator arm and the change of the end position in the base coordinate system. The test results show that when the method of the present application is adopted, the discontinuity of target positioning is significantly overcome. The joints of the manipulator arm rotate according to the position and image of the target, making its end position close to the valve. Figure 8 is the first-person view image during the task execution. From Figure 8 it can be concluded that when the method of the present application is adopted, the underwater vehicle-manipulator can autonomously complete the opening and closing operations of different valves in a random posture.

[0076] Figure 9It is the curve of the change in the position of the target grasping point of the pipeline during the operation process. Figure 10 They are the curves of the rotation angles of the joints of the dual robotic arms and the change in the end position in the base coordinate system. The test results show that when the method of this application is adopted, the hull successfully approaches the target and performs dynamic positioning, and the joints of the dual robotic arms rotate according to the position and image of the target, making the end position close to the pipeline. Figure 11 Show the first-person view images during the task execution process. From Figure 11 it can be concluded that when the method provided by the embodiment of this application is adopted, the underwater vehicle-manipulator can use both hands to complete the grasping of the pipeline.

[0077] In summary, this application discloses a binocular vision servo method for the autonomous operation of an underwater vehicle-dual-arm manipulator. Under the influence of poor underwater lighting conditions, many complex interferences, and the shaking of the hull itself, it is a huge challenge for the underwater vehicle-manipulator system to complete the operation autonomously through its own sensing devices. In order to ensure the rapidity of target detection, a lightweight deep neural network target detector is designed. In order to avoid losing the target position in cases such as parallax holes and missed detections, a method for continuously estimating the three-dimensional coordinates of the target is designed, and at the same time, the influence of the non-linear motion of the hull on positioning is considered to improve the robustness. Then, this application designs a scheme for converting the target position and image into control commands, divides the operation execution process of the underwater vehicle and the manipulator into the hull dynamic positioning based on the target position and the manipulator target grasping based on the target position and image, and designs a manipulator end joint controller based on the neural network to adjust the grasping angle, which can effectively solve the problem of autonomous operation with unknown target postures and improve the operation success rate.

[0078] Figure 12 It is the architecture diagram of the binocular vision servo system for the autonomous operation of the underwater vehicle-dual-arm manipulator provided by the embodiment of this application; as Figure 12 shown, it includes: An image acquisition unit 1210, configured to acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator; A target observation unit 1220, configured to input the two-dimensional image into a target detector to obtain a target recognition frame; then determine a three-dimensional coordinate search area on the depth image according to the target recognition frame, combine the depth image to eliminate the pixel points with invalid depth values in the search area, and retain the pixel points with the middle depth value according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; A target prediction unit 1230 is configured to determine the three-dimensional coordinates of multiple key points on a two-dimensional image with reference to a depth image; obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two consecutive moments; and use the three-dimensional coordinates and three-dimensional velocity of the target as a state vector, and the observed value of the three-dimensional coordinates of the target as an observation vector, and obtain a predicted value of the three-dimensional coordinates of the target by combining the motion parameters. An unbiased estimation unit 1240 is configured to fuse the observed value and the predicted value of the three-dimensional coordinates based on the Kalman filtering method to obtain an unbiased estimated value of the three-dimensional coordinates of the target. A servo control unit 1250 is configured to drive the device to move to a preset position with reference to the unbiased estimated value of the three-dimensional coordinates of the target; then control the actuator of at least one manipulator to move towards the target according to the unbiased estimated value of the three-dimensional coordinates of the target, and drive the end joint of the corresponding actuator to rotate by using the two-dimensional image of the target, so that the actuator of the manipulator can effectively clamp the target after reaching the target.

[0079] It should be understood that the above system is used to execute the method in the above embodiments. For the corresponding program units in the system, their implementation principles and technical effects are similar to those described in the above method. The working process of the system can refer to the corresponding process in the above method and will not be elaborated here.

[0080] Based on the method in the above embodiments, an embodiment of the present application provides an electronic device, as Figure 13 shown. The electronic device may include: a processor 1310, a communication interface 1320, a memory 1330, and a communication bus 1340. Among them, the processor 1310, the communication interface 1320, and the memory 1330 complete mutual communication through the communication bus 1340. The processor 1310 can call the logical instructions in the memory 1330 to execute the method in the above embodiments.

[0081] In addition, when the logical instructions in the above memory 1330 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in various embodiments of the present application.

[0082] Based on the method in the above embodiments, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program runs on a processor, the processor is caused to execute the method in the above embodiments.

[0083] Based on the method in the above embodiments, an embodiment of the present application provides a computer program product. When the computer program product runs on a processor, the processor is caused to execute the method in the above embodiments.

[0084] It can be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. The general-purpose processor may be a microprocessor or any conventional processor.

[0085] The method steps in the embodiments of the present application may be implemented in a hardware manner or by a processor executing software instructions. The software instructions may be composed of corresponding software modules. The software modules may be stored in a random access memory (RAM), flash memory, read-only memory (ROM), programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), registers, hard disk, removable hard disk, CD-ROM, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may be located in an ASIC.

[0086] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)), etc.

[0087] It can be understood that the various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application.

[0088] It is easy for those skilled in the art to understand that the above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A binocular visual servoing method for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, characterized in that: Methods include: Acquire a two-dimensional image and a depth image of the target in front of the underwater vehicle-dual-arm manipulator; Inputting the two-dimensional image into a target detector to obtain a target recognition frame; Then, the three-dimensional coordinate search area on the depth image is determined according to the target recognition frame, and the pixels with invalid depth values ​​in the search area are removed in combination with the depth image, and the pixels with middle depth values ​​are retained according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; Determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two moments before and after; The three-dimensional coordinates and three-dimensional velocity of the target are used as the state vector, the three-dimensional coordinate observation value of the target is used as the observation vector, and the three-dimensional coordinate prediction value of the target is obtained in combination with the motion parameters, and the three-dimensional coordinate observation value and the prediction value are fused based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target; The device is driven to move to a preset position with reference to an unbiased estimate of the target's three-dimensional coordinates; then, at least one actuator of the manipulator is controlled to move toward the target according to the unbiased estimate of the target's three-dimensional coordinates, and the end joint of the corresponding actuator is driven to rotate using a two-dimensional image of the target, so that after the manipulator reaches the target, its actuator effectively clamps the target.

2. The method according to claim 1, characterized in that: The target detector comprises: a feature extractor, an encoder and a decoder; The feature extractor is used to extract the features of the target in the two-dimensional image; The encoder is used to adjust the features extracted by the feature extractor to the same scale and concatenate them in the channel dimension to obtain features. ; Through convolution, features Perform dimensionality reduction to obtain features ; Use fast Fourier transform to transform the features Convert to frequency domain, corresponding to spectrum ; Adjust the spectrum through learnable filters The information in the image is then converted back to the time domain using the inverse fast Fourier transform to obtain new features. The new features are sequentially input into the first multi-layer perceptron, the deformable convolutional network, and the second multi-layer perceptron for processing to obtain the features enhanced by the encoder. The decoder is used to obtain a target recognition frame through depth-separable convolution.

3. The method according to claim 1, characterized in that Determine the three-dimensional coordinate observation values ​​of the target, including: Taking the center point in the target recognition frame as the center of the search area on the depth image, and determining the width and height of the search area according to the size of the target recognition frame; Match the depth image with the two-dimensional image based on pixel coordinates to obtain the depth value of each pixel in the search area; Sort all pixels in the search area according to their depth values; Eliminate invalid pixels in the search area based on a preset sorting threshold to obtain a new pixel set; Keep the pixel points in the middle of the new pixel point set; The three-dimensional coordinates of all retained pixels are averaged to obtain the three-dimensional coordinate observation value of the target.

4. The method according to claim 3, characterized in that: Keep the pixels with the middle depth value in the new pixel set, including: Calculate the upper quartile of the sorted set of new pixels and lower quartile : in, n is the total number of pixels included in the new pixel set; if and If is an integer, then take and The corresponding position depth value is used as the threshold and ; if and If is not an integer, round it up and down to get , , , , and take the four depth values ​​at the corresponding positions , , , , The depth value satisfy The pixels at that time are retained.

5. The method according to claim 1, characterized in that Combine the three-dimensional coordinates of the key points at the previous and next moments to obtain the motion parameters of the device, including: The first i The error function of the key points Set to: in, Indicates the current moment i The three-dimensional coordinates of the key points, Indicates the last moment i The three-dimensional coordinates of the key points, R Represents the three-dimensional motion rotation matrix of the device, Represents the device translation vector; With the goal of minimizing the sum of squares of the error functions of all key points, solve the and .

6. The method according to claim 1, characterized in that Acquiring the predicted three-dimensional coordinates of the target in combination with the motion parameters includes: According to the target k The state vector at time -1 , through the state transfer matrix Get the prediction result of the state vector at time k : Among them, the state vector ; ; t Indicates the time interval between the previous and next moments; Target k The predicted three-dimensional coordinates at the time ; Motion parameters and Combined k -1 and k The three-dimensional coordinates of the key points at the moment determine the device's three-dimensional motion rotation matrix and device translation vector.

7. The method according to claim 1, characterized in that Unbiased estimate of the target's 3D coordinates for: in, is the predicted value of the three-dimensional coordinate, is the three-dimensional coordinate observation value; H is the observation transformation matrix, ; is a filter gain matrix; and / or in, for k -1 moment covariance matrix; is the state transfer matrix; and All represent noise; For combination k -1 and k The three-dimensional motion rotation matrix of the device determined by the three-dimensional coordinates of the key points at the moment; For prediction k The covariance matrix of the moments; After considering the motion parameters k Moment covariance matrix.

8. The method according to claim 1, characterized in that: The two-dimensional image of the target is used to drive the end joint of the manipulator to rotate, including: The two-dimensional image is input into a network model, and the rotation angle of the end joint is output; the network model is trained based on a sample group consisting of a determined target two-dimensional image and the rotation angle of the end joint, or the network model is learned by a mapping relationship between a preset target two-dimensional image and the rotation angle of the end joint; and / or The network model includes: a convolutional neural network, a global average pooling layer and a fully connected layer; the convolutional neural network includes channel-by-channel convolution or point-by-point convolution to extract high-dimensional features in a two-dimensional image, the global average pooling layer is used to fuse the high-dimensional features in the spatial dimension, and the fully connected layer includes an output node for predicting the rotation angle of the end joint based on the fused high-dimensional features; The loss function of the network model is for: in, N represents the number of samples; Indicates the desired rotation angle; represents the rotation angle predicted by the network model, Represents a two-dimensional image.

9. A binocular visual servo system for autonomous operation of an underwater vehicle and a dual-arm manipulator, wherein the underwater vehicle and the dual-arm manipulator constitute a device for performing underwater operations, characterized in that: The system includes: An image acquisition unit, used to acquire a two-dimensional image and a depth image of a target in front of the underwater vehicle-dual-arm manipulator; The target observation unit is used to input the two-dimensional image into the target detector to obtain a target recognition frame; then determine the three-dimensional coordinate search area on the depth image according to the target recognition frame, remove the pixel points with invalid depth values ​​in the search area in combination with the depth image, and retain the pixel points with middle depth values ​​according to the depth value sorting, so as to determine the three-dimensional coordinate observation value of the target; The target prediction unit is used to determine the three-dimensional coordinates of multiple key points on the two-dimensional image with reference to the depth image; and obtain the motion parameters of the device by combining the three-dimensional coordinates of the key points at two moments before and after; and taking the three-dimensional coordinates and three-dimensional velocity of the target as the state vector, taking the three-dimensional coordinate observation value of the target as the observation vector, and combining the motion parameters to obtain the three-dimensional coordinate prediction value of the target; An unbiased estimation unit is used to fuse the three-dimensional coordinate observation value and the predicted value based on the Kalman filter method to obtain an unbiased estimate of the three-dimensional coordinate of the target; A servo control unit is used to drive the device to move to a preset position with reference to an unbiased estimate of the target's three-dimensional coordinates; then, according to the unbiased estimate of the target's three-dimensional coordinates, control at least one actuator of the manipulator to move toward the target, and use a two-dimensional image of the target to drive the end joint of the corresponding actuator to rotate, so that after the manipulator reaches the target, its actuator effectively clamps the target.

10. An electronic device, characterized in that: include: at least one memory for storing a computer program; At least one processor is used to execute the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Rapid sorting method based on deep learning

    CN114693661A

  • Motion information and visual information fused mechanical arm space pose self-adaptive estimation method

    CN119228893A

  • Search And Rescue Unmanned Aerial System

    US20190188906A1

Cited By

  • Operation and maintenance manipulator intelligent control method and system based on visual identification

    CN120680525A