Stack bridge berthing method, medium and equipment based on fusion perception of vision and laser radar
By using a fusion perception method combining vision and lidar, the berthing process of the wave-compensated trestle is automatically controlled, solving the problems of limited field of vision and safety hazards in manual operation, and achieving stable and precise automatic guidance of the trestle.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA OFFSHORE ENG & TECH CO LTD
- Filing Date
- 2026-04-09
- Publication Date
- 2026-07-07
AI Technical Summary
Under current technology, the berthing operation of wave-compensated trestle bridges relies on human experience, which can easily lead to safety accidents due to limited visibility or improper operation.
By employing a vision-and-LiDAR fusion perception method, the target position of the docking area is determined through visual images and LiDAR point clouds, and coordinate transformation and joint parameter calculation are performed to achieve automated motion control of the trestle.
It achieves fully automated guidance and docking of the trestle, reduces operational complexity, avoids safety issues caused by limited visibility or improper operation, and ensures the stability and precise control of the three-degree-of-freedom joints.
Smart Images

Figure CN122345863A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine engineering equipment technology, and in particular to a method, medium and equipment for berthing a jetty using a fusion perception of vision and lidar. Background Technology
[0002] A wave-compensated pier is a device used for transferring personnel or goods at sea. With its dynamic positioning system activated, it can dynamically compensate for the six degrees of freedom motion of a ship caused by wind, waves, and other environmental factors, thereby improving the safety of personnel or goods moving on the pier. However, berthing a wave-compensated pier requires experienced operators to manually move it continuously to the designated docking or hovering point. This requires operators to remain focused and observe the pier's movement from multiple angles, and to accurately estimate the relative distance between the pier's end and the docking area. If the operator's view is obstructed, observation is neglected, or experience is insufficient, equipment damage or even a safety accident may occur.
[0003] Therefore, a technical solution is needed to automatically guide the berthing of wave-compensated trestle bridges. Summary of the Invention
[0004] One objective of this application is to provide a method, medium, and device for docking a trestle using a fusion perception system of vision and lidar, in order to solve the safety problems that operators may encounter when docking trestles due to limited field of vision, improper operation, and other factors under the existing technology.
[0005] To achieve the above objectives, some embodiments of this application provide a gantry docking method based on the fusion of visual and lidar perception, the method comprising:
[0006] Based on the acquired visual images, the target pixel position of the docking area in the visual images is determined. The visual images are obtained from the camera, and the target pixel position includes the pixel coordinates of the center point of the docking area in the image coordinate system, the pixel width, and the pixel height.
[0007] Based on the target pixel position, the acquired laser point cloud, and the spatial filter, the first coordinate of the center point of the docking area in the lidar coordinate system is determined. The laser point cloud comes from the lidar, and the spatial filter is used to determine the target point cloud corresponding to the docking area based on the target pixel position and the laser point cloud.
[0008] Perform coordinate transformation on the first coordinate to determine the second coordinate of the center point of the docking and berthing area in the trestle coordinate system;
[0009] The target joint parameters of the trestle are determined based on the second coordinate. The joint parameters include rotation angle, pitch angle and extension length.
[0010] The motion trajectory of the trestle is determined based on the target joint parameters. The motion trajectory includes multiple sets of joint parameters.
[0011] The motion trajectory is sent to the trestle control device so that the trestle control device can execute the corresponding trestle motion, which includes rotation, pitch and extension motion.
[0012] Furthermore, based on the acquired visual images, the target pixel positions of the docking and berthing areas in the visual images are determined, including:
[0013] The acquired visual images are matched with the target images of the preset docking and berthing areas using feature matching and real-time tracking to determine the target pixel positions of the docking and berthing areas in the visual images.
[0014] Furthermore, the spatial filter is determined based on the mapping relationship between the target pixel position and the camera / LiDAR coordinates, as expressed by the following formula:
[0015] ,
[0016] in, SpatialFilter For spatial filters, The pixel coordinates of the center point of the docking area at the target location in the image coordinate system. To align with the pixel width of the docking area, To align with the pixel height of the berthing area, For the first i The coordinates of the target point cloud in the lidar coordinate system For the first i The coordinates of the target point cloud in the camera coordinate system For the first i The pixel coordinates of the target point cloud in the image coordinate system This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. This is the translation matrix of the lidar coordinate system relative to the camera coordinate system.
[0017] Furthermore, based on the target pixel location, the acquired laser point cloud, and the spatial filter, the first coordinates of the center point of the docking area in the lidar coordinate system are determined, including:
[0018] Based on the target pixel position, the acquired laser point cloud, and the spatial filter, determine the target point cloud corresponding to the docking area;
[0019] Based on the coordinates of the target point cloud in the lidar coordinate system, determine the coordinates of the mean point of the target point cloud;
[0020] Determine the set of nearest neighbors of the mean point based on its coordinates;
[0021] The coordinates of the mean point of the nearest neighbor set are used as the first coordinates of the center point of the docking area in the lidar coordinate system.
[0022] Furthermore, the camera-LiDAR coordinate mapping relationship is determined based on the camera's intrinsic parameter matrix, the rotation matrix of the LiDAR coordinate system relative to the camera coordinate system, and the translation matrix of the LiDAR coordinate system relative to the camera coordinate system, as expressed by the following formula:
[0023] ,
[0024] in, This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. This represents the inverse depth.
[0025] Furthermore, the rotation matrix of the lidar coordinate system relative to the camera coordinate system and the translation matrix of the lidar coordinate system relative to the camera coordinate system are determined based on the least squares error of the camera-lidar coordinate mapping relationship, and are expressed by the following formula:
[0026] ,
[0027] ,
[0028] in, For least squares error, This is the parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. For inverse depth, The objective function is the least squares function.
[0029] Furthermore, the target joint parameters of the trestle are determined based on the second coordinate, and expressed by the following formula:
[0030] ,
[0031] in, This is the turning angle of the pier. The pitch angle of the pier. This refers to the telescopic length of the trestle. This is the second coordinate in the trestle coordinate system.
[0032] Furthermore, based on the target joint parameters, the motion trajectory of the trestle is determined, including:
[0033] Based on the relative distance between the end of the trestle bridge and the docking area, the number of interpolation steps is calculated, expressed by the following formula:
[0034] ,
[0035] ,
[0036] ,
[0037] in, This is the second coordinate in the trestle coordinate system. These are the coordinates of the end of the trestle in the trestle coordinate system. For the total displacement, This refers to the relative distance between the end of the trestle bridge and the docking / berthing area. N For the number of interpolation steps, The average speed of the trestle's extension and retraction. To control the cycle;
[0038] Based on the number of interpolation steps, the coordinates of the corresponding multiple interpolation points are determined, as expressed by the following formula:
[0039] ,
[0040] in, For the first The coordinates of the interpolation points ;
[0041] Based on the coordinates of multiple interpolation points, multiple sets of joint parameters in the motion trajectory of the trestle are determined, and expressed by the following formula:
[0042] ,
[0043] in, For the first The joint parameters corresponding to each interpolation point.
[0044] Some embodiments of this application also provide a computer-readable medium having computer-readable instructions stored thereon, which can be executed by a processor to implement the aforementioned gantry docking method based on visual and lidar fusion perception.
[0045] Some embodiments of this application also provide an electronic device, which includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device performs the aforementioned gantry docking method based on visual and lidar fusion perception.
[0046] Compared with existing technologies, the solution provided in this application can determine the target pixel position of the docking area in the visual image based on the acquired visual image. Then, based on the target pixel position, the acquired laser point cloud, and the spatial filter, the first coordinate of the center point of the docking area in the lidar coordinate system is determined. The first coordinate is then transformed to obtain the second coordinate of the center of the docking area in the trestle coordinate system. The target joint parameters of the trestle are determined based on the second coordinate. Based on the target joint parameters, the motion trajectory of the trestle is determined and sent to the trestle control device so that the trestle control device can execute the corresponding trestle movement. This achieves fully automated trestle guidance and docking, reduces the complexity of trestle operation, avoids safety issues caused by limited vision or improper operation of the operator, and also achieves stability and precise control of the three-degree-of-freedom joint execution during the automatic guidance of the trestle. Attached Figure Description
[0047] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings.
[0048] Figure 1 A flowchart of a gantry docking method based on visual and lidar fusion perception provided for some embodiments of this application.
[0049] Figure 2 This is a system structure diagram of a trestle docking method that realizes fusion perception of vision and lidar, provided for some embodiments of this application. Detailed Implementation
[0050] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.
[0051] Here, the visual and lidar fusion perception method for trestle mooring in this application embodiment is suitable for scenarios where wave-compensated trestles are guided to moor.
[0052] In this scenario, the end of the wave-compensated trestle needs to be moved to the docking area for hovering or mooring. During this process, the trestle operator's view is easily obstructed, making it impossible to know the current position of the trestle. Improper operation may lead to safety accidents.
[0053] The visual and lidar fusion perception-based trestle docking method provided in this application can determine the target pixel position of the docking area in the visual image based on the acquired visual image. Then, based on the target pixel position, the acquired lidar point cloud, and the spatial filter, the first coordinate of the center point of the docking area in the lidar coordinate system is determined. The first coordinate is then transformed to obtain the second coordinate of the center point of the docking area in the trestle coordinate system. The target joint parameters of the trestle are determined based on the second coordinate, and the motion trajectory of the trestle is determined based on the target joint parameters. The motion trajectory is then sent to the trestle control device so that the trestle control device can execute the corresponding trestle movement. This achieves fully automated trestle docking guidance, reduces the complexity of trestle operation, avoids safety issues caused by limited field of vision or improper operation by operators, and also achieves stability and precise control of the three-degree-of-freedom joint execution during the automatic trestle guidance process.
[0054] Figure 1 The following are some embodiments of this application illustrating the flow of a gantry docking method using a combination of visual and lidar perception executed by an electronic device, where the electronic device is the executing entity of the method, such as... Figure 1 As shown, the method may include the following steps:
[0055] Step S101: Based on the acquired visual image, determine the target pixel position of the docking area in the visual image.
[0056] Here, cameras and lidar are installed on the wave-compensated trestle, and the cameras and lidar are used to perceive the environment related to the trestle.
[0057] In some embodiments of this application, a single image frame and a single point cloud frame can be acquired simultaneously using a camera and a LiDAR, respectively. The camera acquires a visual image, while the LiDAR acquires a laser point cloud.
[0058] In some embodiments of this application, the target pixel position of the docking area in the visual image includes the pixel coordinates of the center point of the docking area in the image coordinate system, the pixel width, and the pixel height. For example, the pixel coordinates of the center point are represented as follows: The pixel width is expressed as Pixel height is represented as .
[0059] In some embodiments of this application, the target pixel position of the docking area in the visual image is determined based on the acquired visual image. This can be achieved by performing feature matching and real-time tracking between the acquired visual image and the target image of the preset docking area to determine the target pixel position of the docking area in the visual image.
[0060] Specifically, the acquired visual image can be matched with the target image, including the docking area, stored in the database in advance to identify the location of the docking area in the visual image. Then, the docking area is tracked in real time using a template matching algorithm. After the identification and tracking process, the target pixel position of the docking area in the visual image can be obtained.
[0061] Step S102: Based on the target pixel position, the acquired laser point cloud, and the spatial filter, determine the first coordinate of the center point of the docking area in the lidar coordinate system.
[0062] Here, the spatial filter is used to determine the target point cloud corresponding to the docking area based on the target pixel position and the laser point cloud.
[0063] In some embodiments of this application, the spatial filter can be determined based on the mapping relationship between the target pixel position and the camera LiDAR coordinates, as expressed by the following formula:
[0064] ,
[0065] in, SpatialFilter For spatial filters, The pixel coordinates of the center point of the docking area at the target location in the image coordinate system. To align with the pixel width of the docking area, To align with the pixel height of the berthing area, For the first i The coordinates of the target point cloud in the lidar coordinate system For the first i The coordinates of the target point cloud in the camera coordinate system For the first i The pixel coordinates of the target point cloud in the image coordinate system This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. This is the translation matrix of the lidar coordinate system relative to the camera coordinate system.
[0066] In some embodiments of this application, the camera-LiDAR coordinate mapping relationship can be determined based on the camera's intrinsic parameter matrix, the rotation matrix of the LiDAR coordinate system relative to the camera coordinate system, and the translation matrix of the LiDAR coordinate system relative to the camera coordinate system, as expressed by the following formula:
[0067] ,
[0068] in, This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. This represents the inverse depth.
[0069] Here, the formula for the camera-LiDAR coordinate mapping relationship is constructed based on the principle of rotational kinematics. High-precision checkerboard data from multiple angles is pre-acquired simultaneously by the camera and LiDAR. Then, a feature point extraction algorithm extracts the corner points of the checkerboard in the image and LiDAR point cloud. Based on the corresponding coordinates of multiple pairs of corner points, the coordinate mapping relationship formula for the camera-LiDAR is then established. and Solve the problem.
[0070] In some embodiments of this application, the rotation matrix of the lidar coordinate system relative to the camera coordinate system and the translation matrix of the lidar coordinate system relative to the camera coordinate system are determined based on the least squares error of the camera-lidar coordinate mapping relationship, and are expressed by the following formulas:
[0071] ,
[0072] ,
[0073] in, For least squares error, This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. For inverse depth, The objective function is the least squares function.
[0074] Specifically, the formula for the coordinate mapping relationship between the camera and the LiDAR is transformed into a least squares problem to define the least squares error, by making... Determine by reaching the minimum and .
[0075] Here, the least squares error can be linearized and differentiated using a perturbation model, thus the rotation matrix is... Convert to rotation vector , Linearization can be expressed by the following formula:
[0076] ,
[0077] ,
[0078] in, For left perturbation, This is the transpose of the Jacobian matrix, i.e. about The transpose of the first derivative, Let the coordinates of the point cloud be in the lidar coordinate system. The horizontal focal length of the camera. This is the camera's focal length in the vertical direction.
[0079] Subsequently, through continuous adjustments using the Levenberg-Marquardt method Iterative optimization solution and To minimize the least squares error To reach the minimum, that is Minimum.
[0080] Additionally, the camera's intrinsic parameter matrix It is a matrix consisting of the camera's horizontal focal length, vertical focal length, pixel center coordinates, and distortion coefficients. The camera's intrinsic parameters, extrinsic parameters, and distortion coefficients are parameters that have been pre-calibrated for the camera.
[0081] In some embodiments of this application, calibrating camera intrinsic parameters, extrinsic parameters, and distortion coefficients may include the following steps:
[0082] 1) Based on the principle of pinhole imaging in cameras, the mapping relationship between the checkerboard pattern and the image can be constructed, which can be expressed by the following formula:
[0083] ,
[0084] ,
[0085] ,
[0086] ,
[0087] in, These are the pixel coordinates of a point in the image. These are the coordinates of the point in the camera coordinate system, which is affected by lens refraction. Let the coordinates of this point be in the world coordinate system. For rotation matrix, It is a translation matrix. For camera internal parameters, For inverse depth, The radial distortion coefficient is... The tangential distortion coefficient is... These are the coordinates of the point in the distortion-corrected normalized camera coordinate system. for and the origin of the normalized camera coordinate system The distance between them.
[0088] Here, a world coordinate system is established with the top left corner of the high-precision checkerboard grid as the origin, the vertical axis as the X-axis, the horizontal axis as the Y-axis, and the direction perpendicular to the X-axis and Y-axis as the Z-axis. A camera coordinate system is established with the optical center of the left camera as the origin, the optical center axis as the Z-axis, the horizontal axis as the X-axis, and the direction perpendicular to the X-axis and Z-axis as the Y-axis.
[0089] 2) Acquire multiple high-precision chessboard images from multiple angles using a camera. For each acquired chessboard image, use a chessboard corner detection algorithm to identify the pixel coordinates of each chessboard corner point in the image. and the world coordinates of each corner point of the chessboard. A one-to-one correspondence is established.
[0090] 3) The pinhole imaging model constructed above is transformed into a least squares problem using Zhang Zhengyou's calibration method to solve for the camera intrinsic parameters and distortion coefficients. The least squares error function can be expressed by the following formula:
[0091] ,
[0092] ,
[0093] in, The first in the image coordinate system The coordinates of the corner points For the world coordinate system The coordinates of the corner points n For the known and The number of multiple pairs of points formed, i.e. the number of detected chessboard corner points.
[0094] 4) Solve for the camera intrinsic parameters, extrinsic parameters, and distortion coefficients iteratively using the Levenberg-Marquardt algorithm.
[0095] Here, the camera's internal parameters are: External reference is The distortion coefficient is .
[0096] In some embodiments of this application, determining the first coordinates of the center point of the docking area in the lidar coordinate system based on the target pixel position, the acquired laser point cloud, and the spatial filter may include the following steps:
[0097] 1) Determine the target point cloud corresponding to the docking area based on the target pixel position, the acquired laser point cloud, and the spatial filter.
[0098] Here, all laser point clouds within the docking and berthing area can be obtained through a spatial filter. Since the camera lacks depth information, the spatial filter will include noise point clouds reflected by other objects into the target point cloud during filtering. In some embodiments of this application, a point cloud segmentation algorithm based on RANSAC is used to process the target point cloud and remove noise point clouds from the target point cloud.
[0099] Here, the RANSAC-based point cloud segmentation algorithm removes noise points from the point cloud by fitting a plane. The specific working principle is as follows:
[0100] a) Set the number of iterations N and the distance threshold d;
[0101] b) Each time, three sets of points are randomly selected from the point cloud to fit a plane, and the plane equation is expressed as Ax+By+Cz+D=0;
[0102] c) Substitute all points into the plane fitted in b), calculate the distance from each point to the plane; if the distance is within the threshold d, it is an interior point; otherwise, it is an exterior point (i.e., a noise point), and count the number of interior points K.
[0103] d) After N iterations, select the plane equation with the largest number of interior points K. The interior points of this plane are retained, and the rest are discarded.
[0104] 2) Determine the coordinates of the mean point of the target point cloud based on the coordinates of the target point cloud in the lidar coordinate system.
[0105] This can be expressed by the following formula:
[0106] ,
[0107] ,
[0108] ,
[0109] in, n The number of target point clouds, The coordinates of the mean point of the target point cloud.
[0110] 3) Determine the set of nearest neighbors of the mean point based on its coordinates.
[0111] Here, we first find the nearest neighbor of the mean point. , The distance from the mean point does not exceed a preset threshold. It can be expressed by the following formula:
[0112] ,
[0113] After finding all the nearest neighbors of the mean point, then search for... nearest neighbor , making The distance from the mean point also does not exceed the preset threshold. This process is repeated until the set of nearest neighbors is obtained. .
[0114] 4) Use the coordinates of the mean point of the nearest neighbor set as the first coordinate of the center point of the docking area in the lidar coordinate system.
[0115] Get the set of nearest neighbors Next, calculate the mean of the coordinates of all points in the set, and then use the coordinates of the mean point. This serves as the first coordinate of the center point of the docking and berthing area in the lidar coordinate system.
[0116] Step S103: Perform coordinate transformation on the first coordinate to determine the second coordinate of the center point of the docking and berthing area in the trestle coordinate system.
[0117] Here, a coordinate system for the trestle is established with the center of rotation of the trestle as the origin, the orientation of the trestle when it rotates at 0° and tilts at 0° as the positive X-axis, the direction parallel to the trestle tower and upward as the positive Z-axis, and the direction perpendicular to the X-axis and Z-axis as the Y-axis.
[0118] In some embodiments of this application, a coordinate transformation is performed on the first coordinates to determine the second coordinates of the center point of the docking area in the trestle coordinate system, which can be expressed by the following formula:
[0119] ,
[0120] in, This is the second coordinate in the trestle coordinate system. The first coordinate in the lidar coordinate system. Let be the rotation matrix of the lidar coordinate system relative to the trestle coordinate system. This is the translation vector of the lidar coordinate system relative to the trestle coordinate system.
[0121] The ship's roll angle can be transmitted in real time based on the ship's attitude measurement device (Motion Reference Unit, MRU). The pitch angle is fed back in real time by the trestle mechanism. Turning angle The calculation is as follows, expressed by the formula:
[0122] ,
[0123] The extension length can be adjusted based on real-time feedback from the trestle mechanism. The calculation is as follows, expressed by the formula:
[0124]
[0125] Step S104: Determine the target joint parameters of the trestle based on the second coordinate.
[0126] In some embodiments of this application, joint parameters may include the slewing angle, pitch angle, and telescopic length of the bridge.
[0127] The target joint parameters of the trestle, such as the slewing angle, pitch angle, and extension length, can be calculated using the second coordinate system. The formula is as follows:
[0128] ,
[0129] in, This is the turning angle of the pier. The pitch angle of the pier. This refers to the telescopic length of the trestle. This is the second coordinate.
[0130] Step S105: Determine the motion trajectory of the trestle based on the target joint parameters.
[0131] Here, to ensure the end of the trestle Able to move under control Path planning is required. Path planning can be performed using linear paths. The path is discretized, and Cartesian space linear interpolation is used to interpolate the points on the path. Then, the interpolated points are used to solve for the joint parameters of the trestle.
[0132] In some embodiments of this application, the motion trajectory of the trestle includes multiple sets of joint parameters, each set of joint parameters may include the trestle's rotation angle, pitch angle and extension length.
[0133] In some embodiments of this application, determining the motion trajectory of the trestle based on the target joint parameters may include the following steps:
[0134] 1) Calculate the number of interpolation steps based on the relative distance between the end of the trestle and the docking area.
[0135] This can be expressed by the following formula:
[0136] ,
[0137] ,
[0138] ,
[0139] in, This is the second coordinate in the trestle coordinate system. These are the coordinates of the end of the trestle in the trestle coordinate system. For the total displacement, Where is the relative distance between the end of the trestle and the docking / berthing area, and N is the number of interpolation steps. The average speed of the trestle's extension and retraction. To control the cycle.
[0140] 2) Determine the coordinates of the corresponding interpolation points based on the number of interpolation steps.
[0141] This can be expressed by the following formula:
[0142] ,
[0143] in, For the first The coordinates of the interpolation points .
[0144] 3) Determine multiple sets of joint parameters in the motion trajectory of the trestle based on the coordinates of multiple interpolation points.
[0145] This can be expressed by the following formula:
[0146] ,
[0147] in, For the first The joint parameters corresponding to each interpolation point.
[0148] Step S106: Send the motion trajectory to the trestle control device so that the trestle control device can execute the corresponding trestle motion.
[0149] Here, the trestle's motion includes rotation, pitch, and extension / retraction. The motion trajectory is represented as a discretized sequence of trestle joint parameters. The data is sent to the trestle control device. Upon receiving it, the trestle control device executes the rotation, pitch, and extension movements in sequence according to the time frame. The detailed operation involves uniformly controlling the trestle's rotation angle from... arrive Pitch angle from arrive , stretching amount from arrive .
[0150] Figure 2 The system illustrating a gantry docking method for implementing vision and lidar fusion perception in some embodiments of this application is shown, such as Figure 2As shown, the hardware modules of this system include a vision camera, a LiDAR, and an industrial control computer. The software modules include intrinsic parameter calibration, extrinsic parameter calibration, target localization, path generation, and mechanism control. The data used includes point cloud data and image data. Achieving gantry docking using a fusion perception system combining vision and LiDAR can include the following steps:
[0151] 1) Multiple high-precision checkerboard images from multiple angles are acquired using a vision camera. A checkerboard corner detection algorithm is used to detect the pixel coordinates of the checkerboard corners in the images and correlate them with the world coordinates of each corner. The corresponding coordinates are then substituted into the camera model, and the Levenberg-Marquardt method is used to minimize the error. To optimize camera intrinsics External reference and distortion coefficient This allows for the calibration of camera parameters.
[0152] 2) Simultaneously acquire high-precision checkerboard data from multiple angles using both a camera and a LiDAR. Use a feature point extraction algorithm to extract the corner points of the checkerboard from the image and point cloud. Substitute the corresponding corner point coordinates into the camera-LiDAR coordinate mapping equation, and use the Levenberg-Marquardt nonlinear optimization method to solve for the equation parameters. This enables the joint calibration of the external parameters of the camera and lidar.
[0153] 3) Simultaneously acquire one frame of visual image and one frame of laser point cloud using an industrial control computer. Perform target recognition and tracking on the visual image, and use a spatial filter constructed based on the intrinsic and extrinsic parameters of the camera and LiDAR to filter the laser point cloud to obtain all point clouds corresponding to the docking area. Then, use a RANSAC-based point cloud segmentation algorithm to remove noisy point clouds to obtain the target point cloud corresponding to the docking area. Calculate the mean point of the target point cloud, and finally obtain the nearest neighbor set based on the mean point. .
[0154] 4) For the set of nearest neighbors Then, the mean value is calculated to estimate the coordinates of the center point of the docking and berthing area. .
[0155] 5) Coordinates of the center point of the docking and berthing area conduct Coordinate transformation yields the coordinates of the center point of the docking and berthing area in the trestle coordinate system. Then, the three-degree-of-freedom joint parameters of the trestle are obtained by inverse solving using polar coordinate formulas. , , .
[0156] 6) The obtained trestle joint parameters , , Calculate the number of interpolation steps Then based on the number of steps Calculate the Cartesian space linear interpolation points at each step. These discretized points are converted into parameters for the motion path of the trestle joints. The movement trajectory of the trestle is generated from these parameter points.
[0157] 7) Track the movement of the trestle joints Input the trestle controller, which executes rotation, pitch, and telescopic movements sequentially over time. Specifically, it controls the trestle's rotation angle at a constant speed from... arrive Pitch angle from arrive , stretching amount from arrive .
[0158] 8) During the interval After a certain time, repeat steps 3) to 5) to reread the new frame image and point cloud for processing, and repeat step 6) to plan the trestle motion trajectory and step 7) to discretize and control the trestle motion.
[0159] In summary, the solution provided in this application can determine the target pixel position of the docking area in the visual image based on the acquired visual image. Then, based on the target pixel position, the acquired laser point cloud, and the spatial filter, the first coordinate of the center point of the docking area in the lidar coordinate system is determined. The first coordinate is then transformed to obtain the second coordinate of the center of the docking area in the trestle coordinate system. The target joint parameters of the trestle are determined based on the second coordinate. Based on the target joint parameters, the motion trajectory of the trestle is determined and sent to the trestle control device so that the trestle control device can execute the corresponding trestle movement. This achieves fully automated trestle guidance and docking, reduces the complexity of trestle operation, avoids safety issues caused by limited field of vision or improper operation by operators, and also achieves stability and precise control of the three-degree-of-freedom joint execution during the automatic guidance of the trestle.
[0160] It should be noted that this application can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of this application can be executed by a processor to implement the steps or functions described above. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of this application can be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0161] In a typical configuration of this application, both the terminal and the network device include one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0162] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0163] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0164] Furthermore, a portion of this application can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. The program instructions invoking the methods of this application may be stored in a fixed or removable recording medium, and / or transmitted via a data stream in a broadcast or other signal carrying medium, and / or stored in the working memory of a computer device operating according to the program instructions. Here, one embodiment of this application includes a device comprising a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to run methods and / or technical solutions based on the foregoing embodiments of this application.
[0165] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application. No reference numerals in the claims should be construed as limiting the scope of the claims. Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices recited in the apparatus claims may also be implemented by a single unit or device in software or hardware. The terms "first," "second," etc., are used to indicate names and do not indicate any particular order.
Claims
1. A method for docking a trestle using a fusion perception system combining vision and lidar, characterized in that, The method includes: Based on the acquired visual image, the target pixel position of the docking area in the visual image is determined, wherein the visual image is from a camera, and the target pixel position includes the center pixel coordinates, pixel width, and pixel height of the docking area in the image coordinate system; Based on the target pixel position, the acquired laser point cloud, and the spatial filter, the first coordinate of the center point of the docking area in the lidar coordinate system is determined. The laser point cloud comes from the lidar, and the spatial filter is used to determine the target point cloud corresponding to the docking area based on the target pixel position and the laser point cloud. Perform coordinate transformation on the first coordinate to determine the second coordinate of the center point of the docking area in the trestle coordinate system; The target joint parameters of the trestle are determined based on the second coordinates, wherein the joint parameters include rotation angle, pitch angle and extension length; The motion trajectory of the trestle is determined based on the target joint parameters, wherein the motion trajectory includes multiple sets of the joint parameters; The motion trajectory is sent to the trestle control device so that the trestle control device can execute the corresponding trestle motion, wherein the trestle motion includes rotation, pitch and extension motion.
2. The method according to claim 1, characterized in that, Based on the acquired visual image, determine the target pixel position of the docking area in the visual image, including: The acquired visual image is matched with the target image of the preset docking and berthing area using feature matching and real-time tracking to determine the target pixel position of the docking and berthing area in the visual image.
3. The method according to claim 1, characterized in that, The spatial filter is determined based on the mapping relationship between the target pixel position and the camera / LiDAR coordinates, and can be expressed by the following formula: , in, SpatialFilter For spatial filters, The pixel coordinates of the center point of the docking area at the target location in the image coordinate system. To align with the pixel width of the docking area, To align with the pixel height of the berthing area, For the first i The coordinates of the target point cloud in the lidar coordinate system For the first i The coordinates of the target point cloud in the camera coordinate system For the first i The pixel coordinates of the target point cloud in the image coordinate system This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. This is the translation matrix of the lidar coordinate system relative to the camera coordinate system.
4. The method according to claim 3, characterized in that, Based on the target pixel position, the acquired laser point cloud, and the spatial filter, the first coordinates of the center point of the docking area in the lidar coordinate system are determined, including: Based on the target pixel position, the acquired laser point cloud, and the spatial filter, determine the target point cloud corresponding to the docking area; Based on the coordinates of the target point cloud in the lidar coordinate system, determine the coordinates of the mean point of the target point cloud; Based on the coordinates of the mean point, determine the set of nearest neighbors of the mean point; The coordinates of the mean point of the nearest neighbor set are used as the first coordinates of the center point of the docking area in the lidar coordinate system.
5. The method according to claim 3, characterized in that, The camera-LiDAR coordinate mapping relationship is determined based on the camera's intrinsic parameter matrix, the rotation matrix of the LiDAR coordinate system relative to the camera coordinate system, and the translation matrix of the LiDAR coordinate system relative to the camera coordinate system, and is expressed by the following formula: , in, This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. This represents the inverse depth.
6. The method according to claim 5, characterized in that, The rotation matrix of the lidar coordinate system relative to the camera coordinate system and the translation matrix of the lidar coordinate system relative to the camera coordinate system are determined based on the least squares error of the camera-lidar coordinate mapping relationship, and are expressed by the following formula: , , in, For least squares error, This is the intrinsic parameter matrix of the camera. Let be the rotation matrix of the lidar coordinate system relative to the camera coordinate system. Let be the translation matrix of the lidar coordinate system relative to the camera coordinate system. Let the coordinates of the point cloud be in the lidar coordinate system. These are the pixel coordinates of the point cloud in the image coordinate system. For inverse depth, The objective function is the least squares function.
7. The method according to claim 1, characterized in that, The target joint parameters of the trestle are determined based on the second coordinate, and expressed by the following formula: , in, This is the turning angle of the pier. The pitch angle of the pier. This refers to the telescopic length of the trestle. This is the second coordinate in the trestle coordinate system.
8. The method according to claim 1, characterized in that, Determining the motion trajectory of the trestle based on the target joint parameters includes: Based on the relative distance between the end of the trestle bridge and the docking area, the number of interpolation steps is calculated, expressed by the following formula: , , , in, This is the second coordinate in the trestle coordinate system. These are the coordinates of the end of the trestle in the trestle coordinate system. For the total displacement, This refers to the relative distance between the end of the trestle bridge and the docking / berthing area. N For the number of interpolation steps, The average speed of the trestle's extension and retraction. To control the cycle; Based on the number of interpolation steps, the coordinates of the corresponding multiple interpolation points are determined, as expressed by the following formula: , in, For the first The coordinates of the interpolation points ; Based on the coordinates of multiple interpolation points, multiple sets of joint parameters in the motion trajectory of the trestle are determined, expressed by the following formula: , in, For the first The joint parameters corresponding to each interpolation point.
9. A computer-readable medium having stored thereon computer-readable instructions that can be executed by a processor to implement the method as described in any one of claims 1 to 8.
10. An electronic device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1 to 8.