Container lockpin positioning method and system based on combination of monocular camera and mechanical arm
By using a positioning method combining a monocular camera with a robotic arm in the automatic loading and unloading system of container locking pins, using deep learning and pole constraint calculation, the problem of insufficient robustness and real-time performance of the existing system is solved, and more efficient locking pin recognition and positioning is achieved, reducing costs.
Patent Information
- Application Number
- CN202510442835.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-27
AI Technical Summary
The existing container locking pin automatic loading and unloading system has insufficient robustness, poor real-time performance, strong dependence on operational scenarios, and high construction costs.
The container lock pin positioning method based on the combination of monocular camera and robotic arm is adopted, and the target detection is performed through the deep learning model. The motion parameters of the monocular camera obtained by the robotic arm are used to calculate the pole constraints, generate a set of pixel pairs of the same name, reconstruct the 3D detection frame point cloud, and calculate the positioning results.
Improve the accuracy and speed of container lock pin identification and position estimation, reduce the hardware cost of the system, and enhance the robustness of positioning and calculation speed.
Smart Images

Figure CN120206523A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of container locking pin loading and unloading, and more specifically, to a container locking pin positioning method and system based on the combination of a monocular camera and a robotic arm. Background Art
[0002] Automated container locking pin loading and unloading systems play a crucial role in enhancing the automation level of ports and reducing the physical burden on operators. Such systems are typically equipped with one or more robotic arms that use vision positioning technology to identify and locate the locking pins, enabling automatic loading and unloading operations of container locking pins. Among them, the vision system, as the core component for realizing the automatic loading and unloading of container locking pins by the robotic arm, plays a very important role. In the application of robot vision in the prior art, binocular or multi-view stereo cameras are usually used to directly obtain the three-dimensional information of the target object (such as the three-dimensional point cloud or the spatial coordinates of feature corner points, etc.). For the obtained three-dimensional information, point cloud processing algorithms are used to calculate and obtain the three-dimensional positioning information for target object positioning. The vision system technologies applied to the automatic loading and unloading of container locking pins mainly include two types: binocular camera locking pin positioning technology and the combination technology of monocular camera and laser ranging. However, both of these two vision system technologies have problems such as insufficient system robustness, poor real-time performance, strong dependence on the operation scenario, and high construction costs required. Summary of the Invention
[0003] This application provides a container locking pin positioning method and system based on the combination of a monocular camera and a robotic arm, which improves the accuracy and recognition speed of container locking pin recognition and position estimation of the monocular vision system, and at the same time reduces the hardware cost of the system.
[0004] The specific technical solutions are as follows:
[0005] In a first aspect, an embodiment of this application provides a container locking pin positioning method based on the combination of a monocular camera and a robotic arm. The container locking pin positioning method includes:
[0006] Based on a deep learning model, and using a monocular camera installed at the end of the robotic arm, perform target detection on the side holes of the container corner fittings in the left and right view images of the corner fittings to be detected, and obtain the left image detection frame pixel set and the right image detection frame pixel set of the minimum enclosing.
[0007] Traverse each point in the left image detection frame pixel set, calculate the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set, and generate a set of corresponding pixel point pairs of the corner fittings to be detected in the left and right view images according to the corresponding constraints of the same-name detection frame pixel points in the left and right view images.
[0008] According to the set of corresponding pixel point pairs, calculate the depth value of each point in the left image detection frame pixel set. Based on the depth value of each point in the left image detection frame pixel set, reconstruct the 3D detection frame point cloud of the corner fitting to be detected, and calculate the coordinates and Euler angles of the center of the 3D detection frame point cloud.
[0009] In some embodiments of the present application, based on the deep learning model, and using a monocular camera installed at the end of the robotic arm, target detection is respectively performed on the side holes of the container corner fitting in the left and right perspective images of the corner fitting to be detected, to obtain the left image detection frame pixel set and the right image detection frame pixel set of the minimum enclosing, specifically including:
[0010] Use the monocular camera installed at the end of the robotic arm to obtain a frame of side hole left RGB image of the corner fitting to be detected in the left perspective, and input it into the deep learning model, and calculate the left image detection frame pixel set S of the minimum enclosing img.left ;
[0011] Control the robotic arm to move horizontally to the right by a preset distance and rotate by a preset angle;
[0012] Use the monocular camera to obtain a frame of side hole right RGB image of the corner fitting to be detected in the right perspective, input it into the deep learning model, and calculate the right image detection frame pixel set S of the minimum enclosing img.right , and record the translation vector t and rotation matrix R of the robotic arm.
[0013] In some embodiments of the present application, the calculation formula for the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set is:
[0014] l right =(K -1 ) T [t]×BK -1 X img,left (1)
[0015] Wherein, l right is the epipolar line of the point X img,left on the side hole left RGB image on the side hole right RGB image, K is the internal parameter matrix of the monocular camera, t is the translation vector of the robotic arm when shooting in the right perspective of the monocular camera relative to shooting in the left perspective, and R is the rotation matrix of the robotic arm when shooting in the right perspective of the monocular camera relative to shooting in the left perspective;
[0016] Denote the positional relationship of the point X img,left on the side hole left RGB image on the diagonal line of the detection frame of the container corner fitting side hole as:
[0017]
[0018] Take the above formulas (1) and (2) as the corresponding constraints for the pixel points of the same-name detection frames in the left and right perspective images.
[0019] In some embodiments of the present application, generating the set of corresponding pixel point pairs of the corner fitting to be detected in the left and right perspective images according to the corresponding constraints of the pixel points of the same-name detection frames in the left and right perspective images specifically includes:
[0020] According to the corresponding constraints of the pixel points of the same-name detection frames in the left and right perspective images, traverse all the points X img.left in the left image detection frame pixel set S img,left , and search for the corresponding pixel point X img.right of the point X img,left in the right image detection frame pixel set S img,right , and generate the set of corresponding pixel point pairs P_Set(X img,left , X img,right ).
[0021] In some embodiments of the present application, the calculation formula for the depth value Dp of each point X img,left in the left image detection frame pixel set is:
[0022]
[0023] where X img,left represents a point in the left image detection frame pixel set S img.left , represents the skew-symmetric matrix of the corresponding pixel point X img,left of the point X img,right , t represents the translational vector of the robotic arm when the monocular camera takes a right perspective shot relative to when it takes a left perspective shot, and R represents the rotational matrix of the robotic arm when the monocular camera takes a right perspective shot relative to when it takes a left perspective shot.
[0024] In some embodiments of the present application, reconstructing the 3D detection frame point cloud of the corner fitting to be detected based on the depth value of each point in the left image detection frame pixel set specifically includes:
[0025] Based on the depth value of each point in the left image detection frame pixel set, calculate the point P(X, Y, Z) on the 3D detection frame of the side hole of the container corner fitting. The calculation formula for the point P(X, Y, Z) is:
[0026]
[0027] where D p is the depth value of each point X img,left in the left image detection frame pixel set, and (x, y) is each point X in the left image detection frame pixel setimg,left coordinates, f x is the focal length of the monocular camera on the x-axis, f y is the focal length of the monocular camera on the y-axis, and (u0, v0) are the coordinates of the principal point of the monocular camera;
[0028] Traverse all points in the left image detection box pixel set to reconstruct the 3D detection box point cloud P_Cloud.
[0029] In a second aspect, an embodiment of the present application provides a container locking pin positioning system based on the combination of a monocular camera and a robotic arm. The container locking pin positioning system includes:
[0030] A detection box pixel set generation module, configured to perform target detection on the side holes of the container corner fittings in the left and right perspective images of the to-be-detected corner fittings respectively by using a monocular camera installed at the end of the robotic arm based on a deep learning model, and obtain a left image detection box pixel set and a right image detection box pixel set that are the smallest enclosing;
[0031] A corresponding pixel point pair set generation module, configured to traverse each point in the left image detection box pixel set, calculate the epipolar constraint of each point in the left image detection box pixel set in the right image detection box pixel set, and generate a corresponding pixel point pair set of the to-be-detected corner fittings in the left and right perspective images according to the corresponding constraints of the corresponding detection box pixels in the left and right perspective images;
[0032] A 3D detection box point cloud reconstruction module, configured to calculate the depth value of each point in the left image detection box pixel set according to the corresponding pixel point pair set, reconstruct the 3D detection box point cloud of the to-be-detected corner fitting based on the depth value of each point in the left image detection box pixel set, and calculate the coordinates and Euler angles of the center of the 3D detection box point cloud.
[0033] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the container locking pin positioning method based on the combination of a monocular camera and a robotic arm as described in the first aspect.
[0034] In a fourth aspect, an embodiment of the present application provides an electronic device. The electronic device includes a processor and a memory coupled to the processor. The memory is used to store a computer program. When the computer program is executed by the processor, the electronic device implements the container locking pin positioning method based on the combination of a monocular camera and a robotic arm as described in the first aspect.
[0035] Fifth aspect, an embodiment of the present application provides a computer program product, which contains instructions that, when running on a computer or a processor, cause the computer or the processor to execute the container lock pin positioning method based on the combination of a monocular camera and a robotic arm as described in the first aspect.
[0036] The beneficial effects of the embodiments of the present application are as follows:
[0037] Based on the deep learning 2D object detection box as a feature, combined with the epipolar constraint calculated from the monocular camera motion parameters obtained by a high-precision robotic arm, search and match the corresponding feature points of the same corner fitting side hole under different perspectives to obtain depth information, and perform three-dimensional reconstruction on the 3D detection box of the corner fitting side hole, so as to calculate and obtain the accurate pose information of the center of the corner fitting side hole. Compared with the traditional scheme based on a 3D camera or a binocular stereo camera, the present application gives full play to the ability of the high-precision control robotic arm, simplifies the feature complexity and the search process of feature points, improves the robustness and calculation speed of positioning, and also reduces the sensor cost of the entire system. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0039] Figure 1 It is a schematic flowchart of a container lock pin positioning method based on the combination of a monocular camera and a robotic arm provided by an embodiment of the present application;
[0040] Figure 2 It is a schematic diagram of epipolar geometry in a container lock pin positioning method based on the combination of a monocular camera and a robotic arm provided by an embodiment of the present application;
[0041] Figure 3 It is a block diagram of the composition of a container lock pin positioning system based on the combination of a monocular camera and a robotic arm provided by an embodiment of the present application. Detailed Embodiments
[0042] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.
[0043] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The terms "including" and "having" in the embodiments and the accompanying drawings of the present application and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.
[0044] A complete visual system for loading and unloading container locking pins mainly includes two key modules: visual detection and distance measurement. The visual detection module analyzes images through deep learning algorithms to identify different types of container locking pins and their installation states on the corner fittings, providing accurate target state information for the automated loading and unloading system and guiding the robotic arm to replace the corresponding locking and unloading tools. The distance measurement module is responsible for calculating the spatial position transformation between the robotic arm gripper and the target locking pin to ensure that the robotic arm can accurately install or remove the locking pin. During the entire automated loading and unloading process, the accurate identification of the visual detection module ensures the correct state judgment of the system, enabling the robotic arm to correctly perform the loading and unloading tasks, while the accurate pose estimation of the distance measurement module ensures the accuracy of the robotic arm's actions, enabling successful completion of the locking pin loading and unloading and avoiding damage to the fixture, locking pin, and the robotic arm itself.
[0045] In the prior art, a stereo vision positioning system based on binocular cameras mainly identifies various locking pins through deep learning algorithms and uses these identification results to guide the robotic arm to select appropriate loading and unloading tools. Then, the system calculates the spatial distance between the loading and unloading tools and the locking pins through the stereo modeling technology implemented by the binocular cameras. On the one hand, the corner point detection technology used in this ranging method relies on the optical characteristics of the image, and its accuracy and robustness are easily affected by the on-site shooting lighting conditions. Therefore, special lighting conditions, such as appropriate light brightness and contrast, need to be set in the operating environment. On the other hand, when performing large-scale corner point stereo modeling, the processing speed may not be sufficient to meet the real-time operation requirements. For the detection of specific corner points, higher requirements are placed on the photography environment and lighting conditions, which usually means more costs need to be invested in the initial stage to ensure a stable photography environment.
[0046] In the strategy of using a monocular camera combined with a laser distance measurement device, the responsibility of the monocular camera is to identify the types of locking pins using a deep learning model, while the laser distance measurement device mainly measures the precise positions of the container angle irons to complete the positioning work. This strategy usually requires the container to be placed at a pre-set position or requires the use of special positioning tools. As a result, this technology depends on a specific environment or requires the use of large mechanical positioning equipment, which increases the deployment cost. At the same time, due to single-point laser ranging, the measured distance may not be the point to be measured, resulting in the problem of ambiguity.
[0047] It can be seen that the existing binocular camera locking pin positioning technology and the combination technology of monocular camera and laser ranging have common problems such as insufficient system robustness, poor real-time performance, strong dependence on the operation scenario, and relatively high construction cost required. To address this problem, the embodiments of the present application disclose a container locking pin positioning method based on the combination of a monocular camera and a robotic arm. Based on the characteristics of the simple scenario of container locking pins in ports and the relatively single features of the detected targets, a method of using the target detection frame as the target matching feature, combining the camera motion parameters fed back by the robotic arm with high-precision control to calculate the explicit epipolar constraint, and accelerating the multi-view three-dimensional positioning is adopted to improve the accuracy and recognition speed of container locking pin recognition and position estimation of the monocular vision system, while reducing the hardware cost of the system. The following will be described in detail respectively.
[0048] Figure 1 FIG. shows a container locking pin positioning method based on the combination of a monocular camera and a robotic arm according to an embodiment of the present application. As Figure 1 shown, the container locking pin positioning method includes the following steps:
[0049] Step S110: Based on a deep learning model and using a monocular camera installed at the end of the robotic arm, perform target detection on the side holes of the container angle parts in the left and right view images of the angle parts to be detected, and obtain the left image detection frame pixel set and the right image detection frame pixel set of the minimum enclosing.
[0050] The embodiments of the present application are mainly applicable to the container locking pin positioning scenario. For the positioning of container locking pins, since the positional relationship between the lower hole position and the side hole position of the angle part where the container locking pin is located is based on the national standard dimensions, after three-dimensional pose positioning of the center of the side hole of the angle part, the three-dimensional poses of the lower hole position of the angle part and the locking pin can be obtained through conversion.
[0051] Specifically, the monocular camera in this application is installed at the end of a robotic arm with high-precision control function. By controlling the horizontal translation and rotation of the robotic arm, RGB images of the side holes of the same angle part to be detected can be obtained from the left and right perspectives. In some embodiments, a monocular camera installed at the end of the robotic arm is used to obtain a frame of left-side hole RGB image of the angle part to be detected from the left perspective and input it into the deep learning model, and the pixel set S of the left detection frame with the smallest enclosing rectangle is calculated. img.left The robotic arm is controlled to move horizontally to the right by a preset distance and rotate by a preset angle; the translated and rotated monocular camera is used to obtain a frame of right-side hole RGB image of the angle part to be detected from the right perspective, input it into the deep learning model, and the pixel set S of the right detection frame with the smallest enclosing rectangle is calculated. img.right And the translation vector t and rotation matrix R of the robotic arm are recorded.
[0052] In the specific implementation process, after the container arrives at the position, at the initial position, the monocular camera on the end of the robotic arm takes a frame of RGB image (left image), inputs it into the deep learning model, and calculates the pixel set S of the horizontal detection frame with the smallest enclosing rectangle. img.left At the same time, the robotic arm translates horizontally to the right by 20 cm to combine the FOV of the monocular camera and the distance between the monocular camera and the target, ensuring that the target to be detected is within the camera's field of view when photographed from both the left and right perspectives. The monocular camera then takes another frame of RGB image (right image) and inputs it into the deep learning model to calculate S. img.right And the translation vector t and rotation matrix R fed back from the robotic arm are recorded. This container locking pin positioning method uses a deep learning model to perform target detection on the side holes of the same angle part to be recognized respectively, and draws the horizontal detection frame with the smallest enclosing rectangle of the angle part side holes. And the rotation and translation amounts (i.e., translation vector t and rotation matrix R) of the monocular camera at two positions obtained from the robotic arm are recorded to facilitate the subsequent reconstruction of the 3D detection frame point cloud of the angle part to be detected. It should be noted that the deep learning model in this application is a well-trained model, whose input is an RGB image, and the output result is the center coordinates, length, and width of the smallest enclosing rectangle of the target to be detected in the image coordinates, so as to reconstruct the horizontal detection frame in the image and calculate the pixel set of the horizontal detection frame.
[0053] Step S120: Traverse each point in the pixel set of the left detection frame, calculate the epipolar constraint of each point in the pixel set of the left detection frame in the pixel set of the right detection frame, and generate a set of corresponding pixel point pairs of the angle part to be detected in the left and right perspective images according to the corresponding constraints of the same-name detection frame pixel points in the left and right perspective images.
[0054] As Figure 2 shown, according to the principle of epipolar geometry, a point P on the virtual 3D detection frame has a corresponding projection point X on the left image (i.e., the left-side hole RGB image). img,left, its corresponding projection point X on the right figure (i.e., the right RGB image of the side hole) img,right must be located on the epipolar line l right . According to the calculation principle of epipolar geometry, it can be obtained that:
[0055] l right = FX img,left (5)
[0056] In the above formula (5), F is the fundamental matrix in epipolar geometry, which represents the epipolar constraint from the left image pixel point to the corresponding epipolar line in the right figure. Its calculation formula is as follows:
[0057] F = (K -1 ) T [t]×RK -1 (6)
[0058] In the above formula (6), K is the internal parameter matrix of the monocular camera, K -1 represents the inverse matrix of the internal parameter matrix K, (K -1 ) T represents the transpose of the inverse matrix of the internal parameter matrix K, t is the translational vector of the robotic arm when the monocular camera takes pictures from the right view relative to the left view, and R is the rotation matrix of the robotic arm when the monocular camera takes pictures from the right view relative to the left view.
[0059] From the above formulas (5) and (6), the calculation formula for the epipolar constraint of each point in the left figure detection frame pixel set in the right figure detection frame pixel set can be obtained as:
[0060] l right = (K -1 ) T [t]×RK -1 X img,left (1)
[0061] In the above formula (1), l right is the epipolar line of the point X img,left on the left RGB image of the side hole on the right RGB image of the side hole. K is the internal parameter matrix of the monocular camera, t is the translational vector of the robotic arm when the monocular camera takes pictures from the right view relative to the left view, and R is the rotation matrix of the robotic arm when the monocular camera takes pictures from the right view relative to the left view.
[0062] The projection point X img,right is the intersection of the epipolar line l right and the right figure detection frame pixel set S img.right . Since the detection frame is rectangular, the number of intersection points n of the epipolar line l right and the detection frame belongs to [1, 2]. Therefore, when traversing the left figure detection frame pixel set S img.left , record the point X img,leftThe positional relationship of the diagonal of the detection frame on the side hole of the container corner fitting is as follows:
[0063]
[0064] Thus, through point X img,left with respect to the positional relationship of the diagonal of the detection frame and the epipolar line l right , determine the number of intersection points of the epipolar line l right and the detection frame, and then specifically obtain its corresponding projection point X img,right on the right figure. In summary, formulas (1) and (2) are the corresponding constraints for the pixel points of the same-name detection frames in the left and right view images.
[0065] In the embodiment of the present application, according to the corresponding constraints of the pixel points of the same-name detection frames in the left and right view images, search for the pixel point pairs of the same-name detection frames in the left and right images. The specific method is: according to the constraints of the above formulas (1) and (2), traverse all points X img.left in the pixel set S img,left of the left figure detection frame, and search for the same-name pixel point X img.right of point X img,left in the pixel set S img,right of the right figure detection frame, and form a set P_Set(X img,left , X img,right ) of the same-name pixel point pairs.
[0066] Among them, combining the image target detection result with the epipolar geometry principle is one of the innovative points of the embodiment of the present application, which simplifies the complexity of finding the same-name feature points in binocular vision matching and improves the robustness and calculation speed of positioning.
[0067] Step S130: Calculate the depth value of each point in the pixel set of the left figure detection frame according to the set of the same-name pixel point pairs. Based on the depth value of each point in the pixel set of the left figure detection frame, reconstruct the 3D detection frame point cloud of the corner fitting to be detected, and calculate the coordinates and Euler angles of the center of the 3D detection frame point cloud.
[0068] In the embodiment of the present application, according to the triangulation principle, the calculation formula for the depth value Dp of each point X img,left in the pixel set of the left figure detection frame is:
[0069]
[0070] Among them, X img,left represents the point in the pixel set S img.left of the left figure detection frame, represents the same-name pixel point X img,left of point X img,rightThe anti-symmetric matrix, t represents the translational vector of the robotic arm when the monocular camera takes a right-view image relative to when it takes a left-view image, and R represents the rotational matrix of the robotic arm when the monocular camera takes a right-view image relative to when it takes a left-view image.
[0071] In some embodiments, based on the depth values of each point in the left-image detection box pixel set, the 3D detection box point P(X, Y, Z) of the side hole of the container corner fitting is calculated, where (X, Y, Z) are the three-dimensional coordinates of point P in the left-view monocular camera coordinate system. For the point P(X, Y, Z) on the virtual 3D detection box, its corresponding pixel point X in the left image can be calculated by the above step S120. img,left (x, y, D p ) combined with the internal parameters (f x , f y , u0, v0) of the monocular camera, can be calculated as follows:
[0072]
[0073] where D p is the depth value of each point X img,left in the left-image detection box pixel set, (x, y) are the coordinates of each point X img,left in the left-image detection box pixel set, f x is the focal length of the monocular camera on the x-axis, f y is the focal length of the monocular camera on the y-axis, and (u0, v0) are the coordinates of the principal point of the monocular camera.
[0074] Traverse all points in the left-image detection box pixel set according to the above formulas (3) and (4) to reconstruct the 3D detection box point cloud P_Cloud. Then, calculate the coordinates and Euler angles of the center of the 3D detection box point cloud P_Cloud, which are the positioning results of the center of the side hole of the corner fitting to be detected.
[0075] By obtaining the rotation and translation amounts of the monocular camera at two positions of the left and right views from the robotic arm, the epipolar line of each detection box pixel point in the left image in the right image can be calculated, which is used as the constraint for searching the corresponding detection box feature points in the right image. Search for the homologous points of each pixel of the left-image detection box on the right-image detection box. According to the parallax of the homologous points in the left and right images on the image, calculate the depth of each pixel point of the left-image detection box. Finally, convert the depth into the coordinates of 3D points to reconstruct the 3D point cloud of the side hole detection box of the corner fitting. The 3D coordinates and Euler angles of its center are the positioning results required. In the embodiments of the present application, the deep learning 2D detection box is used as the feature of the side hole of the corner fitting, and the points on the deep learning detection box are used as feature points twice. The high-precision motion parameters of the monocular camera are directly obtained by using the robotic arm with high-precision control. The epipolar geometry is introduced, which simplifies the feature complexity and the search process of feature points, speeds up the search speed of feature points, and improves the robustness and calculation speed of positioning.
[0076] Corresponding to the above method embodiments, another embodiment of the present application provides a container locking pin positioning system based on the combination of a monocular camera and a robotic arm, as Figure 3 shown. The container locking pin positioning system 200 based on the combination of a monocular camera and a robotic arm mainly includes: a detection frame pixel set generation module 210, a homologous pixel point pair set generation module 220, and a 3D detection frame point cloud reconstruction module 230.
[0077] Specifically, the detection frame pixel set generation module 210 can be used to perform target detection on the side holes of the container corner fittings in the left and right perspective images of the to-be-detected corner fitting based on a deep learning model and using a monocular camera installed at the end of the robotic arm, so as to obtain a left image detection frame pixel set and a right image detection frame pixel set of the minimum enclosing.
[0078] The homologous pixel point pair set generation module 220 can be used to traverse each point in the left image detection frame pixel set, calculate the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set, and generate a homologous pixel point pair set of the to-be-detected corner fitting in the left and right perspective images according to the corresponding constraints of the homologous detection frame pixel points in the left and right perspective images.
[0079] The 3D detection frame point cloud reconstruction module 230 can be used to calculate the depth value of each point in the left image detection frame pixel set according to the homologous pixel point pair set, reconstruct the 3D detection frame point cloud of the to-be-detected corner fitting based on the depth value of each point in the left image detection frame pixel set, and calculate the coordinates and Euler angles of the center of the 3D detection frame point cloud.
[0080] In a specific application scenario, the detection frame pixel set generation module 210 can specifically be used to: use the monocular camera installed at the end of the robotic arm to obtain a frame of left RGB image of the side hole of the to-be-detected corner fitting in the left perspective, and input it into the deep learning model to calculate and obtain the left image detection frame pixel set S img.left ; control the robotic arm to move horizontally to the right by a preset distance and rotate by a preset angle; use the monocular camera to obtain a frame of right RGB image of the side hole of the to-be-detected corner fitting in the right perspective, input it into the deep learning model, and calculate and obtain the right image detection frame pixel set S img.right , and record the translation vector t and rotation matrix R of the robotic arm.
[0081] In a specific application scenario, the calculation formula for the epipolar constraint of each point in the left image detection frame pixel set calculated by the homologous pixel point pair set generation module 220 in the right image detection frame pixel set is:
[0082] l right =(K -1 ) T [t]×RK -1X img,left (1)
[0083] where l right is the point X on the left RGB image of the side hole img,left is the epipolar line on the right RGB image of the side hole, K is the internal parameter matrix of the monocular camera, t is the translational vector of the robotic arm when shooting from the right view of the monocular camera relative to shooting from the left view, and R is the rotational matrix of the robotic arm when shooting from the right view of the monocular camera relative to shooting from the left view;
[0084] Denote the positional relationship of the point X on the left RGB image of the side hole img,left on the diagonal of the detection frame of the corner fitting side hole of the container as:
[0085]
[0086] The corresponding pixel point pair set generation module 220 uses the above formula (1) and formula (2) as the corresponding constraints for the detected pixel points of the same name in the left and right view images.
[0087] In a specific application scenario, the 3D detection frame point cloud reconstruction module 230 calculates the depth value Dp of each point X img,left in the pixel set of the left image detection frame, and the calculation formula is:
[0088]
[0089] where X img,left represents the point in the pixel set S img.left of the left image detection frame, represents the anti-symmetric matrix of the corresponding pixel point X img,left of the point X img,right , t represents the translational vector of the robotic arm when shooting from the right view of the monocular camera relative to shooting from the left view, and R represents the rotational matrix of the robotic arm when shooting from the right view of the monocular camera relative to shooting from the left view.
[0090] It should be noted that for other corresponding descriptions of the various functional modules involved in the container locking pin positioning system based on the combination of a monocular camera and a robotic arm provided in the embodiments of the present application, reference can be made to the corresponding descriptions in the above method embodiments, and details are not described herein again.
[0091] In summary, the embodiments of the present application provide a container lock pin positioning method and system based on the combination of a monocular camera and a robotic arm. Based on the deep learning 2D object detection box as a feature, combined with the epipolar constraint calculated from the monocular camera motion parameters obtained by the high-precision robotic arm, the same-name feature points of the side holes of the same corner fitting under different perspectives are searched and matched to obtain depth information, and the 3D detection box of the side holes of the corner fitting is three-dimensionally reconstructed, so as to calculate and obtain the accurate pose information of the center of the side holes of the corner fitting. Compared with the traditional solutions based on 3D cameras or binocular stereo cameras, the present application gives full play to the capabilities of the high-precision control robotic arm, simplifies the feature complexity and the feature point search process, improves the robustness and calculation speed of positioning, and at the same time reduces the sensor cost of the entire system.
[0092] Based on the above method embodiments, another embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in the above method embodiments of the container lock pin positioning method based on the combination of a monocular camera and a robotic arm is implemented.
[0093] Based on the above method embodiments, another embodiment of the present application provides an electronic device, which includes a processor and a memory. The processor is coupled to the memory, and the memory is used to store a computer program. When the computer program is executed by the processor, the electronic device implements the method described in the above method embodiments of the container lock pin positioning method based on the combination of a monocular camera and a robotic arm.
[0094] Based on the above method embodiments, another embodiment of the present application provides a computer program product, which contains instructions. When the instructions run on a computer or a processor, the computer or the processor is made to execute the method described in the above method embodiments of the container lock pin positioning method based on the combination of a monocular camera and a robotic arm.
[0095] Those of ordinary skill in the art can understand that the drawings are only schematic diagrams of one embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present application. In addition, the modules in the devices in the embodiments can be distributed in the devices in the embodiments according to the description of the embodiments, or can be correspondingly changed to be located in one or more devices different from the present embodiment. The modules in the above embodiments can be combined into one module, or further split into multiple sub-modules.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A container lock pin positioning method based on a monocular camera combined with a mechanical arm, characterized in that: The container lock pin positioning method comprises: Based on the deep learning model and using the monocular camera installed at the end of the robotic arm, the side holes of the container corner fittings to be inspected in the left and right perspective images are detected to obtain the minimum enclosing left image detection frame pixel set and right image detection frame pixel set; Traversing each point in the left image detection frame pixel set, calculating the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set, and generating a set of pixel pairs of the same name of the corner piece to be detected in the left and right view images according to the corresponding constraints of the same name detection frame pixel points in the left and right view images; According to the set of pixel pairs with the same name, the depth value of each point in the left image detection frame pixel set is calculated, and based on the depth value of each point in the left image detection frame pixel set, the 3D detection frame point cloud of the corner piece to be detected is reconstructed, and the coordinates and Euler angles of the center of the 3D detection frame point cloud are calculated.
2. The container lock pin positioning method based on the combination of a monocular camera and a mechanical arm according to claim 1 is characterized in that: The method is based on a deep learning model and uses a monocular camera installed at the end of the robotic arm to perform target detection on the side holes of the container corner fittings in the left and right perspective images of the corner fittings to be detected, and obtains the minimum enclosing left image detection frame pixel set and right image detection frame pixel set, specifically including: The monocular camera installed at the end of the robotic arm is used to obtain a frame of the left RGB image of the side hole of the corner piece to be detected from the left perspective, and input it into the deep learning model to calculate the minimum surrounding pixel set S of the left image detection frame. img.left ; Controlling the robotic arm to move horizontally to the right by a preset distance and to rotate by a preset angle; Using the monocular camera, obtain a frame of the right RGB image of the side hole of the corner piece to be detected in the right perspective, input it into the deep learning model, and calculate the minimum surrounding pixel set S of the right image detection frame img.right , and record the translation vector t and rotation matrix R of the robotic arm.
3. The container lock pin positioning method based on the combination of a monocular camera and a mechanical arm according to claim 2 is characterized in that: The calculation formula for the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set is: l right =(K -1 ) T [t]×RK -1 X img,left (1) Among them, l right is the point X on the left RGB image of the side hole img,left The epipolar line on the right RGB image of the side hole, K is the intrinsic parameter matrix of the monocular camera, t is the translation vector of the manipulator when the monocular camera is shooting from the right perspective relative to the left perspective, and R is the rotation matrix of the manipulator when the monocular camera is shooting from the right perspective relative to the left perspective; Note that point X on the left RGB image of the side hole img,left The position relationship of the diagonal lines of the detection frame of the side hole of the container corner fitting is: The above formula (1) and formula (2) are used as the corresponding constraints for the pixel points of the same-name detection frames of the left and right viewing angle images.
4. The container lock pin positioning method based on the combination of a monocular camera and a mechanical arm according to claim 1 is characterized in that: The generating of a set of pixel pairs of the same name of the corner piece to be detected in the left and right viewing angle images according to the corresponding constraints of the pixel points of the same name of the detection frame of the left and right viewing angle images specifically includes: According to the corresponding constraints of the detection frame pixels of the same name in the left and right perspective images, traverse the detection frame pixel set S of the left image img.left All points X in img,left , in the right image detection box pixel set S img.right Search for the point X img,left The pixel with the same name X img,right , generate the same-name pixel point pair set P_Set(X img,left , X img,right ).
5. The container lock pin positioning method based on the combination of a monocular camera and a mechanical arm according to claim 4 is characterized in that: Each point X in the left image detection frame pixel set img,left The calculation formula of the depth value Dp is: Among them, X img,left Represents the pixel set S of the left image detection box img.left The point in Indicates point X img,left The pixel with the same name X img,right The antisymmetric matrix of t represents the translation vector of the robotic arm when the monocular camera is shooting from the right perspective relative to the left perspective, and R represents the rotation matrix of the robotic arm when the monocular camera is shooting from the right perspective relative to the left perspective.
6. The container lock pin positioning method based on the combination of a monocular camera and a mechanical arm according to claim 1 is characterized in that: The step of reconstructing the 3D detection frame point cloud of the corner piece to be detected based on the depth value of each point in the left image detection frame pixel set specifically includes: Based on the depth value of each point in the pixel set of the left image detection frame, the point P (X, Y, Z) on the 3D detection frame of the container corner fitting side hole is calculated. The calculation formula of the point P (X, Y, Z) is: Among them, D p For each point X in the left image detection box pixel set img,left The depth value of (x, y) is the depth value of each point X in the pixel set of the detection box on the left. img,left The coordinates of x is the focal length of the monocular camera on the x-axis, f y is the focal length of the monocular camera on the y-axis, (u0, v0) is the principal point coordinate of the monocular camera; Traverse all points in the pixel set of the left image detection box to reconstruct the 3D detection box point cloud P_Cloud.
7. A container lock pin positioning system based on a monocular camera combined with a mechanical arm, characterized in that: The container lock pin positioning system comprises: The detection frame pixel set generation module is used to perform target detection on the container corner fitting side holes of the corner fittings to be detected in the left and right perspective images based on the deep learning model and using the monocular camera installed at the end of the robotic arm, and obtain the minimum enclosing left image detection frame pixel set and right image detection frame pixel set; A same-name pixel point pair set generation module is used to traverse each point in the left image detection frame pixel set, calculate the epipolar constraint of each point in the left image detection frame pixel set in the right image detection frame pixel set, and generate the same-name pixel point pair sets of the corner piece to be detected in the left and right perspective images according to the corresponding constraints of the same-name detection frame pixel points in the left and right perspective images; The 3D detection frame point cloud reconstruction module is used to calculate the depth value of each point in the left image detection frame pixel set according to the set of pixel pairs with the same name, reconstruct the 3D detection frame point cloud of the corner piece to be detected based on the depth value of each point in the left image detection frame pixel set, and calculate the coordinates and Euler angles of the center of the 3D detection frame point cloud.
8. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor, the container lock pin positioning method based on the combination of a monocular camera and a robotic arm as described in any one of claims 1-6 is implemented.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory coupled to the processor, the memory is used to store a computer program, and the computer program is executed by the processor, so that the electronic device implements the container lock pin positioning method based on the combination of a monocular camera and a robotic arm as described in any one of claims 1-6.
10. A computer program product, characterized in that The computer program product includes instructions, which, when executed on a computer or a processor, cause the computer or the processor to execute the container lock pin positioning method based on the combination of a monocular camera and a robotic arm as described in any one of claims 1-6.
Citation Information
Patent Citations
Monocular vision-based dense point cloud reconstruction method and system for triangulation measurement depth
CN111798505A
Point cloud fusion method of double-monocular three-dimensional imaging system
CN113781305A
Container corner fitting positioning method, device and equipment and storage medium
CN119540348A
Hole site information detection method and apparatus, and device and storage medium
WO2024255550A1
Cited By
Lightweight container correction method and system based on machine vision
CN120472324A
High-flux high-precision three-dimensional behavior trajectory tracking system and method for whole life cycle of zebra fish
CN122049061A