A visual guidance method, system and computer device based on multi-camera joint positioning
By using a multi-camera joint positioning method, the problem of high-precision visual guidance for large-sized parts that cannot be met by a single camera was solved. Stable visual positioning and accurate grasping of parts without obvious features were achieved, with an accuracy better than ±0.5mm, especially under complex lighting conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-22
- Publication Date
- 2026-04-07
AI Technical Summary
In the existing technology, a single camera cannot meet the high-precision visual guidance of large-sized parts, and it is difficult to stably perform visual positioning for parts that do not have obvious visual recognition features.
A multi-camera joint positioning method is adopted to obtain the pose information of a local area of a large component captured by multiple cameras, calculate the transformation relationship, and combine it with the offset information of the end effector gripper of the robotic arm to achieve high-precision visual guidance.
It achieves high-precision visual positioning and stable grasping of large-sized parts, especially under complex lighting conditions, it can also achieve accurate positioning guidance for parts without obvious features, with a grasping accuracy better than ±0.5mm.
Smart Images

Figure CN119399279B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine vision guidance technology, and in particular to a vision guidance method, system and computer device based on multi-camera joint positioning. Background Technology
[0002] Existing visual positioning solutions typically utilize a single 2D / 3D camera to photograph parts and collect grayscale or depth map data. Identification and positioning are then performed using features within this data, such as round holes, square holes, and corner points within the parts. As the imaging principle of cameras tells us, camera accuracy depends on the camera's resolution and field of view. With the same field of view, higher resolution means smaller pixel size, resulting in higher accuracy, and vice versa. Therefore, when positioning large parts, high-precision cameras struggle to capture complete feature information, while cameras with large fields of view often lack sufficient accuracy. Using high-precision cameras only captures features from small, localized areas of the part. Positioning based on these localized features can lead to significant deviations in the overall positioning of larger parts due to errors in feature localization, thus failing to meet the requirements for precise workpiece grasping.
[0003] Furthermore, most current feature recognition and localization algorithms struggle to maintain stable recognition performance under complex lighting conditions and working environments. To achieve stable recognition of component features, it is often necessary to manually select features that are easy for visual algorithms to recognize, such as circular holes. However, in some scenarios, components lack obvious features similar to circular holes or screw holes, and using corner points or other features can lead to unstable recognition algorithms due to the influence of shadows or other factors in the scene.
[0004] It is evident that existing technologies suffer from several problems: a single camera cannot provide high-precision visual guidance for large-sized components; and for components without obvious visually recognizable features, stable visual positioning guidance is difficult to achieve. Summary of the Invention
[0005] This invention provides a visual guidance method, system, and computer device based on multi-camera joint positioning, to solve the problems in the prior art where a single camera cannot meet the high-precision visual guidance of large-sized parts; and where it is difficult to stably perform visual positioning guidance for parts without obvious visual recognition features.
[0006] To achieve the above objectives, the present invention employs the following technical solution:
[0007] In a first aspect, the present invention provides a visual guidance method based on multi-camera joint localization, comprising:
[0008] S1: Obtain the local area of a large component captured by N cameras, and calculate the component pose information of the local area under each camera.
[0009] S2: Calculate the pose information of the entire large component online based on the component pose information of the local area under each camera;
[0010] S3: Determine the transformation relationship between the pose information of large-sized parts calculated offline and the pose information of large-sized parts calculated online;
[0011] S4: Calculate the required offset information of the end gripper based on the transformation relationship, and guide the robotic arm to complete the corresponding offset to achieve grasping based on the end gripper offset information.
[0012] Secondly, this application provides a vision guidance system based on multi-camera joint positioning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method described in the first aspect.
[0013] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method of the first aspect.
[0014] Beneficial effects:
[0015] The present invention provides a vision guidance method based on multi-camera joint localization. In the offline template stage, local point cloud data of large-sized components is acquired. In the online guidance stage, component data acquired by the same camera is registered with the offline local point cloud, and the rotation and translation matrix from the offline template point cloud to the online guidance point cloud is calculated to determine the local deflection of the large-sized component in the camera coordinate system. After determining the local deflection data of the large-sized component under each camera, the multiple local deflection information can be fused and calculated to obtain the overall positioning pose of the component, which guides the robotic arm to complete the grasping operation. This application can solve the problem that a single camera cannot meet the high-precision vision guidance requirements for large-sized components.
[0016] A further proposed solution utilizes local point cloud data of components acquired from 3D cameras, matching it with point cloud data collected during the offline configuration phase. By fusing visual positioning information from multiple cameras with different features, it enables precise positioning guidance for grasping or placing large components. This provides stable visual positioning guidance even for components lacking obvious visually identifiable features. Attached Figure Description
[0017] Figure 1A flowchart illustrating a visual guidance method based on multi-camera joint localization according to a preferred embodiment of the present invention;
[0018] Figure 2 A flowchart of a visual recognition and positioning method for a local area surface of a component without obvious features, according to a preferred embodiment of the present invention;
[0019] Figure 3 This is a flowchart of a visual guidance method for feature fusion according to a preferred embodiment of the present invention. Detailed Implementation
[0020] The technical solution of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0021] Unless otherwise defined, the technical or scientific terms used in this invention shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms "an" or "a" and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms "connected" or "linked" and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. "Up," "down," "left," "right," etc., are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship also changes accordingly.
[0022] Please see Figure 1 This application provides a visual guidance method based on multi-camera joint localization, comprising:
[0023] S1: Obtain the local area of a large component captured by N cameras, and calculate the component pose information of the local area under each camera.
[0024] In this step, N is a positive integer. In one example, N is 3, and in another example, N is 4. This is just an example and is not a limitation.
[0025] In this context, obtaining a local area of a large component captured by N cameras means using N cameras to capture images of the large component. During the capture process, each camera captures a portion of the large component, and the images captured by the N cameras are combined to obtain a complete image of the large component.
[0026] S2: Calculate the pose information of the entire large component online based on the component pose information of the local area under each camera.
[0027] S3: Determine the transformation relationship between the pose information of large-sized parts calculated offline and the pose information of large-sized parts calculated online.
[0028] In this step, the relative relationship between the pose information of the large-size component calculated offline and the pose information of the large-size component calculated online is T(R, t).
[0029] S4: Calculate the required offset information of the end gripper based on the transformation relationship, and guide the robotic arm to complete the corresponding offset to achieve grasping based on the end gripper offset information.
[0030] The aforementioned vision guidance method based on multi-camera joint localization acquires local point cloud data of large-sized parts during the offline template stage. During the online guidance stage, it registers the part data acquired by the same camera with the offline local point cloud and calculates the rotation and translation matrix from the offline template point cloud to the online guidance point cloud, thereby determining the local deflection of the large-sized part in the camera coordinate system. After determining the local deflection data of the large-sized part under each camera, the multiple local deflection information can be fused and calculated to obtain the overall positioning pose of the part, which guides the robotic arm to complete the grasping operation. This application can solve the problem that a single camera cannot meet the high-precision vision guidance requirements for large-sized parts.
[0031] The steps of the high-precision visual guidance method based on multi-camera joint localization described above will be described in detail below with a complete example:
[0032] In this example, firstly for each camera, such as Figure 2 As shown, the main steps of the visual recognition and localization method for localized areas of components without obvious features are as follows:
[0033] During the offline phase:
[0034] The robotic arm, with its gripper, approaches the component, bringing various local positions of the component into the assumed depth of field of view of each camera, and the cameras collect local data.
[0035] The depth map data acquired by the camera is converted into point cloud data through camera intrinsic parameters, and noise in the point cloud is removed using a denoising algorithm. Point cloud data of non-part parts is also manually removed.
[0036] Calculate the centroid of the point cloud data in the camera coordinate system, and calculate the three principal axis vectors of the point cloud using the principal component analysis algorithm. Establish a coordinate system with the centroid as the origin and the three principal axis vectors as the coordinate axes. After transforming the data to the robot base coordinate system using camera extrinsic parameters, save the data.
[0037] During the online phase:
[0038] The robotic arm, carrying a gripper and a camera, approaches the component and brings various local areas of the component into the depth of field of the camera on the gripper, allowing the camera to collect data.
[0039] The depth map in the online phase is converted into online point cloud data P by calculating the camera intrinsic parameters;
[0040] Load the offline point cloud data Q, and calculate the point in the online point cloud data P that is closest to every point in the offline point cloud data Q, i.e. Where R is the rotation matrix of the point cloud, t is the translation vector, and q is a point in the offline point cloud data, which is a 4x4 identity matrix in the initial step;
[0041] Based on the relative relationships between corresponding points in the offline point cloud and the online point cloud obtained from the above steps, the SVD decomposition algorithm is used to solve for Rt, i.e. The online point cloud data P is transformed using the calculated Rt. Specifically, the transformation can be performed by calculating the pose relationship (rotation matrix R, translation vector t) between point clouds Q and P using point cloud registration algorithms such as ICP. Point cloud P is a three-dimensional matrix composed of a set of XYZ vectors, and the transformation is based on R×P+t.
[0042] Because the above steps for solving Rt can lead to local optima in the matching of point clouds, resulting in poor matching performance, this invention introduces the Welsch function to add weights to the distances between each point. This results in a smaller weight for points that are far apart, making them less susceptible to noise in the point cloud and thus preventing the above steps from getting stuck in a local optimum.
[0043] Calculate the distance from each point in the transformed point cloud data P to the nearest point in the offline point cloud data Q, and determine whether the offline point cloud and the online point cloud match accurately. If they do not match accurately, recalculate the relative relationship between the corresponding points in the offline point cloud and the online point cloud based on the above steps, and repeat the above matching steps.
[0044] In this step, the method to determine whether the offline point cloud and the online point cloud match accurately can be: the point cloud distance value of each point is less than a threshold, which is considered convergence; if the proportion of the converged point cloud to the total number of point cloud data P is greater than a threshold (e.g., 90%), it is considered a correct match. This is only an example and is not a limitation.
[0045] Based on the calculated relative relationship T(R,t) transformation matrix, hereinafter referred to as the Rt rotation transformation matrix, the offline stored coordinate system is transformed into the newly captured point cloud, and this coordinate system is used as the centroid and axis of the online captured data;
[0046] Complete the local visual recognition and localization algorithm for a single high-precision small field-of-view camera;
[0047] By offline acquisition of local point cloud data of large-sized components, and during the identification and guidance phase, component data acquired by the same camera is registered with the aforementioned offline local point cloud. The rotation and translation data from the offline point cloud to the point cloud acquired during the guidance phase are calculated to determine the local deflection of the large-sized component in this camera coordinate system. After determining the local deflection data of the large-sized component under a high-precision, small-field-of-view camera, the local deflection information of the same component acquired by multiple cameras at other locations can be fused and calculated to obtain the precise position of the entire component. This position is then used to guide the gripper to complete the guided grasping operation.
[0048] Through the aforementioned offline and online steps, precise local feature positioning of large-sized components captured online can be achieved within the coordinate system of a single camera. However, using positioning data from a single camera can easily lead to significant overall guidance errors for large-sized components. This is because even slight angular shifts in a local area can cause substantial displacements in parts of the large-sized component that are far from that local area. Therefore, it is necessary to use multiple high-precision cameras for image positioning of large-sized components, combining the overall positioning information to achieve high-precision overall guidance for the large-sized component.
[0049] Furthermore, a high-precision guidance method for joint positioning using data from multiple cameras is as follows:
[0050] As previously mentioned, for high-precision grasping scenarios involving large-sized parts, this application will employ multiple high-precision 3D cameras with small fields of view. Each camera will locate the part's features within a localized area, calculate the localized pose information of the part, and convert this data to the robot's base coordinate system based on calibrated extrinsic parameters. To achieve high-precision guided grasping, such as... Figure 3 As shown, this high-precision guidance method for feature fusion also consists of two stages:
[0051] Offline phase:
[0052] Through the above steps, the local positioning data of each camera captured by the large-sized component is calculated and obtained;
[0053] After storing the above data, the robot flange end is manually taught so that the designed complex gripper can meet the set requirements when grasping large-sized parts. The set requirements are that it can stably and accurately grasp large-sized parts and record the flange end position data at this time.
[0054] Based on the hand-eye calibration results of N cameras, the component pose information of the local area captured by N cameras is transformed into the coordinate system of the end flange of the robotic arm, and fused based on the least squares method to obtain the overall pose information of the component, which is then saved.
[0055] Online phase:
[0056] Each camera captures a local area of a large component, and the aforementioned featureless component surface visual recognition and localization algorithm is used to calculate the component pose information of the local area under each camera.
[0057] Using the same least squares method as in the offline stage, the pose information of the entire large-sized component is calculated;
[0058] Calculate the transformation relationship Rt between the offline pose information and the online pose information of large-sized parts, and use this Rt data to calculate the data information of the gripper that needs to be offset.
[0059] The robot's end effector guides the gripper to make the corresponding offset, which can then accurately guide and position large-sized parts and complete the grasping action.
[0060] In summary, to address the need for high-precision visual guidance and positioning of large-sized components, this invention proposes a method for recognizing positioning features using a high-precision camera with a small field of view in complex lighting scenarios, and for achieving multi-camera joint positioning. Using this method, we have successfully completed high-precision positioning guidance and grasping of large components such as front and rear doors, rear covers, front covers, and chassis in multiple automotive OEMs, with positioning guidance accuracy better than ±0.5mm.
[0061] This application also provides a high-precision visual guidance system based on multi-camera joint positioning, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the high-precision visual guidance method based on multi-camera joint positioning described above.
[0062] This application also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the high-precision visual guidance method based on multi-camera joint positioning described above. This computer device can implement various embodiments of the high-precision visual guidance method based on multi-camera joint positioning described above, and can achieve the same beneficial effects, which will not be elaborated here.
[0063] The preferred embodiments of the present invention have been described in detail above. It should be understood that those skilled in the art can make numerous modifications and variations based on the concept of the present invention without creative effort. Therefore, all technical solutions that can be obtained by those skilled in the art based on the concept of the present invention through logical analysis, reasoning, or limited experimentation on the basis of existing technology should be within the scope of protection defined by the claims.
Claims
1. A visual guidance method based on multi-camera joint localization, characterized in that, include: S1: Obtain the local area of a large component captured by N cameras, and calculate the component pose information of the local area under each camera. S2: Calculate the pose information of the entire large component online based on the component pose information of the local area under each camera; S3: Determine the transformation relationship between the pose information of large-sized parts calculated offline and the pose information of large-sized parts calculated online; S4: Calculate the required offset information of the end gripper according to the transformation relationship, and guide the robotic arm to complete the corresponding offset according to the end gripper offset information to achieve grasping; The calculation of component pose information for a local region under each camera includes: S11: For each camera, store its offline coordinate system and convert the depth map data acquired during the offline phase into offline point cloud data. The depth map data collected during the online phase is converted into online point cloud data. ; S12: Determine offline point cloud data With online point cloud data The relative relationship T(R, t) between corresponding points is used to analyze online point cloud data. Perform the transformation; S13: Introduce the Welsch function to add weights to the distance between each point; S14: Calculate the distance from each point in the transformed point cloud data P to the nearest point in the offline point cloud data Q, and determine whether the offline point cloud and the online point cloud are accurately matched. If they are not accurately matched, repeat S12 to S14 above. S15: Based on the relative relationship T(R, t) transformation matrix, transform the offline stored coordinate system to the newly captured point cloud, and use this coordinate system as the centroid and axis of the online captured data; S16: Calculate the component pose information of the local area under the camera based on the centroid and axis of the data captured online.
2. The visual guidance method based on multi-camera joint localization according to claim 1, characterized in that, For each camera, its offline coordinate system is stored, and the depth map data acquired during the offline phase is converted into offline point cloud data. ,include: During the offline phase, the robotic arm with grippers approaches the component, bringing various local positions of the component into the field of view and depth of field of each camera, and enabling the cameras to collect local data. The depth map data acquired by the camera is converted into point cloud data using camera intrinsic parameters. Noise removal algorithms are then used to remove noise from the point cloud, and non-partial part point cloud data is removed to obtain offline point cloud data. ; Calculate offline point cloud data The centroid in the camera coordinate system is determined, and the three principal axis vectors of the point cloud are calculated using the principal component analysis algorithm. A coordinate system with the centroid as the origin and the three principal axis vectors as the coordinate axes is established. The data is then transformed to the robot base coordinate system using camera extrinsic parameters and saved.
3. The visual guidance method based on multi-camera joint localization according to claim 1, characterized in that, The determination of offline point cloud data With online point cloud data The relative relationship T(R, t) between corresponding points includes: Calculate online point cloud data Mid-range offline point cloud data The point closest to each point in the array; Determine offline point cloud data based on the nearest point. With online point cloud data The relative relationship between corresponding points is T(R, t).
4. The visual guidance method based on multi-camera joint localization according to claim 1, characterized in that, Before step S3, the method further includes: The pose information of large-sized components calculated in the offline phase.
5. The visual guidance method based on multi-camera joint localization according to claim 4, characterized in that, The pose information of the large-sized component calculated in the offline stage includes: In the offline phase, acquire the pose information of large-sized components in local areas captured by N cameras; The component pose information of the local area is stored, and the end of the robot flange is taught by the human to make the end gripper meet the set requirements when grasping large-sized components, and the flange end pose data at this time is recorded. Based on the hand-eye calibration results of N cameras, the component pose information of the local area captured by the N cameras is transformed into the coordinate system of the end flange of the robotic arm, and then fused based on the least squares method to obtain the overall pose information of the component, which is then saved.
6. A visual guidance system based on multi-camera joint positioning, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor, when executing the computer program, implements the steps of any one of the methods described in claims 1 to 5.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Vehicle-mounted visual positioning method and system based on multiple combined cameras and storage medium
CN111243021A
Robot visual global positioning method under guidance of laser ranging
CN117889861A