A visual positioning method, a terminal and a storage medium

By acquiring and transforming the coordinates of image feature points from the robot's vision device, and using optimization functions to update camera intrinsic parameters and matrices, the problems of low accuracy and poor robustness in traditional visual positioning are solved, achieving higher-precision visual positioning.

CN116051634BActive Publication Date: 2026-01-02SHENZHEN YOUIBOT ROBOTICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211696821.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-01-02
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

Traditional visual positioning technology suffers from low accuracy and poor robustness.

Method used

By acquiring images captured by the robot's vision device, the pixel coordinates and real-world coordinates of feature points in the images are obtained. Coordinate transformation is performed using camera intrinsics, rotation matrices, and translation matrices. An optimization function is established to update the camera intrinsics and matrices to improve positioning accuracy.

Benefits of technology

It improves the accuracy and robustness of visual positioning, adapts to the parameter requirements of actual application scenarios, and achieves more precise robot motion control.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116051634B_ABST
    Figure CN116051634B_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the field of machine vision, and specifically provide a visual positioning method, a terminal and a storage medium. The method comprises: acquiring an image captured by a vision device on a robot, and acquiring a first pixel coordinate value of a feature point in the image and a corresponding first real coordinate value of the feature point in a real world coordinate system; based on camera intrinsic parameters, a rotation matrix and a translation matrix of the vision device, acquiring a second real coordinate value corresponding to the first pixel coordinate value in the real world coordinate system according to the first pixel coordinate value of the feature point; determining a motion position of the robot according to the second real coordinate value; based on the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device, acquiring a second pixel coordinate value corresponding to the second real coordinate value in the image according to the second real coordinate value; and establishing an optimization function according to the first pixel coordinate value and the second pixel coordinate value of the feature point in the image, and updating the camera intrinsic parameters, the rotation matrix and the translation matrix.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of machine vision, and particularly relates to a visual positioning method, a terminal and a storage medium. BACKGROUND

[0002] In recent years, machine vision technology has developed rapidly, and the combination of robots and machine vision technology has been relatively mature in industrial applications. Machine vision technology gives robots the ability to perceive changes in the working environment. Through the vision system, the robot can accurately identify and locate the target in the working space and autonomously move to the target position. In the traditional visual positioning technology, the camera calibration result is used for visual positioning, and there are problems of low visual positioning accuracy and poor robustness. SUMMARY

[0003] The main purpose of the embodiments of the present application is to provide a visual positioning method, a terminal and a storage medium, aiming to improve the accuracy and robustness of visual positioning.

[0004] In a first aspect, the embodiments of the present application provide a visual positioning method, comprising:

[0005] obtaining an image photographed by a vision device on a robot, and obtaining a first pixel coordinate value of a feature point in the image and a first real coordinate value corresponding to the first pixel coordinate value in a real world coordinate system;

[0006] based on camera intrinsic parameters, a rotation matrix and a translation matrix of the vision device, obtaining a second real coordinate value corresponding to the first pixel coordinate value in the real world coordinate system according to the first pixel coordinate value of the feature point;

[0007] determining a motion position of the robot according to the second real coordinate value;

[0008] based on the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device, obtaining a second pixel coordinate value corresponding to the second real coordinate value in the image according to the second real coordinate value;

[0009] establishing an optimization function according to the first pixel coordinate value and the second pixel coordinate value corresponding to the feature point in the image;

[0010] updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device according to the optimization function.

[0011] In a second aspect, the embodiments of the present application also provide a robot terminal, which comprises a processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing the connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of any one of the visual positioning methods provided in the present application are realized.

[0012] In a third aspect, the embodiments of the present application further provide a storage medium for computer readable storage, and the storage medium stores one or more programs, and the one or more programs are executable by one or more processors to implement the steps of the visual positioning method according to any one of the embodiments of the present application.

[0013] The embodiments of the present application provide a visual positioning method, a terminal and a storage medium. In the visual positioning method, an image captured by a vision device on a robot is obtained, and a first pixel coordinate value of a feature point in the image and a first real coordinate value corresponding to the first pixel coordinate value in a real world coordinate system are obtained. Based on camera intrinsic parameters, a rotation matrix and a translation matrix of the vision device, the second real coordinate value corresponding to the first pixel coordinate value in the real world coordinate system is obtained according to the first pixel coordinate value of the feature point. The motion position of the robot is determined according to the second real coordinate value. Based on the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device, the second pixel coordinate value corresponding to the feature point in the image is obtained according to the second real coordinate value. An optimization function is established according to the first pixel coordinate value and the second pixel coordinate value of the feature point in the image. The camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device are updated according to the optimization function. In the present application, the pixel coordinate value of the feature point in the image and the real coordinate value corresponding to the pixel coordinate value in the real world coordinate system are obtained. Then, the predicted coordinate value corresponding to the pixel coordinate value in the real world coordinate system is obtained based on the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device. The predicted pixel coordinate value is derived according to the predicted coordinate value. The optimization function is constructed by using the predicted pixel coordinate value and the pixel coordinate value. The camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device are updated by solving the optimization function. Therefore, the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device can be updated by using the data in the actual application scene, which is more suitable for the parameter requirements in the actual application scene, and the accuracy of visual positioning is improved. BRIEF DESCRIPTION OF DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating labor.

[0015] Figure 1 A flowchart of a visual positioning method provided by the embodiments of the present application is shown in the figure.

[0016] Figure 2 A fixed-pitch array flat plate calibration plate sample provided by the embodiments of the present application is shown in the figure.

[0017] Figure 3 A schematic diagram of an image coordinate system and a pixel coordinate system provided for an embodiment of the present application is shown in the following figure;

[0018] Figure 4 For Figure 1 A specific implementation of step S6 in the embodiment of the present application corresponds to the following step flow chart;

[0019] Figure 5 A structural schematic block diagram of a terminal provided for an embodiment of the present application is shown in the following figure. DETAILED DESCRIPTION

[0020] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0021] The flow chart shown in the accompanying drawings is only an example and does not necessarily include all the contents and operations / steps, nor does it necessarily have to be executed in the order described. For example, some operations / steps can be further decomposed, combined or partially merged, so the actual execution order can be changed according to the actual situation.

[0022] It should be understood that the terms used in the present application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the present application specification and the appended claims, unless otherwise clearly indicated by the context, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0023] The embodiments of the present application provide a visual positioning method, a terminal and a storage medium. The visual positioning method can be applied to a terminal, which can be a robot.

[0024] The embodiments of the present application provide a visual positioning method, a terminal and a storage medium, wherein the visual positioning method obtains pixel coordinate values of feature points in an image and real coordinate values corresponding to the pixel coordinate values in a real world coordinate system, then obtains predicted coordinate values corresponding to the pixel coordinate values in the real world coordinate system based on camera intrinsic parameters, a rotation matrix and a translation matrix of a vision device, derives predicted pixel coordinate values according to the predicted coordinate values, constructs an optimization function using the predicted pixel coordinate values and the pixel coordinate values, and updates the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device through solving the optimization function, so that the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device can be updated using data in an actual application scenario, which is more consistent with the parameter requirements in the actual application scenario, thereby improving the accuracy of visual positioning.

[0025] Some embodiments of the present application will be described in detail below with reference to the drawings. The following embodiments and features of the embodiments can be combined with each other in the case of no conflict.

[0026] Please refer to Figure 1 , Figure 1 A flowchart of a visual positioning method provided by an embodiment of the present application.

[0027] As Figure 2 shown, the visual positioning method includes steps S1 to S6.

[0028] Step S1: obtaining an image captured by a visual device on a robot, and obtaining a first pixel coordinate value of a feature point in the image and a first real coordinate value corresponding to the first pixel coordinate value in a real world coordinate system.

[0029] Exemplarily, an industrial robot is a high-end manufacturing device in manufacturing industry, and the requirements for stability and positioning accuracy are very high. Therefore, image processing by means of machine vision technology is needed to realize guided positioning and pattern recognition by a camera, so as to quickly obtain the centroid and boundary of an object, meet the self-positioning requirements of the industrial robot system, shorten the gap between the expected position and the end position, and further realize the accuracy of positioning.

[0030] In order to improve the accuracy of visual positioning, the parameters need to be continuously optimized according to the actual use data. Therefore, in the actual use process, the image captured by the visual device on the robot and the corresponding pixel coordinate value of the feature point in the image and the first real coordinate value corresponding to the pixel coordinate value in the real world coordinate system are needed. With the increase of the required data collection samples, the analyzed data also increases. That is, the image captured by the visual device and the corresponding pixel coordinate value of the feature point in the image and the first real coordinate value corresponding to the pixel coordinate value in the real world coordinate system also increase, which is more conducive to subsequent optimization analysis.

[0031] In some embodiments, the obtaining of the image captured by the visual device on the robot includes: controlling the visual device of the robot to capture a preset template to obtain the image, the preset template including at least one feature point; and identifying the feature point in the image to obtain the feature point in the image.

[0032] Exemplarily, the image itself is a matrix composed of only brightness and color, and it is difficult to consider the information of the camera from the matrix level alone, so generally, representative points are first selected from the image, which are called key points. When the light in the environment changes or the distance between the camera and the object changes, the position of the original key point may no longer be a key point. In order to achieve the robustness of the key points, the pixels around the key points are generally described by a vector, which is called a descriptor. The key points and the descriptors together form a feature point. Researchers in the field of computer vision have designed many more stable local image features in years of research, such as scale-invariant feature transform (SIFT), speeded up robust features (SURF), and ORB algorithm.

[0033] In order to improve the reliability of the pixel coordinate values corresponding to the feature points in the image and the first real coordinate values corresponding to the pixel coordinate values in the real world coordinate system, that is, to reduce the model effect deviation caused by the error of the training data and the test data itself in the model training process. The template, i.e. the preset template, can be photographed, and the feature points can be identified more specifically and accurately when the preset template is used, and the reliability of the pixel coordinate values corresponding to the feature points and the first real coordinate values corresponding to the pixel coordinate values in the real world coordinate system is greatly increased, which provides good support for subsequent parameter optimization.

[0034] For example, in the field of image processing, there are many different categories of feature points that can meet different scene requirements. Therefore, in machine vision detection, a calibration board is often used. A calibration board with a fixed pitch array is photographed by a camera and a calibration algorithm is calculated to obtain a high-precision measurement result, as shown in Figure 3 .

[0035] The analysis of visual positioning cannot be separated from the reference coordinate system, including but not limited to pixel coordinate system, image coordinate system, camera coordinate system, and world coordinate system, wherein the pixel coordinate system uses pixels as units, and the coordinate origin is in the upper left corner; the image coordinate system uses physical units to represent the position of the pixel, and the coordinate origin is the intersection position of the camera optical axis and the image physical coordinate system, with units of mm; the camera coordinate system takes the camera optical center as the origin, with the z-axis coinciding with the optical axis, i.e. the z-axis pointing to the front of the camera, with units of m; the world coordinate system can represent any object according to the situation, with units of m.

[0036] For example, the mapping relationship between the coordinate in the pixel coordinate system and the coordinate feature point in the world coordinate system is mainly obtained in this step, such as collecting pictures of different angles of the position of the calibration board by the visual acquisition system, analyzing the collected pictures by using the SIFT technology, and then obtaining the feature point information in the picture to obtain the pixel coordinates of the feature points in the pixel coordinate system. Then, the position of the calibration board is selected as the world coordinate system, and then the coordinates of the corresponding feature points in the world coordinate system are obtained, and the mapping relationship between the pixel coordinates of the feature points and the coordinates in the real world is constructed, and then the pixel coordinate value of the feature points in the image and the corresponding real coordinate value of the pixel coordinate value in the real world coordinate system are obtained.

[0037] Step S2: based on the camera intrinsic parameter, the rotation matrix and the translation matrix of the visual device, the second real coordinate value corresponding to the first pixel coordinate value in the real world coordinate system is obtained according to the first pixel coordinate value of the feature point.

[0038] Exemplarily, the image pixel coordinate system P d (u d ,v d ) is converted into the point P(X W ,Y W ,Z W ) in the real world coordinate system. The transformation relationship from the pixel coordinate system to the camera coordinate system is obtained by the camera intrinsic parameter, and the transformation relationship of the camera coordinate system relative to the world coordinate system, i.e. the rotation matrix and the translation matrix, are obtained, and then the pixel coordinate value to the second real coordinate value of the real world coordinate system is obtained.

[0039] In some specific embodiments, the initial intrinsic parameter of the visual device after calibration is obtained as the camera intrinsic parameter; and the rotation matrix and the translation matrix of the visual device are obtained according to the camera intrinsic parameter of the visual device and the hand-eye calibration principle.

[0040] Exemplarily, the hand-eye calibration is a conversion matrix from two-dimensional coordinates to three-dimensional coordinates, that is, a conversion matrix from the pixel coordinate system to the world coordinate system. In the actual control process, after the visual device detects the pixel position of the target in the image, the pixel coordinates of the camera are transformed to the spatial coordinate system through the calibrated coordinate conversion matrix. The composition of the conversion matrix used depends on the camera intrinsic parameter, the rotation matrix and the translation matrix.

[0041] The camera intrinsic parameter can be understood as a model for converting the camera coordinate system and the pixel coordinate system. First, the conversion process of the data in the camera coordinate system to the image coordinate system is that, in the camera coordinate system, the coordinates of a point P1 in space are (X1, Y1, Z1), and the coordinates of the imaging point P2 in the image coordinate system are (X2, Y2, Z2). According to the perspective projection relationship, then:

[0042]

[0043] Wherein, f is the camera focal length, f = Z2, then through the formula 1 can be obtained:

[0044]

[0045]

[0046] Through the formula 2 and formula 3 can be obtained:

[0047]

[0048] Again, the conversion from the image coordinate system to the pixel coordinate system, the pixel coordinate system and the image coordinate system are both in the imaging plane, only the origin and the unit of measurement are different. The unit of the image coordinate system is mm, which belongs to the physical coordinate system, while the unit of the pixel coordinate system is pixel, and the description of a pixel point in the pixel coordinate system is several rows and several columns. Therefore, the conversion relationship between the image coordinate system and the pixel coordinate system can be understood as follows: dx and dy represent how many mm each column and each row represents, that is, 1 pixel = dx mm. Further, the coordinates of the imaging point P2 in the image coordinate system are (X2, Y2, Z2), as shown in Figure 4 The conversion relationship of the coordinates of the imaging point in the pixel coordinate system is (u, v) as follows:

[0049]

[0050] In formula 5, u0 and v0 represent the center of the image plane in the image coordinate system. Convert formula 5 to homogeneous coordinates and matrix form as follows:

[0051]

[0052] Therefore, the relationship between the pixel coordinate system and the camera coordinate system is as follows:

[0053]

[0054] According to formula 7, the camera intrinsic parameters can be represented as:

[0055]

[0056] The conversion relationship between the camera coordinate system and the world coordinate system can be understood as the conversion between the coordinate system and the coordinate system, and the conversion relationship between the coordinate system and the coordinate system can be realized by rotation and translation. Therefore, the conversion relationship between the camera coordinate system and the world coordinate system is composed of a rotation matrix and a translation matrix. The camera coordinate system is X C, the world coordinate system is X, then Xc = RX + T can be obtained, wherein R represents a 3*3 rotation matrix, and T represents a 3*1 translation matrix or a translation vector.

[0057] In some embodiments, the camera intrinsic parameters of the calibrated visual device are acquired, including: constructing a matrix equation according to a camera calibration principle to acquire the camera intrinsic parameters and distortion parameters of the visual device; and performing distortion calibration on the pixel coordinate values of the key points according to the distortion parameters to acquire the calibrated first pixel coordinate values.

[0058] Exemplarily, due to the influence of internal factors such as process, installation level, etc. in the manufacturing process of the camera, and external factors such as temperature, humidity, pressure, etc. in the use process, the camera is not an ideal pinhole imaging model in the imaging process, and there are distortion parameters in the camera sensor. Therefore, there is a certain distortion between the actual imaging and the ideal imaging in the image coordinate system. Therefore, it is necessary to introduce distortion parameters for distortion calibration of the pixel coordinates.

[0059] When the point under the world coordinate system corresponds to its position in the image coordinate system, a nonlinear transformation occurs on the model, and the deviation is different with the position of the point on the image plane, according to the way the lens affects the image geometry. The optical distortion can be divided into radial distortion and tangential distortion. The model of the nonlinear distortion of the optical system is:

[0060]

[0061] In the above formula, x = x'-x 0 , y = y'-y0, x', y' are the measured coordinates of the image point, and x0, y0 are the image center point offset in the camera intrinsic parameters, that is, the c x ,,c y in the camera model intrinsic parameter matrix is as follows

[0062] In some embodiments, the matrix equation is constructed according to the camera calibration principle to acquire the camera intrinsic parameters and distortion parameters of the visual device, including: capturing a black and white calibration board from multiple angles, selecting multiple images captured from multiple angles; and constructing a matrix equation by using Zhang Zhengyou camera calibration method using the multiple images, solving the matrix equation, and obtaining the camera intrinsic parameters and distortion parameters of the visual device.

[0063] Exemplarily, a Zhang Zhengyou calibration method chessboard is prepared, the size of the chessboard is known, different angle shooting is performed on the chessboard by a camera to obtain a group of images; feature points in the images such as the corner points of the calibration board are detected to obtain pixel coordinate values of the corner points of the calibration board, and the physical coordinate system of the corner points of the calibration board is calculated according to the known size of the chessboard and the origin of the world coordinate system; and then the camera intrinsic parameter, the rotation matrix, the translation matrix and the distortion coefficient are solved.

[0064] For example, the points on the calibration board plane and the points on the image plane have a one-to-one correspondence, the points on the calibration board plane are denoted as M=(x, y, z) T , the points on the image plane are denoted as m=(u, v) T , and the corresponding homogeneous coordinates are

[0065] The rotation matrix is a 3*3 matrix, and the translation matrix is a 3*1 matrix, which is defined as

[0066] r1, r2 and r3 are direction vectors of the three coordinate axes of the camera coordinate in the world coordinate, and are perpendicular to each other, and t is a translation vector from the origin of the world coordinate system to the optical center. The camera intrinsic parameter formula 8 is denoted by A. Then the following formula can be obtained through coordinate transformation s is a constant.

[0067] Suppose that the calibration board plane is located on the xy plane of the world coordinate system, i.e. z=0, and the following formula can be obtained:

[0068]

[0069] Still, M represents the coordinates of the points on the calibration board plane, M=(x, y) T , In this way, the following formula can be obtained H= *A*(r1, r2, t)=(h1, h2, h3).

[0070] Through a plurality of groups of corresponding points, the least square method is applied to solve H, wherein the calculation of H is a process of minimizing the error between the actual image coordinates and the predicted coordinates calculated through M. After H is solved, two basic constraints of the camera intrinsic parameter are constructed by using H=λ*A*(r1, r2, t)=(h1, h2, h3) and the orthogonality of r1 and r2. Then the camera intrinsic parameter is solved, and when the camera intrinsic parameter is determined, the rotation matrix and the translation matrix are also determined. Finally, the distortion parameter is evaluated by the least square method, and a more accurate value is obtained.

[0071] Step S3: determining the motion position of the robot according to the second real coordinate value.

[0072] Exemplarily, the motion parameters of each joint of the robot are adjusted according to the second real coordinates, so that the robot can reach the specified motion position, thereby achieving the control of the robot to reach the established target position and posture.

[0073] Step S4: Based on the camera intrinsic parameters, rotation matrix and translation matrix of the vision device, the second pixel coordinate value corresponding to the key point in the image is obtained according to the second real coordinate value.

[0074] Exemplarily, the second pixel coordinate value corresponding to the key point in the image is obtained according to the second real coordinate value and the camera intrinsic parameters, rotation matrix and translation matrix of the vision device, that is, the predicted pixel coordinate value, and the formula is as shown in formula (11). It represents the process of projecting the point P (X W ,Y W ,Z W ) in the real world coordinate system into the image pixel coordinate system P d (u d ,v d ).

[0075]

[0076] Where R represents the rotation matrix, t represents the translation matrix, K represents the camera intrinsic parameters, Δx and Δy represent the distortion parameters of the camera, and the specific meanings are as shown in formula 9.

[0077] Step S5: An optimization function is established according to the first pixel coordinate value and the second pixel coordinate value corresponding to the feature points in the image.

[0078] Exemplarily, according to the difference between the first pixel coordinate and the second pixel coordinate value corresponding to the feature points in the image, the difference between the first pixel coordinate and the second pixel coordinate value is continuously reduced through optimization, and the accuracy of the robot motion is improved.

[0079] In some specific embodiments, the optimization function is established according to the first pixel coordinate value and the second pixel coordinate value corresponding to the feature points in the image, including: calculating the error between the imaging point coordinate of the key point in the image on the pixel plane and the visual projection point coordinate obtained by measurement according to the first pixel coordinate value and the second pixel coordinate value, and constructing the optimization function.

[0080] Exemplarily, in order to solve the problem by using the iterative optimization method, the optimization content comes from the deviation between the imaging point coordinate of the point in the real world after transformation on the camera pixel plane and the visual projection point coordinate obtained by measurement, which is used to evaluate the fitness degree of the model and the actual data. In an ideal case, the deviation should be zero, but in practice, due to the error of the model itself, the smaller the deviation, the higher the fitness degree of the model and the actual situation.

[0081] For example, the optimization function can be designed as the two-norm of the deviation between the coordinates of the imaging points on the camera pixel plane after the point transformation in the real world and the coordinates of the visual projection points measured by the measurement, so as to evaluate the fitting degree of the model and the actual data. The two-norm is intuitive relative to the square of the distance in the Euclidean space.

[0082] In some embodiments, the optimization function is constructed according to the error, including: constructing the optimization function according to the error and a robust kernel function.

[0083] Exemplarily, the combination of the error and the robust kernel function can be used to improve the robust performance of the optimization solving problem. The robust kernel function is introduced to further process the error, artificially reduce the error term that is too large, and reduce the influence of the error data by reducing the weight of some points. Many robust kernel functions are piecewise functions that give a linear growth rate when the input is large, such as cauchy kernel, huber kernel, etc.

[0084] Step S6: updating the camera intrinsic parameter, the rotation matrix and the translation matrix of the visual device according to the optimization function.

[0085] Exemplarily, the camera intrinsic parameter, the rotation matrix and the translation matrix are taken as optimization variables according to the optimization function, and an optimization algorithm is used for iterative optimization to obtain the optimal values of the camera intrinsic parameter, the rotation matrix and the translation matrix under the current condition.

[0086] Please refer to Figure 5 In some embodiments, step S6 includes steps S61 to S62.

[0087] Step S61: constructing a nonlinear optimization model according to the optimization function.

[0088] Exemplarily, the initial value of the camera intrinsic parameter, the rotation matrix and the translation matrix obtained by calculation is taken as the initial value of iteration of the nonlinear optimization algorithm, the camera intrinsic parameter, the rotation matrix and the translation matrix are taken as optimization variables according to the optimization function, and the nonlinear optimization algorithm is used for iterative optimization, so as to effectively improve the calculation speed of iteration and the calibration accuracy. The nonlinear optimization algorithm can be used for iterative optimization to obtain the optimized camera intrinsic parameter, the rotation matrix and the translation matrix.

[0089] For example, the optimization function is designed as the two-norm of the deviation between the coordinates of the imaging points on the camera pixel plane after the point transformation in the real world and the coordinates of the visual projection points measured by the measurement: Wherein x represents the deviation between the coordinates of the imaging points on the camera pixel plane after the point transformation in the real world and the coordinates of the visual projection points measured by the measurement, and f is an arbitrary nonlinear function.

[0090] Step S62: updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device by using nonlinear optimization according to the nonlinear optimization model.

[0091] Exemplarily, let the reciprocal of the objective function be zero, and then solve the optimal value of x, which is the same as solving the extreme value of a binary function. The extreme value is obtained at the point where the derivative is zero. They can be the values at the maximum, minimum or saddle point, and it is only necessary to compare their function values one by one. However, the equation is not convenient to solve directly. Therefore, an iterative method can be used, starting from an initial value, and constantly updating the current optimization variable to make the objective function decrease. Let the solution of the derivative function being zero become a process of constantly finding the gradient and descending. Until the increment is very small and the function cannot be further decreased. At this time, it can be judged that the algorithm converges, and the objective is reached a minimum, and the process of selecting the minimum value is completed. In this process, only the gradient direction of the iterative point needs to be found, and the global derivative function being zero does not need to be found.

[0092] For example, in the prior art, Gauss-Newton method or Levenberg-Marquardt method can be used to solve the optimization problem, and then the optimized camera intrinsic parameters, rotation matrix and translation matrix are obtained.

[0093] Please refer to Figure 5 , Figure 4 The structure of the terminal provided in the embodiments of the present application is shown in a schematic block diagram.

[0094] As shown in Figure 4 , the terminal 300 includes a processor 301 and a memory 302, and the processor 301 and the memory 302 are connected through a bus 303, such as an I2C (Inter-integrated Circuit) bus.

[0095] Specifically, the processor 301 is configured to provide computing and control capabilities to support the operation of the entire server. The processor 301 can be a central processing unit (CPU), and the processor 301 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor.

[0096] Specifically, the memory 302 can be a Flash chip, a Read-Only Memory (ROM) disk, an optical disk, a U disk, or a mobile hard disk, etc.

[0097] Those skilled in the art can understand that, ​ The structure shown in FIG. 3 is only a block diagram of part of the structure related to the embodiments of the present application, and does not constitute a limitation on the terminal to which the embodiments of the present application are applied. Specifically, the terminal can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0098] The processor 301 is configured to run the computer program stored in the memory and implement the visual positioning method provided in any of the embodiments of the present application when the computer program is executed.

[0099] In some embodiments, the processor 301 is configured to run the computer program stored in the memory and implement the following steps when the computer program is executed:

[0100] obtain an image captured by a visual device on a robot, and obtain a first pixel coordinate value of a feature point in the image and a first real coordinate value corresponding to the first pixel coordinate value in a real world coordinate system;

[0101] obtain a second real coordinate value corresponding to the first pixel coordinate value in the real world coordinate system based on the camera intrinsic parameter, the rotation matrix, and the translation matrix of the visual device, according to the first pixel coordinate value of the feature point;

[0102] determine a motion position of the robot according to the second real coordinate value;

[0103] obtain a second pixel coordinate value corresponding to the feature point in the image based on the camera intrinsic parameter, the rotation matrix, and the translation matrix of the visual device, according to the second real coordinate value;

[0104] establish an optimization function according to the first pixel coordinate value and the second pixel coordinate value corresponding to the feature point in the image;

[0105] update the camera intrinsic parameter, the rotation matrix, and the translation matrix of the visual device according to the optimization function.

[0106] In some embodiments, the processor 301 performs the following steps in the process of obtaining the image captured by the visual device on the robot:

[0107] controls the visual device of the robot to capture a preset template to obtain the image, the preset template including at least one feature point;

[0108] Identify feature points in the image to obtain feature points in the image.

[0109] In some embodiments, the processor 301 further performs, during the method:

[0110] Obtain the initial intrinsic parameters of the vision device after calibration as the camera intrinsic parameters of the vision device.

[0111] According to the camera intrinsic parameters of the vision device and the hand-eye calibration principle, obtain the rotation matrix and the translation matrix of the vision device.

[0112] In some embodiments, the processor 301 performs, during the process of obtaining the camera intrinsic parameters of the calibrated vision device:

[0113] According to the camera calibration principle, construct a matrix equation to obtain the camera intrinsic parameters and distortion parameters of the vision device.

[0114] According to the distortion parameters, perform distortion correction on the pixel coordinate values of the feature points of the vision device to obtain the first pixel coordinate values after correction.

[0115] In some embodiments, the processor 301 performs, during the process of constructing a matrix equation according to the camera calibration principle to obtain the camera intrinsic parameters and distortion parameters of the vision device:

[0116] Capture the black and white calibration board from multiple angles, and select multiple images captured from multiple angles.

[0117] Use the multiple images to construct a matrix equation using the Zhang Zhengyou camera calibration method, solve the matrix equation, and obtain the camera intrinsic parameters and distortion parameters of the vision device.

[0118] In some embodiments, the processor 301 performs, during the process of establishing an optimization function according to the first pixel coordinate values and the second pixel coordinate values corresponding to the feature points in the image:

[0119] According to the first pixel coordinate values and the second pixel coordinate values, calculate the error between the imaging point coordinates of the feature points in the image on the pixel plane and the visual projection point coordinates obtained by measurement, and construct the optimization function.

[0120] In some embodiments, the processor 301 performs, during the process of constructing the optimization function according to the error:

[0121] According to the error and the robust kernel function, construct the optimization function.

[0122] In some embodiments, the processor 301 performs the following in the process of updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the visual device according to the optimization function:

[0123] constructing a nonlinear optimization model according to the optimization function;

[0124] updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the visual device by using nonlinear optimization according to the nonlinear optimization model.

[0125] It should be noted that, for the convenience and brevity of description, the specific working process of the terminal described above can refer to the corresponding process in the foregoing visual positioning method embodiments, which will not be described here.

[0126] The embodiments of the present application also provide a storage medium for computer readable storage, the storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the visual positioning methods provided by the embodiments of the present application.

[0127] The storage medium can be an internal storage unit of the terminal, such as a terminal memory. The storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc.

[0128] Those skilled in the art can understand that all or some of the steps in the methods disclosed above and the functional modules / units in the devices can be implemented by software, firmware, hardware, or a combination thereof. In hardware embodiments, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on computer-readable media, which can include computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and can include any information delivery media.

[0129] It should be understood that the term "and / or" as used herein refers to any combination of associated listed items, and all possible combinations, and includes these combinations. It should be noted that the terms "comprising", "including", or any other variant thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article or system that comprises a list of elements does not include only those elements, but can also include other elements not expressly listed or inherent to such process, method, article or system. Without more limitations, an element defined by the phrase "comprising a" does not exclude the presence of additional identical elements in the process, method, article or system including the element.

[0130] The above sequence numbers of the embodiments of the present application are only for description, and do not represent advantages or disadvantages of the embodiments. The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of visual positioning, characterized by The method comprises: acquiring an image captured by a vision device on a robot, and acquiring first pixel coordinate values of feature points in the image and first real coordinate values corresponding to the first pixel coordinate values in a real world coordinate system; based on camera intrinsic parameters, a rotation matrix and a translation matrix of the vision device, acquiring second real coordinate values corresponding to the first pixel coordinate values in the real world coordinate system according to the first pixel coordinate values of the feature points; determining a motion position of the robot according to the second real coordinate values; based on the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device, acquiring second pixel coordinate values corresponding to the second real coordinate values in the image according to the second real coordinate values; establishing an optimization function according to the first pixel coordinate values and the second pixel coordinate values of the feature points in the image; updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device according to the optimization function; the establishing of the optimization function according to the first pixel coordinate values and the second pixel coordinate values of the feature points in the image comprises: calculating errors between imaging point coordinates of the feature points in the image on a pixel plane and vision projection point coordinates obtained through measurement according to the first pixel coordinate values and the second pixel coordinate values, and constructing the optimization function; the constructing of the optimization function according to the errors comprises: constructing the optimization function according to the errors and a robust kernel function; the updating of the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device according to the optimization function comprises: constructing a nonlinear optimization model according to the optimization function, and updating the camera intrinsic parameters, the rotation matrix and the translation matrix of the vision device by using nonlinear optimization according to the nonlinear optimization model.

2. The method of claim 1, wherein, the acquiring of the image captured by the vision device on the robot comprises: controlling the vision device of the robot to capture a preset template to obtain the image, the preset template comprising at least one feature point; identifying the feature points in the image to obtain the feature points in the image.

3. The method of claim 1, wherein, the method further comprises: acquiring initial intrinsic parameters of the vision device after calibration as the camera intrinsic parameters; acquiring the rotation matrix and the translation matrix of the vision device according to the camera intrinsic parameters of the vision device and a hand-eye calibration principle.

4. The method of claim 3, wherein, the acquiring of the camera intrinsic parameters of the vision device after calibration comprises: constructing a matrix equation according to a camera calibration principle to acquire the camera intrinsic parameters and distortion parameters of the vision device; performing distortion calibration on the pixel coordinate values of the feature points of the vision device according to the distortion parameters to acquire the first pixel coordinate values after calibration.

5. The method of claim 4, wherein, the constructing of the matrix equation according to the camera calibration principle to acquire the camera intrinsic parameters and the distortion parameters of the vision device comprises: capturing a black and white calibration board from multiple angles to select multiple images captured from multiple angles; constructing a matrix equation by using the multiple images by using a camera calibration method of Zhang Zhengyou to solve the matrix equation to obtain the camera intrinsic parameters and the distortion parameters of the vision device.

6. A robot terminal, characterized in that, the terminal comprises a processor and a memory; The memory is configured to store a computer program. The processor is configured to execute the computer program and implement the visual positioning method according to any one of claims 1 to 5 when the computer program is executed.

7. A computer storage medium for computer storage, comprising: The computer storage medium stores one or more programs, and the one or more programs are executable by one or more processors to implement the steps of the visual positioning method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Hand-eye system self-calibration method based on active visual sense

    CN106940894A

  • Camera calibration method and device, user equipment and storage medium

    CN109754434A