A method and apparatus for determining a target point
By introducing a third camera into the binocular vision positioning system and performing multi-dimensional stereo matching and camera parameter calibration, the problem of positioning error of non-real target objects in the binocular vision positioning system was solved, and the accurate positioning of the target point was achieved.
Patent Information
- Application Number
- CN202211112502.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-13
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2042-09-13
AI Technical Summary
In the process of using a binocular vision positioning system to determine a target object, there are positioning or tracking errors of non-real target objects, resulting in inaccurate positioning.
A third camera is introduced. The plane formed by the optical centers of the third camera and the left and right cameras does not intersect with the effective field of view of the binocular vision positioning device. Spatial points are obtained through multi-dimensional stereo matching processing. Combined with camera parameter calibration, the world coordinates of the target point are determined.
It improves the accuracy of target point positioning, can judge spatial points from different dimensions, reduces errors from non-real spatial points, and achieves accurate positioning of target points.
Smart Images

Figure CN115546261B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computer vision, and more particularly, to a method and device for determining a target point in the field of computer vision. BACKGROUND
[0002] With the high demand for location information technology, positioning technology has been well developed. Among them, a binocular vision positioning system based on the principle of parallax.
[0003] The binocular vision positioning system includes two cameras, the optical centers of which are located on the same baseline. The two cameras respectively acquire a target object in the same field of view at different shooting angles to obtain the projection points of the target object on the respective images of the two cameras. The binocular vision positioning system performs stereo matching on the projection points of the target object on the respective images captured by the two cameras to obtain the spatial position of the target object; and obtains the spatial coordinates of the target object according to the positions of the projection points of the target object on the respective images captured by the two cameras and the principle of similar triangle measurement, so as to position or track the target object.
[0004] In the process of determining the target object by using the binocular vision positioning system, the stereo matching processing of the projection points collected by the two cameras using the binocular vision positioning system and having epipolar constraint relationship can obtain multiple target objects, and there may be non-real target objects in the multiple target objects, which causes errors in the positioning or tracking of the target object. SUMMARY
[0005] The present application provides a method and device for determining a target point, which can identify the position of a real target object and accurately position the target point.
[0006] In a first aspect, a method for determining a target point is provided, which is applied to a visual positioning system, the visual positioning system comprising a binocular visual positioning device and a third camera, the binocular visual positioning device comprising a left camera and a right camera for capturing images, and a plane formed by an optical center of the third camera, an optical center of the left camera and an optical center of the right camera is not intersected with an effective field of view of the binocular visual positioning device, the method further comprising: obtaining N1 left feature points in a left image of a target picture captured by the left camera, and obtaining M1 right feature points in a right image of the target picture captured by the right camera, the N1 left feature points and the M1 right feature points having a epipolar constraint relationship, and M1 and N1 are integers greater than 0; performing stereo matching processing on the N1 left feature points and the M1 right feature points to obtain P spatial points, and P is an integer greater than 0; obtaining T1 feature points of the target picture captured by the third camera, and T1 is an integer greater than 0; performing stereo matching processing on the N1 left feature points and the T1 feature points to obtain Q spatial points, and Q is an integer greater than 0; and determining the target point according to the P spatial points and the Q spatial points.
[0007] In the above technical solution, the third camera is introduced in the case of obtaining the target point by the original binocular visual positioning device, the optical center of the third camera, the optical center of the left camera and the optical center of the right camera form a first plane, and the first plane is not intersected with the effective field of view of the binocular visual positioning device, so that any point in the effective field of view is not on the first plane. In this way, the N1 left feature points obtained by the left camera and the M1 right feature points obtained by the right camera can be used for stereo matching processing to obtain P spatial points, and the N1 left feature points obtained by the left camera and the T1 feature points obtained by the third camera can be used for stereo matching processing to obtain Q spatial points, so that the target point can be determined according to the P spatial points and the Q spatial points. That is, the spatial points are obtained by cameras from different dimensions, and the obtained spatial points are judged by combining multiple dimensions to obtain the target point.
[0008] In combination with the first aspect, in some possible implementation manners, the method further comprises: obtaining N left feature points in a left image of a target picture captured by the left camera, and obtaining M right feature points in a right image of the target picture captured by the right camera, the N left feature points comprising the N1 left feature points, and the M right feature points comprising M1 right feature points; performing stereo matching processing on N-N1 left feature points and M-M1 right feature points to obtain O spatial points, the N-N1 left feature points and the M-M1 right feature points not having an epipolar constraint relationship, and O is an integer greater than 0; and determining the target point according to the P spatial points and the Q spatial points, comprising: determining the target point according to the P spatial points, the Q spatial points and the O spatial points.
[0009] With reference to the first aspect, in some possible implementation manners, the target point is determined according to the P spatial points, the Q spatial points and the O spatial points, including: determining, as R spatial points, spatial points that have a distance less than or equal to a preset value from any one of the P spatial points and the Q spatial points, R being an integer greater than 0; and determining the R spatial points and the O spatial points as the target point.
[0010] In the technical solution, spatial points that have a distance less than or equal to a preset value from any one of the P spatial points and the Q spatial points are determined as R spatial points. This means that the same spatial point in the P spatial points and the Q spatial points is regarded as a point in the R spatial points, that is, the same spatial point is seen from different dimensions, and thus the spatial point can be regarded as a point in the R spatial points. In addition, the spatial points obtained by stereo matching of the feature points that do not have the epipolar constraint relationship do not have the case of non-real spatial points, and thus the feature points that do not have the epipolar constraint relationship obtained by the left camera and the right camera can be stereo matched to obtain the O spatial points, and the O spatial points are also regarded as the target point.
[0011] With reference to the first aspect and the implementation manners above, in some possible implementation manners, the method further includes: calibrating the left camera and the right camera to determine camera parameters of the left camera and the right camera, with the optical center of the left camera as the origin of a world coordinate system; and calibrating the left camera and the third camera to determine camera parameters of the third camera, with the optical center of the left camera as the origin of the world coordinate system.
[0012] Alternatively,
[0013] calibrating the left camera and the right camera to determine camera parameters of the left camera and the right camera, with the optical center of the right camera as the origin of a world coordinate system; and calibrating the right camera and the third camera to determine camera parameters of the third camera, with the optical center of the right camera as the origin of the world coordinate system.
[0014] It should be understood that the camera parameters refer to the extrinsic and intrinsic parameters of the cameras. The intrinsic parameters of the cameras are determined by the cameras themselves and only related to the cameras themselves, and specifically include the focal length of the cameras, camera lens distortion parameters and pixel size. Among them, the focal lengths of the left camera, the right camera and the third camera are generally equal; the camera lens distortion parameters represent the magnitude of radial distortion; and the pixel size refers to the length and width actually represented by a pixel. The extrinsic parameters of the cameras refer to the poses of the cameras in the world coordinate system, which are determined by the relative pose relationship between the cameras and the world coordinate system, and specifically include a rotation vector and a translation vector. Among them, the rotation vector describes the direction in which the coordinate axes of the world coordinate system correspond to the coordinate axes of the cameras; and the translation vector describes the position of the spatial origin of the world coordinate system in the camera coordinate system.
[0015] In the above technical solution, the optical center of the left camera is taken as the origin of the world coordinate system, and the left camera and the right camera, and the left camera and the third camera are calibrated in order to obtain the spatial geometric position relationship of the left camera and the right camera, the intrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera and the intrinsic parameters of the third camera. Among them, the spatial geometric position relationship of the left camera and the right camera refers to the extrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera refers to the extrinsic parameters of the third camera. The optical center of the right camera is taken as the origin of the world coordinate system, and the left camera and the right camera, and the left camera and the third camera are calibrated in order to obtain the spatial geometric position relationship of the left camera and the right camera, the intrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera and the intrinsic parameters of the third camera. Among them, the spatial geometric position relationship of the left camera and the right camera refers to the extrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera refers to the extrinsic parameters of the third camera. Furthermore, each calibration in the two calibrations needs to use the origin of the same camera as the origin of the world coordinate system, so that unified world coordinate information can be obtained. Thus, the obtained camera parameters can contribute to the subsequent determination of the world coordinates of the target point.
[0016] In combination with the first aspect and the above implementation manners, in some possible implementation manners, the method further includes: determining the world coordinates of the target point according to the pixel coordinates of the target point on the image in the left camera, the pixel coordinates of the target point on the image in the right camera, the camera parameters of the left camera and the camera parameters of the right camera; or determining the world coordinates of the target point according to the pixel coordinates of the target point on the image in the left camera, the pixel coordinates of the target point on the image in the third camera, the camera parameters of the left camera and the camera parameters of the third camera.
[0017] In the technical solution, the world coordinates of the target point can be determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the right camera, the camera parameters of the left camera and the camera parameters of the right camera, or the world coordinates of the target point can be determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the third camera, the camera parameters of the left camera and the camera parameters of the third camera. That is, through the above process, the accurate position of the target point can be determined, which has great application value in the medical field. For example, in the medical field, the position of the patient's diseased tissue can be positioned, and the position of the surgical instrument can be tracked in real time.
[0018] In combination with the first aspect and the above implementation, in some possible implementation, the left camera and the right camera are both black-and-white cameras, and the third camera is a color camera.
[0019] In summary, the present application provides a method for determining a target point. By introducing a third camera when the original binocular vision positioning device obtains the target point, the optical center of the third camera, the optical center of the left camera and the optical center of the right camera form a first plane, which does not intersect with the effective field of view of the binocular vision positioning device. In this way, it can be ensured that any point in the effective field of view does not fall on the first plane. In this way, the N1 left feature points obtained by the left camera and the M1 right feature points obtained by the right camera which have epipolar line constraint relationship can be used for stereo matching processing to obtain P space points; the N1 left feature points obtained by the left camera and the T1 feature points obtained by the third camera can be used for stereo matching processing to obtain Q space points; and thus, the target point can be determined according to the P space points and the Q space points. That is, the space points can be obtained from different dimensions, and the obtained space points can be judged by combining multiple dimensions to obtain the target point.
[0020] In addition, any one of the P space points and the Q space points whose distance is less than or equal to a preset value is determined as an R space point. That is, the same space point in the P space points and the Q space points is regarded as a point in the R space points, that is, the same space point is seen from different dimensions, and thus the space point can be regarded as a point in the R space points. In addition, the space points obtained by stereo matching of the feature points which do not have epipolar line constraint relationship do not have the case of non-real space points, and thus the feature points which do not have epipolar line constraint relationship obtained by the left camera and the right camera can be stereo matched to obtain O space points, and the O space points are also regarded as the target point.
[0021] It should be understood that before the left camera, the right camera and the third camera included in the visual positioning system capture images, the left camera and the right camera, and the left camera and the third camera need to be calibrated. This is to obtain the camera parameters, that is, the positional relationship of the left camera, the right camera and the third camera in space geometry and the intrinsic parameters of the left camera, the right camera and the third camera. Therefore, the obtained camera parameters can contribute to the subsequent determination of the world coordinates of the target point.
[0022] Finally, the world coordinates of the target point can be determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the right camera, the camera parameters of the left camera and the camera parameters of the right camera. That is, through the above process, the accurate position of the target point can be determined, which has great application value in the medical field. For example, it can assist in positioning the position of the patient's diseased tissue in the medical field, and can track the position of the surgical instrument in real time.
[0023] In a second aspect, a device for determining a target point is provided. The device includes: an obtaining module configured to obtain N1 left feature points in a left image of a target picture captured by a left camera, and M1 right feature points in a right image of the target picture captured by a right camera, the N1 left feature points and the M1 right feature points having a epipolar constraint relationship, M1 and N1 being integers greater than 0; a processing module configured to perform stereo matching processing on the N1 left feature points and the M1 right feature points to obtain P spatial points, P being an integer greater than 0; the obtaining module is further configured to obtain T1 feature points of the target picture captured by the third camera, T1 being an integer greater than 0; the processing module is further configured to perform stereo matching processing on the N1 left feature points and the T1 feature points to obtain Q spatial points, Q being an integer greater than 0; and a determining module configured to determine a target point based on the P spatial points and the Q spatial points.
[0024] In combination with the second aspect, in some possible implementation manners, the obtaining module is further configured to obtain N left feature points in the left image of the target picture captured by the left camera, and M right feature points in the right image of the target picture captured by the right camera, the N left feature points including the N1 left feature points, and the M right feature points including the M1 right feature points; the processing module is further configured to perform stereo matching processing on N-N1 left feature points and M-M1 right feature points to obtain O spatial points, the N-N1 left feature points and the M-M1 right feature points not having an epipolar constraint relationship, and O being an integer greater than 0; and the determining module is specifically configured to determine the target point based on the P spatial points, the Q spatial points and the O spatial points.
[0025] With reference to the second aspect, in some possible implementation manners, the determining module, in particular, is further configured to determine, as the R spatial points, spatial points in the P spatial points and spatial points in the Q spatial points that are less than or equal to a preset value away from any one of the spatial points; and determine the R spatial points and the O spatial points as the target points.
[0026] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determining module is further configured to calibrate the left camera and the right camera, to determine camera parameters of the left camera and the right camera, with the optical center of the left camera as the origin of the world coordinate system; and calibrate the left camera and the third camera, to determine camera parameters of the third camera, with the optical center of the left camera as the origin of the world coordinate system; or calibrate the left camera and the right camera, to determine camera parameters of the left camera and the right camera, with the optical center of the right camera as the origin of the world coordinate system; and calibrate the right camera and the third camera, to determine camera parameters of the third camera, with the optical center of the right camera as the origin of the world coordinate system.
[0027] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the determining module is further configured to determine the world coordinates of the target point according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the right camera, the camera parameters of the left camera, and the camera parameters of the right camera; or determine the world coordinates of the target point according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the third camera, the camera parameters of the left camera, and the camera parameters of the third camera.
[0028] With reference to the second aspect and the foregoing implementation manners, in some possible implementation manners, the left camera and the right camera are both black-and-white cameras, and the third camera is a color camera.
[0029] The third aspect provides a device for determining target points, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program, so that the device for determining target points executes the method in the first aspect or any possible implementation manner of the first aspect.
[0030] The fourth aspect provides a computer-readable storage medium, which stores instructions, and the instructions, when running on a computer or a processor, cause the computer or the processor to execute the method in the first aspect or any possible implementation manner of the first aspect.
[0031] In a fifth aspect, a computer program product containing instructions, which, when the computer program product runs on the computer or the processor, enables the computer or the processor to execute the method in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF DRAWINGS
[0032] Figure 1 is a schematic diagram of a stereo matching process provided by an embodiment of the present application;
[0033] Figure 2 is a schematic diagram of a description of epipolar constraint relationships of multiple projection points provided by an embodiment of the present application;
[0034] Figure 3 is a schematic diagram of a structure of a visual positioning system provided by an embodiment of the present application;
[0035] Figure 4 is a schematic flowchart of a method for determining a target point provided by an embodiment of the present application;
[0036] Figure 5 is a schematic diagram of a determination of a target point provided by an embodiment of the present application;
[0037] Figure 6 is a structural schematic diagram of a device for determining a target point provided by an embodiment of the present application;
[0038] Figure 7 is a structural schematic diagram of another device for determining a target point provided by an embodiment of the present application. DETAILED DESCRIPTION
[0039] The technical solutions in the present application will be described in detail below with reference to the drawings. In the description of the embodiments of the present application, two or more than two are meant unless otherwise specified, and in the description of the embodiments of the present application, “multiple” means two or more than two.
[0040] Hereinafter, the terms “first” and “second” are used only for descriptive purposes, and cannot be understood as implying or suggesting relative importance or implicitly indicating the number of indicated technical features. Therefore, the features defined with “first” and “second” can explicitly or implicitly include one or more of the features.
[0041] It should be understood that the implementation principle of the binocular visual positioning device is the same as the visual perception principle of the human eye. The projection points corresponding to the same object seen by the left eye and the right eye are different, and the projection points seen by the left eye and the right eye need to be matched, and the position of the target object is determined according to the matching result. Therefore, stereo matching processing is introduced in the binocular visual positioning device.
[0042] The "stereo matching processing" is a process of determining a target object by matching corresponding one or more projection points captured by a left camera and one or more projection points captured by a right camera. Taking an example of the number of projection points of the target object captured by the left camera and the right camera in respective images as 2, the process of determining the target object is described in detail.
[0043] Figure 1 is a schematic diagram of the stereo matching processing provided by an embodiment of the present application.
[0044] Exemplarily, when the number of projection points of the left image and the right image is 2, as shown in Figure 1 , it is assumed that P1 and P2 are two target objects in space, the projection points of the target objects captured by the left camera in the left image are P l1 and P l2 , and the projection points of the target objects captured by the right camera in the right image are P r1 and P r2 . P l1 and P l2 have an epipolar constraint relationship with P r1 and P r2 . P l1 is matched with P r1 to obtain P1, that is, C l forms a first optical axis with P l1 , P r1 forms a second optical axis with C r , and the intersection of the first optical axis and the second optical axis determines P1; similarly, P l1 is matched with P r2 to obtain M, P l2 is matched with P r1 to obtain N, and P l2 is matched with P r2 to obtain P2.
[0045] It should also be understood that the "epipolar constraint" describes a constraint formed by the projection points, the optical centers of the cameras under the projection model when the target object is projected onto two images with different viewing angles. The "epipolar constraint" can be described as: the first plane formed by the projection points of the target object on the left image, the optical center of the left camera and the optical center of the right camera, and the intersection line formed by the intersection of the first plane and the right image, then the projection point on the right image captured by the right camera must be located on the intersection line.
[0046] Figure 2 is a schematic diagram provided by an embodiment of the present application for describing the epipolar constraint relationship of multiple projection points.
[0047] Exemplarily, as shown in Figure 2 , the projection point on the left image is P l1 , Pl1 C l and C r A first plane is formed, and the line of intersection between the first plane and the right image is l. The projection point on the right image captured by the right camera must lie on the line of intersection l, i.e., P. r1 P r2 and P r3 Located on l.
[0048] There are common pole line constraints between multiple projection points on the left image and multiple projection points on the right image. During the process of determining the target point using stereo matching, it is possible to obtain non-true target points. For example, Figure 1 In this diagram, P1 and P2 are two target objects in space. Stereo matching yields four target points, meaning that M and N, obtained through matching, are not the actual target points. Therefore, it is concluded that processing multiple projection points captured by the left and right cameras using stereo matching may result in non-real target objects.
[0049] To address the aforementioned problems, this application provides a visual positioning system for determining target points.
[0050] Figure 3 This is a schematic diagram of the structure of a visual positioning system provided in an embodiment of this application. The visual positioning system includes a binocular visual positioning device and a third camera. The binocular visual positioning device includes a left camera and a right camera for acquiring images. The plane formed by the optical center of the third camera, the optical center of the left camera, and the optical center of the right camera does not intersect with the effective field of view of the binocular visual positioning device.
[0051] Among them, Figure 3 In the visual positioning system shown, the left and right cameras can form a set of binocular visual positioning devices to capture images of a target object within the same field of view. Similarly, the left and third cameras can also form a set of binocular visual positioning devices to capture images of the target object within the same field of view. Finally, the two cameras in the binocular visual positioning device can obtain the world coordinates of the target object based on the position of the projection point of the target object on their respective images and the principle of similar triangles.
[0052] For example, such as Figure 3 As shown, C l C is the optical center of the left camera. r C is the optical center of the right camera. dThe optical center of the third camera, the projection point of the target object P in the left image captured by the left camera is P1, the projection point of the target object P in the right image captured by the right camera is P2, and the projection point of the target object P in the third image captured by the third camera is P3. According to P1, P2 and the principle of similar triangles, or according to P1, P3 and the principle of similar triangles, the world coordinates of the target object P are obtained.
[0053] Figure 4 is a schematic flowchart of a method for determining a target point provided by an embodiment of the present application.
[0054] It should be understood that the method for determining a target point provided by an embodiment of the present application can be applied to a visual positioning system as shown in the figure. Figure 2 Alternatively, the method for determining a target point can also be applied to a device using the visual positioning system, which can be a robot or various mechanical devices, etc.
[0055] For example, as shown in the figure, the method 400 includes the following steps. Figure 4
[0056] 401, obtaining N1 left feature points in the left image of the target picture captured by the left camera and M1 right feature points in the right image of the target picture captured by the right camera, the N1 left feature points and the M1 right feature points having an epipolar constraint relationship, M1 and N1 being integers greater than 0.
[0057] It should be understood that the left feature points in the above scheme refer to the projection points in the left image as shown in the figure, and the right feature points refer to the projection points in the right image as shown in the figure. Figure 1 Figure 1 It should be understood that the left feature points in the above scheme refer to the projection points in the left image as shown in the figure, and the right feature points refer to the projection points in the right image as shown in the figure.
[0058] It should also be understood that the binocular visual positioning device can further include a bracket connecting the left camera and the right camera. The left camera and the right camera in the binocular visual positioning device are both cameras with adjustable shooting angles. The type of the left camera, the right camera and the third camera is not limited in the present application. The type of the left camera, the right camera and the third camera can be the same. For example, the type of the left camera, the right camera and the third camera is infrared camera or low-illumination camera or wide-dynamic camera, etc. At least one of the types of the left camera, the right camera and the third camera is different. For example, when the third camera is a black-and-white camera, the left camera and the right camera are both infrared cameras, etc. For another example, the left camera is an infrared camera, the right camera is a low-illumination camera, and the third camera is a wide-dynamic camera.
[0059] In a possible implementation, before step 401, N left feature points in the left image of the target picture captured by the left camera are acquired, and M right feature points in the right image of the target picture captured by the right camera are acquired, the N left feature points include the N1 left feature points, the M right feature points include M1 right feature points, and N and M are integers greater than 0.
[0060] It should be understood that the N1 left feature points that have the epipolar constraint relationship with the M1 right feature points in the M right feature points acquired by the right camera can be determined from the N left feature points acquired by the left camera by using an existing epipolar constraint algorithm. For example, the ransac algorithm can be used to determine that the N1 left feature points have the epipolar constraint relationship with the M1 right feature points in the M right feature points.
[0061] In a possible implementation, before the N left feature points in the left image of the target picture captured by the left camera are acquired, and the M right feature points in the right image of the target picture captured by the right camera are acquired, the left camera and the right camera are calibrated with the optical center of the left camera as the origin of the world coordinate system, the camera parameters of the left camera and the right camera are determined, and the left camera and the third camera are calibrated with the optical center of the left camera as the origin of the world coordinate system, the camera parameters of the third camera are determined; or the left camera and the right camera are calibrated with the optical center of the right camera as the origin of the world coordinate system, the camera parameters of the left camera and the right camera are determined, and the right camera and the third camera are calibrated with the optical center of the right camera as the origin of the world coordinate system, the camera parameters of the third camera are determined.
[0062] It should be understood that the camera parameters refer to the extrinsic parameters and intrinsic parameters of the camera. The intrinsic parameters of the camera are determined by the camera itself and only related to the camera itself, and specifically include the focal length of the camera, the camera lens distortion parameter, and the pixel size. The focal lengths of the left camera, the right camera, and the third camera are generally equal. The camera lens distortion parameter represents the magnitude of radial distortion. The pixel size refers to the length and width actually represented by one pixel. The extrinsic parameters of the camera refer to the pose of the camera in the world coordinate system, which is determined by the relative pose relationship between the camera and the world coordinate system, and specifically includes a rotation vector and a translation vector. The rotation vector describes the direction in which the coordinate axis of the world coordinate system corresponds to the coordinate axis of the camera. The translation vector describes the position of the spatial origin of the world coordinate system in the camera coordinate system.
[0063] In the above technical solution, the optical center of the left camera is taken as the origin of the world coordinate system, and the left camera and the right camera, and the left camera and the third camera are calibrated to obtain the spatial geometric position relationship of the left camera and the right camera, the intrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera and the intrinsic parameters of the third camera. The spatial geometric position relationship of the left camera and the right camera refers to the extrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera refers to the extrinsic parameters of the third camera. In the second calibration, the optical center of the right camera is taken as the origin of the world coordinate system, and the left camera and the right camera, and the left camera and the third camera are calibrated to obtain the spatial geometric position relationship of the left camera and the right camera, the intrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera and the intrinsic parameters of the third camera. The spatial geometric position relationship of the left camera and the right camera refers to the extrinsic parameters of the left camera and the right camera, and the spatial geometric position relationship of the third camera refers to the extrinsic parameters of the third camera. Furthermore, the origin of the same camera is used as the origin of the world coordinate system in each calibration, so that unified world coordinate information can be obtained. Finally, the obtained camera parameters can contribute to the subsequent determination of the world coordinates of the target point.
[0064] 402, the N1 left feature points and the M1 right feature points are subjected to stereo matching processing to obtain P spatial points, P being an integer greater than 0.
[0065] It should be understood that the stereo matching processing of the N1 left feature points and the M1 right feature points is to match each of the N1 left feature points with each of the M1 right feature points to obtain P spatial points.
[0066] 403, T1 feature points of a target picture captured by the third camera are obtained, T1 being an integer greater than 0.
[0067] 404, the N1 left feature points and the T1 feature points are subjected to stereo matching processing to obtain Q spatial points, Q being an integer greater than 0.
[0068] It should be understood that the stereo matching processing of the N1 left feature points and the T1 feature points is to match each of the N1 left feature points with each of the T1 feature points to obtain Q spatial points.
[0069] 405, a target point is determined according to the P spatial points and the Q spatial points.
[0070] In one possible implementation, before step 405, stereo matching is performed on N-N1 left feature points and M-M1 right feature points to obtain O spatial points. There is no common pole line constraint relationship between the N-N1 left feature points and the M-M1 right feature points, and O is an integer greater than 0. Step 405 includes: determining the target point based on the P spatial points, the Q spatial points and the O spatial points.
[0071] It should be understood that performing stereo matching on N-N1 left feature points and M-M1 right feature points involves matching each of the N-N1 left feature points with each of the M-M1 right feature points to obtain O spatial points.
[0072] In one possible implementation, determining the target point based on the P spatial points, the Q spatial points, and the O spatial points includes: determining the spatial points whose distance from any one of the P spatial points and any one of the Q spatial points is less than or equal to a preset value as R spatial points, where R is an integer greater than 0; and determining the R spatial points and the O spatial points as the target point.
[0073] In the above technical solution, spatial points whose distance from any one of the P spatial points and any one of the Q spatial points is less than or equal to a preset value are determined as R spatial points. This is equivalent to treating the same spatial point from the P spatial points and the Q spatial points as one of the R spatial points; that is, if a spatial point is seen from different dimensions, then that spatial point can be considered one of the R spatial points. Furthermore, since there are no non-real spatial points obtained from stereo matching between feature points without common-pole line constraints, stereo matching can be performed on feature points without common-pole line constraints acquired by the left and right cameras to obtain O spatial points, which are also used as target points.
[0074] Figure 5 This is a schematic diagram of determining a target point provided in an embodiment of this application.
[0075] For example, such as Figure 5 As shown, assume there are three target points in space, P1, P2, and P3. The three left feature points on the left image captured by the left camera are P1, P2, and P3. l1 P l2 and P l3 The two right feature points on the right image captured by the right camera are P. r1 and P r2 And two feature points P in the image captured by the third camera. d1 and P d2 And P l1 P l3 P r1 and Pr2 There is an epipolar constraint relationship, P l2 and P r2 There is no epipolar constraint relationship. To P l1 and P l3 , stereo matching processing is performed with P r1 and P r2 to obtain P spatial points, specifically P1, P 14 , P 23 and P2. To P l1 and P l2 , stereo matching processing is performed with P d1 and P d2 to obtain Q spatial points, specifically P 15 , P1, P2 and P 26 . P1 and P2, which exist simultaneously in the P spatial points and the Q spatial points, are taken as points in R spatial points; to P l2 and P r2 , stereo matching is performed to obtain P3 as a point in the Q spatial points, at this time the target point is a point in the R spatial points and the Q spatial points, specifically P1, P2 and P3. It should be understood that the simultaneously existing spatial points can be understood as having a distance of 0 between them.
[0076] In a possible implementation, after step 407, the world coordinates of the target point are determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the right camera, the camera parameters of the left camera and the camera parameters of the right camera; or the world coordinates of the target point are determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the third camera, the camera parameters of the left camera and the camera parameters of the third camera.
[0077] It should be understood that the world coordinates of the finally obtained target point can be determined by the pixel coordinates of the target point on the left image, the pixel coordinates of the target point on the right image, the camera parameters of the left camera and the camera parameters of the right camera, or can be determined by the pixel coordinates of the target point on the left image, the pixel coordinates of the target point on the image of the third camera, the camera parameters of the left camera and the camera parameters of the third camera.
[0078] In the technical solution, the world coordinates of the target point can be determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the right camera, the camera parameters of the left camera and the camera parameters of the right camera, or the world coordinates of the target point can be determined according to the pixel coordinates of the target point on the image of the left camera, the pixel coordinates of the target point on the image of the third camera, the camera parameters of the left camera and the camera parameters of the third camera. That is, through the above process, the accurate position of the target point can be determined, which has great application value in the medical field. For example, in the medical field, the position of the patient's diseased tissue can be positioned, and the position of the surgical instrument can be tracked in real time.
[0079] For example, the process of determining the world coordinates of the target point P1 in the above Figure 5 will be described. The pixel coordinates of the left feature point P l1 on the left image obtained by the left camera are (u1, v1), the pixel coordinates of the right feature point P r1 on the right image obtained by the right camera are (u2, v2), and the focal lengths of the left camera and the right camera are both f. The world coordinates (x w , y w , z w ) of the target point P1 are determined according to the following formula.
[0080]
[0081]
[0082] wherein z1 and z2 are the distances from the left image to the origin of the left camera coordinate system and from the right image to the origin of the right camera coordinate system, A l and A r are the camera intrinsic parameters of the left camera and the right camera, R l , R r , t l and t r are the rotation vectors of the left camera and the right camera and the translation vectors of the left camera and the right camera.
[0083] For example, the process of determining the world coordinates of the target point P1 in the above Figure 5 will be described. The pixel coordinates of the left feature point P l1 on the left image obtained by the left camera are (u1, v1), the pixel coordinates of the feature point P d1 on the image obtained by the third camera are (u3, v3), and the focal lengths of the left camera and the third camera are both f. The world coordinates (x w , yw , z w ).
[0084]
[0085]
[0086] wherein z1, z3 are the distances from the left image to the origin of the left camera coordinate system, the distance from the third camera image to the origin of the third camera coordinate system, A l , A d are the camera intrinsic parameters of the left camera and the third camera, R l , R d , t l , t d are the rotation vector of the left camera, the rotation vector of the third camera, the translation vector of the left camera and the translation vector of the third camera, respectively.
[0087] Figure 6 is a structural schematic diagram of a device for determining a target point provided by an embodiment of the present application.
[0088] Exemplarily, as shown in FIG. 6, the device 600 includes: Figure 6 an acquisition module 601, configured to acquire N1 left feature points in a left image of a target picture captured by a left camera and M1 right feature points in a right image of the target picture captured by a right camera, the N1 left feature points and the M1 right feature points having a epipolar line constraint relationship, M1 and N1 being integers greater than 0;
[0089] a processing module 602, configured to perform stereo matching processing on the N1 left feature points and the M1 right feature points to obtain P spatial points, P being an integer greater than 0;
[0090] The acquisition module 601 is further configured to acquire T1 feature points of the target picture captured by a third camera, T1 being an integer greater than 0;
[0091] The processing module 602 is further configured to perform stereo matching processing on the N1 left feature points and the T1 feature points to obtain Q spatial points, Q being an integer greater than 0;
[0092] a determination module 603, configured to determine a target point according to the P spatial points and the Q spatial points.
[0093]
[0094] Optionally, the acquisition module 601 is further configured to acquire N left feature points in a left image of a target picture captured by the left camera, and M right feature points in a right image of the target picture captured by the right camera, the N left feature points including the N1 left feature points, and the M right feature points including M1 right feature points.
[0095] Optionally, the processing module 602 is further configured to perform stereo matching processing on N-N1 left feature points and M-M1 right feature points to obtain O spatial points, the N-N1 left feature points and the M-M1 right feature points not having epipolar constraint relationship, and O being an integer greater than 0.
[0096] Optionally, the determining module 603 is specifically configured to determine the target point according to the P spatial points, the Q spatial points, and the O spatial points.
[0097] Optionally, the determining module 603 is specifically further configured to determine, as R spatial points, spatial points having a distance less than or equal to a preset value from any one of the P spatial points and the Q spatial points, and R being an integer greater than 0; and determine the R spatial points and the O spatial points as the target point.
[0098] Optionally, the determining module 603 is further configured to calibrate the left camera and the right camera with the optical center of the left camera as the origin of a world coordinate system to determine camera parameters of the left camera and the right camera; and calibrate the left camera and the third camera with the optical center of the left camera as the origin of the world coordinate system to determine camera parameters of the third camera; or calibrate the left camera and the right camera with the optical center of the right camera as the origin of the world coordinate system to determine camera parameters of the left camera and the right camera; and calibrate the right camera and the third camera with the optical center of the right camera as the origin of the world coordinate system to determine camera parameters of the third camera.
[0099] Optionally, the determining module 603 is further configured to determine the world coordinates of the target point according to pixel coordinates of the target point on an image of the left camera, pixel coordinates of the target point on an image of the right camera, the camera parameters of the left camera, and the camera parameters of the right camera; or determine the world coordinates of the target point according to pixel coordinates of the target point on an image of the left camera, pixel coordinates of the target point on an image of the third camera, the camera parameters of the left camera, and the camera parameters of the third camera.
[0100] Optionally, the left camera and the right camera are both black-and-white cameras, and the third camera is a color camera.
[0101] Figure 7is a structural schematic diagram of a device for determining a target point provided in the embodiment of the present application.
[0102] As shown in the example, Figure 7 The device 700 for determining a target point comprises a memory 701, a processor 702, and a computer program 703 stored in the memory 701 and running on the processor 702, wherein the processor 702 executes the computer program 703, so that the device for determining a target point can perform any one of the methods for determining a target point described above.
[0103] The embodiment can divide the device for determining a target point into functional modules according to the above method examples, for example, each functional module can be provided, or two or more functions can be integrated into one processing module, and the integrated module can be implemented in the form of hardware. It should be noted that the division of the modules in the embodiment is illustrative, and is only a logical function division, and another division mode can be used in actual implementation.
[0104] In the case of dividing each functional module according to each function, the device for determining a target point can comprise an acquisition module, a processing module, and a determination module, etc. It should be noted that all related contents of each step involved in the above method embodiments can be referred to the function description of the corresponding functional module, and will not be repeated here.
[0105] The device for determining a target point provided in the embodiment is used to execute the above method for determining a target point, and thus can achieve the same effect as the above implementation method.
[0106] In the case of using an integrated unit, the device for determining a target point can comprise a processing module and a storage module. The processing module can be used to control and manage the actions of the device for determining a target point. The storage module can be used to support the device for determining a target point to execute program codes and data, etc.
[0107] The processing module can be a processor or a controller, which can realize or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure of the present application. The processor can also be a combination of computing functions, such as one or more microprocessor combinations, a combination of digital signal processing (DSP) and microprocessor, etc. The storage module can be a memory.
[0108] The device for determining a target point provided in the embodiments of the present application can be a chip, a component or a module. The device for determining a target point can include a processor and a memory connected to each other. The memory is configured to store instructions. When the device for determining a target point is running, the processor can invoke and execute the instructions to enable the chip to perform any one of the methods for determining a target point described above.
[0109] The embodiments of the present application provide a computer readable storage medium having instructions stored therein. When the instructions are run on a computer or a processor, the computer or the processor performs any one of the methods for determining a target point described above.
[0110] The embodiments of the present application also provide a computer program product having instructions. When the computer program product is run on a computer or a processor, the computer or the processor performs the related steps to implement any one of the methods for determining a target point described above.
[0111] The device for determining a target point, the computer readable storage medium, the computer program product having instructions or the chip provided in the embodiments of the present application are used to perform the corresponding methods provided above, and thus the beneficial effects achieved by the device for determining a target point, the computer readable storage medium, the computer program product having instructions or the chip can refer to the beneficial effects of the corresponding methods provided above, which will not be described here.
[0112] From the above description of the embodiments, those skilled in the art can understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.
[0113] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division, and there can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another device, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the units shown or discussed can be indirect coupling or communication connection through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0114] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method of determining a target point, characterized by, The method is applied to a visual positioning system, the visual positioning system comprising a binocular visual positioning device and a third camera, the binocular visual positioning device comprising a left camera and a right camera for capturing images, a plane formed by an optical center of the third camera, an optical center of the left camera and an optical center of the right camera being non-intersected with an effective field of view of the binocular visual positioning device, the method comprising: obtaining N1 left feature points in a left image of a target picture captured by the left camera and M1 right feature points in a right image of the target picture captured by the right camera, the N1 left feature points and the M1 right feature points having a epipolar constraint relationship, M1 and N1 being integers greater than 0; performing stereo matching processing on the N1 left feature points and the M1 right feature points to obtain P spatial points, P being an integer greater than 0; obtaining T1 feature points of the target picture captured by the third camera, T1 being an integer greater than 0; performing the stereo matching processing on the N1 left feature points and the T1 feature points to obtain Q spatial points, Q being an integer greater than 0; determining a target point according to the P spatial points and the Q spatial points.
2. The method of claim 1, wherein, The method further comprises: obtaining N left feature points in a left image of a target picture captured by the left camera and M right feature points in a right image of the target picture captured by the right camera, the N left feature points comprising the N1 left feature points, the M right feature points comprising the M1 right feature points, N and M being integers greater than 0; performing stereo matching processing on N-N1 left feature points and M-M1 right feature points to obtain O spatial points, the N-N1 left feature points and the M-M1 right feature points not having an epipolar constraint relationship, O being an integer greater than 0; and the determining the target point according to the P spatial points and the Q spatial points comprises: determining the target point according to the P spatial points, the Q spatial points and the O spatial points.
3. The method of claim 2, wherein, The determining the target point according to the P spatial points, the Q spatial points and the O spatial points comprises: determining, as R spatial points, spatial points having a distance less than or equal to a preset value from any one of the P spatial points and the Q spatial points, R being an integer greater than 0; determining the R spatial points and the O spatial points as the target point.
4. The method of claim 1, wherein, The method further comprises: calibrating the left camera and the right camera with an optical center of the left camera as an origin of a world coordinate system to determine camera parameters of the left camera and the right camera; and calibrating the left camera and the third camera with the optical center of the left camera as the origin of the world coordinate system to determine camera parameters of the third camera; or, calibrating the left camera and the right camera with an optical center of the right camera as an origin of a world coordinate system to determine camera parameters of the left camera and the right camera; And, taking the optical center of the right camera as the origin of a world coordinate system, calibrating the right camera and the third camera to determine camera parameters of the third camera.
5. The method of claim 4, wherein, The method further includes: determining the world coordinates of the target point according to the pixel coordinates of the target point on the image in the left camera, the pixel coordinates of the target point on the image in the right camera, the camera parameters of the left camera and the camera parameters of the right camera; Or, determining the world coordinates of the target point according to the pixel coordinates of the target point on the image in the left camera, the pixel coordinates of the target point on the image in the third camera, the camera parameters of the left camera and the camera parameters of the third camera.
6. The method according to claim 4 or 5, characterized in that, The left camera and the right camera are both black-and-white cameras, and the third camera is a color camera.
7. An apparatus for determining a target point, the apparatus comprising: The device includes: an acquisition module configured to acquire N1 left feature points in a left image of a target picture captured by a left camera and M1 right feature points in a right image of the target picture captured by a right camera, the N1 left feature points and the M1 right feature points having a epipolar constraint relationship, M1 and N1 being integers greater than 0; a processing module configured to perform stereo matching processing on the N1 left feature points and the M1 right feature points to obtain P spatial points, P being an integer greater than 0; the acquisition module is further configured to acquire T1 feature points of the target picture captured by a third camera, T1 being an integer greater than 0; the processing module is further configured to perform the stereo matching processing on the N1 left feature points and the T1 feature points to obtain Q spatial points, Q being an integer greater than 0; a determination module configured to determine a target point according to the P spatial points and the Q spatial points.
8. A device for determining a target point, characterized in that, The computer readable storage medium stores instructions, when the instructions are run on a computer or a processor, causing the computer or the processor to execute the method for determining a target point according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, when the instructions are run on a computer or a processor, causing the computer or the processor to execute the method for determining a target point according to any one of claims 1 to 6.
10. A computer program product comprising instructions, characterized in that, The computer program product, when run on a computer or a processor, causes the computer or the processor to execute the method for determining a target point according to any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual reality and vision-based positioning system
CN106251357A
Dual-target positioning method for simulated medical instrument and virtual simulation medical teaching system
CN108830905A