Camera pose determination method and monocular camera vision system
By using the absolute distance constraint value of the double-sided target and the PnP algorithm to correct the camera pose in the monocular camera vision SLAM, the accumulated pose error and position drift problems in the monocular camera vision SLAM are solved, and high-precision positioning and mapping are achieved.
Patent Information
- Application Number
- CN202410040854.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-10
- Publication Date
- 2025-07-11
AI Technical Summary
In monocular camera visual SLAM technology, the existing loopback constraint method is difficult to correct the accumulated pose error and position drift in time, affecting the accuracy of visual SLAM.
Using absolute distance constraint values based on the double-sided target, by updating the world coordinates and camera poses corresponding to the image coordinates in the image frame, the initial camera pose is corrected using the PnP algorithm to provide accurate scale information to improve positioning and map construction accuracy.
It effectively solves the trajectory deviation and map drift problems in monocular camera visual SLAM, improves the accuracy and efficiency of visual SLAM without additional sensors, is low cost and strong robustness.
Smart Images

Figure CN120298486A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a method for determining the pose of a camera and a monocular camera vision system. Background Art
[0002] The Simultaneous Localization and Mapping (SLAM) technology based on visual information is widely used in fields such as robotics, virtual reality, augmented reality, and unmanned driving. Its applications include the localization and mapping of sensors themselves, as well as subsequent path planning and scene understanding. Among them, the visual SLAM technology based on a monocular camera can locate the monocular camera according to the camera pose and obtain the coordinate transformation relationship between the image coordinate system and the world coordinate system, and then convert the image coordinates in the image frame collected by the monocular camera to the world coordinate system to complete the mapping process.
[0003] However, it is difficult for a monocular camera to obtain direct distance information (or depth information) relative to the environment from the collected two-dimensional images. It is necessary to estimate its own pose change through two or more frames of images, and then calculate the current position by accumulating the pose change. Therefore, cumulative pose errors are likely to occur during the movement of the camera, which in turn reduces the accuracy of visual SLAM. Currently, targets such as QR codes, checkerboards, and AprilTags with scale information and feature point information (such as corner point information) are usually used as scale markers and visual markers to assist in positioning to reduce pose errors. However, problems such as incorrect corner point matching or position drift are still likely to occur during the process of detecting corner points and obtaining scale information, affecting the accuracy of visual SLAM. Existing technologies usually use loop constraints to correct the cumulative errors and position drift generated during the movement of the monocular camera. By extracting repeated or similar features between all image frames for comparison, a closed loop is formed to optimize the camera trajectory and map. However, this common loop constraint method is difficult to obtain a high-precision camera pose, and cannot correct the drift of the camera trajectory and the cumulative errors during the mapping process in a timely manner, thereby affecting the accuracy of monocular camera visual SLAM. Summary of the Invention
[0004] To improve the accuracy of monocular camera visual SLAM, the present invention provides a method for determining the pose of a camera and a monocular camera vision system.
[0005] In a first aspect, an embodiment of the present application provides a method for determining the pose of a camera, including:
[0006] Based on a monocular camera to collect an image sequence, the image sequence includes a plurality of image frames, wherein one image frame includes image information of at least one double-sided target, the target surface of the double-sided target includes a first surface and a second surface, the first surface includes a plurality of labels, the second surface includes a plurality of labels, and the label includes a plurality of corner points;
[0007] Based on the image frame, obtain the image coordinates of each corner point and the world coordinates corresponding to the image coordinates;
[0008] Based on the image coordinates and the world coordinates corresponding to the image coordinates, obtain the initial camera pose, where the initial camera pose corresponds to an image frame;
[0009] Update the initial camera pose according to the absolute distance constraint value, where the absolute distance constraint value is set based on the distance between corner points.
[0010] In some embodiments, the absolute distance constraint value is the distance between the corner points of the labels on the same first surface or the distance between the corner points of the labels on the same second surface.
[0011] In some embodiments, updating the initial camera pose according to the absolute distance constraint value includes:
[0012] According to the absolute distance constraint value, update the world coordinates corresponding to the image coordinates in the image frame to obtain the updated world coordinates corresponding to the image coordinates;
[0013] Update the initial camera pose according to the image coordinates and the updated world coordinates corresponding to the image coordinates.
[0014] In some embodiments, the absolute distance constraint value is the distance between the corner points of the label on the first surface and the corner points of the label on the second surface.
[0015] In some embodiments, updating the initial camera pose according to the absolute distance constraint value includes:
[0016] According to the image sequence, obtain the first image frame and the second image frame collected in chronological order, where the first image frame contains the image information of the label of one target surface, and the second image frame contains the image information of the label of another target surface;
[0017] According to the absolute distance constraint value, update the world coordinates corresponding to the image coordinates in the second image frame to obtain the updated world coordinates corresponding to the image coordinates in the second image frame;
[0018] Update the initial camera pose according to the image coordinates in the second image frame and the updated world coordinates corresponding to the image coordinates in the second image frame.
[0019] Based on the accurate pose of the monocular camera and the updated world coordinates obtained by the above method, the problems of trajectory deviation and map drift in positioning and mapping can be effectively solved, and the accuracy of monocular camera visual SLAM can be improved.
[0020] In some embodiments, the double-sided target is one or a combination of a rectangular plate, a cube, or a frustum. Double-sided targets of different sizes can be used as visual markers according to actual test requirements. Labels can also be made on multiple faces of the cube or frustum to adapt to different SLAM application scenarios.
[0021] In some embodiments, the double-sided target is a combination of one or more materials such as paper, foam, plastic, or metal. The double-sided target has a simple structure and low cost, and labels can be made on the target surface of the target through common processes such as printing or laser processing, and the processing accuracy is controllable.
[0022] In a second aspect, an embodiment of the present application further provides a monocular camera vision system, which includes:
[0023] A double-sided target, with at least one double-sided target. The target surface of the double-sided target includes a first surface and a second surface. The first surface includes multiple labels, and the second surface includes multiple labels. The labels include multiple corner points;
[0024] A monocular camera for collecting an image sequence. The image sequence includes multiple image frames. Among them, one image frame includes image information of at least one double-sided target;
[0025] A processing unit for obtaining the image coordinates of each corner point and the world coordinates corresponding to the image coordinates according to the image frame;
[0026] The processing unit is further configured to obtain an initial camera pose according to the image coordinates and the world coordinates corresponding to the image coordinates, where the initial camera pose corresponds to one image frame;
[0027] The processing unit is further configured to update the initial camera pose according to the absolute distance constraint value, where the absolute distance constraint value is set based on the distance between corner points.
[0028] The present application discloses a method for determining the pose of a camera and a monocular camera vision system. The method sets an absolute distance constraint value based on a double-sided target, providing accurate scale information for the pose correction of a monocular camera. It can effectively solve the problems of trajectory drift and map drift in positioning and mapping, and improve the accuracy of visual SLAM. The method performs pose correction based on a double-sided target with a simple structure, without the need for additional sensors, and has the characteristics of low cost, high accuracy, and strong robustness, and can be applied to different monocular camera visual SLAM scenarios. Description of the Drawings
[0029] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0030] Figure 1 is a schematic flowchart of a method for determining the pose of a camera in an embodiment of the present application;
[0031] Figure 2 is a schematic diagram of the target surface of a double-sided target in an embodiment of the present application;
[0032] Figure 3 is a schematic diagram of the positional relationship between the moving trajectory of a camera and a double-sided target in an embodiment of the present application;
[0033] Figure 4 is a schematic diagram of the positional relationship between the moving trajectory of a camera and a double-sided target in an embodiment of the present application;
[0034] Figure 5 is a schematic diagram of the positional relationship between a monocular camera and a double-sided target in an embodiment of the present application;
[0035] Figure 6 is a schematic diagram of the positional relationship between the moving trajectory of a camera and two double-sided targets in an embodiment of the present application;
[0036] Figure 7 is a schematic diagram of the positional relationship between the moving trajectory of a camera and two double-sided targets in an embodiment of the present application. Detailed implementation manners
[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe in detail the embodiments of the present application in conjunction with the drawings. When the following description involves the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are merely examples of methods and devices consistent with some aspects of the present application as detailed in the appended claims.
[0038] Existing monocular camera visual SLAM technology usually corrects the cumulative pose error and map drift based on the loop closure constraint method. However, ordinary loop closure constraint methods require the monocular camera to pass through an observed scene or target and retrieve and match all image frames in the image sequence, making it difficult to obtain the accurate camera pose in a timely manner. Moreover, the above loop closure constraint method is relatively sensitive to changes in the surrounding environment, making the detection process still prone to feature mismatch, trajectory drift, and map drift problems.
[0039] The embodiment of the present application discloses a method for determining the pose of a camera. This method sets an absolute distance constraint value based on a double-sided target, and can update the world coordinates corresponding to the image coordinates of the corner points and the camera pose under the image frame in a timely manner. As Figure 1 shown, the method includes the following steps:
[0040] S101. Acquire an image sequence based on a monocular camera.
[0041] In one embodiment, at least one double-sided target is included in the moving field of view of the monocular camera. In one example, the image sequence includes multiple image frames, and one image frame includes the image information of at least one double-sided target. The target surface of the double-sided target includes a first surface and a second surface. The first surface includes multiple labels, and the second surface includes multiple labels. The positions of the labels on each target surface are fixed. The multiple labels on the first surface do not overlap, and the multiple labels on the second surface do not overlap. The label includes multiple corner points. In one example, as Figure 2 shown, the double-sided target is a rectangular target board, and the target surface includes a first surface and a second surface. The first surface shown in the figure includes at least two non-overlapping AprilTags. The second surface includes at least two non-overlapping AprilTags, and the second surface is not shown in the figure. In one example, Figure 2 the scale information of the double-sided target shown is known. The scale information includes the size of the double-sided target, the size of the AprilTag on the target surface, the distance between the AprilTags, and the orientation relationship between the AprilTags, etc. According to the scale information of the double-sided target, the distance between different corner points and the relative position relationship between the corner points can be obtained.
[0042] In one embodiment, before step S102, it further includes: establishing a world coordinate system with any point on the double-sided target as the origin, and obtaining the world coordinates of each corner point according to the scale information of the double-sided target. In one example, as Figure 2 shown, when the label is an AprilTag, each AprilTag on the target surface of the double-sided target also has a unique coding information. The corresponding relationship between the coding information of the AprilTag and the world coordinates of the corner points on the AprilTag can be recorded in the processing unit.
[0043] S102. Obtain the image coordinates of each corner point and the world coordinates corresponding to the image coordinates according to the image frame.
[0044] In one embodiment, for the image frames in the image sequence collected by a monocular camera, an image coordinate system is established with the center of the image frame as the origin. The processing unit detects and extracts each corner point in the image frame using common feature point descriptor algorithms such as Scale-Invariant Feature Transform (SIFT), Orient Fast and Rotated Brief (ORB), or Speeded-Up Robust Features (SURF) according to the image frames in the image sequence, and obtains the image coordinates of each corner point and the descriptor of each corner point. In one example, the descriptor of a corner point includes one or a combination of information such as the gray values of the pixels around the corner point in the image frame, the gradient amplitude, or the change in gray gradient, and is used to characterize the distribution position of each corner point in the image coordinate system and the texture information around the corner point. The processing unit matches each corner point in the image coordinate system with each corresponding corner point in the world coordinate system according to the descriptor of each corner point and the scale information of the double-sided target, and obtains the world coordinates corresponding to the image coordinates of each corner point. In one example, when the tag is AprilTag, the processing unit can also obtain the world coordinates corresponding to the image coordinates of each corner point according to the correspondence between the encoding information of AprilTag and the world coordinates of the corner points on AprilTag.
[0045] S103. Obtain the initial camera pose according to the image coordinates and the world coordinates corresponding to the image coordinates.
[0046] In one embodiment, the processing unit uses the Perspective-n-Point (PnP) algorithm to obtain the initial camera pose under the image frame according to the image coordinates of each corner point in the image frame and the world coordinates corresponding to the image coordinates. Among them, the PnP method can be one or a combination of Direct Linear Transform (DLT), Perspective-3-Point (P3P) based on three pairs of feature points, or Efficient Perspective-n-Point (EPnP).
[0047] S104. Update the initial camera pose according to the absolute distance constraint value, where the absolute distance constraint value is set based on the distance between corner points.
[0048] To prevent problems such as incorrect corner matching or cumulative errors from affecting the accuracy of monocular camera visual SLAM, it is also necessary to correct the initial camera pose obtained in step S103.
[0049] In one embodiment, the absolute distance constraint value includes a first type of absolute distance constraint value, and the first type of absolute distance constraint value is the distance between the corner points of the labels on the same first surface or the distance between the corner points of the labels on the same second surface. In an example, as Figure 3 shown, the monocular camera moves on one side of the double-sided target, and the image frames in the image sequence contain the image information of the labels on the first surface. The first surface includes a first label and a second label. The first label includes four corner points, and the second label includes four corner points. The first type of absolute distance constraint value is the distance between the corner points on the first label and the corner points on the second label. The processing unit updates the world coordinates corresponding to the image coordinates of each corner point on the first label and the world coordinates corresponding to the image coordinates of each corner point on the second label in the image frame according to the first type of absolute distance constraint value, and obtains the updated world coordinates corresponding to the image coordinates. The calculated value of the distance between each corner point obtained according to the updated world coordinates is the same as the corresponding first type of absolute distance constraint value. The processing unit updates the initial camera pose based on the PnP algorithm according to the image coordinates of each corner point and the updated world coordinates corresponding to the image coordinates. During the movement of the monocular camera on one side of the double-sided target, the initial camera pose of the current image frame can be updated in a timely manner through the first type of absolute distance constraint value, without having to retrieve and match all the image frames in the image sequence, which can improve the efficiency of monocular camera visual SLAM and effectively solve the problems of cumulative error and position drift. The updated initial camera pose and the updated world coordinates corresponding to the image coordinates can also be used to update the camera movement trajectory and optimize the map, improving the accuracy of monocular camera positioning and mapping.
[0050] In one embodiment, the absolute distance constraint value further includes a second type of absolute distance constraint value, and the second type of absolute distance constraint value is the distance between the corner points of the label on the first surface and the corner points of the label on the second surface. In an example, as Figure 4As shown, the monocular camera moves from one side of the double-sided target to the other side. The image sequence collected by the monocular camera during the above movement includes a first image frame and a second image frame collected in chronological order. In one example, the first image frame contains the image information of the label on the first side, and the second image frame contains the image information of the label on the second side. In another example, the first image frame contains the image information of the label on the second side, and the second image frame contains the image information of the label on the first side. In one example, according to the second type of absolute distance constraint value, the world coordinates corresponding to the image coordinates of each corner point in the second image frame are updated to obtain the updated world coordinates corresponding to the image coordinates. According to the image coordinates in the second image frame and the updated world coordinates corresponding to the image coordinates, the initial camera pose under the second image frame is updated based on the PnP algorithm. In one embodiment, before updating the initial camera pose under the second image frame according to the second type of absolute distance constraint value, it further includes: updating the initial camera pose under the first image frame according to the first type of absolute distance constraint value. By updating the initial camera poses under different image frames through the absolute distance constraint value, the error can be corrected in a timely manner, and the accuracy of the monocular camera visual SLAM can be improved.
[0051] In one embodiment, there is actually an image frame in the image sequence collected by the monocular camera that does not contain the image information of the label on any target surface of the double-sided target. As Figure 5 shown, when the monocular camera moves to the position shown in the figure, the target surface cannot be observed within the field of view of the monocular camera, and the corresponding image frame does not contain the image information of the label on any target surface. The current image frame is recorded as a blind area image frame. At this time, it is necessary to combine the image information of other image frames in the image sequence to obtain the camera pose under the blind area image frame. In one example, according to the image sequence, the first image frame, the blind area image frame, and the second image frame collected in chronological order are obtained. Among them, the first image frame contains the image information of the label on one target surface, and the second image frame contains the image information of the label on the other target surface. In one example, the initial camera pose under the first image frame is updated according to the first type of absolute distance constraint value, and the initial camera pose under the second image frame is updated according to the second type of absolute distance constraint value. In another example, the initial camera pose under the first image frame is updated according to the first type of absolute distance constraint value, and the initial camera pose under the second image frame is updated according to the first type of absolute distance constraint value. The processing unit updates the initial camera pose under the blind area image frame between the first image frame and the second image frame according to the updated initial camera pose under the first image frame and the updated initial camera pose under the second image frame. The pose determination method in the above embodiment can effectively solve the problem that when the double-sided target is in the field of view blind area of the monocular camera, the camera pose cannot be obtained in a timely manner or the camera pose error is large.
[0052] In one embodiment, the moving field of view of the monocular camera includes at least two double-sided targets. In one example, asFigure 6 As shown, the image frame captured by the monocular camera includes the image information of the first target and the second target. Both the first target and the second target are double-sided targets. A world coordinate system is established with any point on one of the double-sided targets as the origin. When the monocular camera moves to different positions, the moving field of view of the monocular camera can observe the target surface of at least one double-sided target, and the image frame captured by the monocular camera contains the image information of the labels on at least one target surface. In one example, the processing unit can update the initial camera pose under the image frame according to the first type of absolute distance constraint value.
[0053] In one embodiment, the first target and the second target are placed in parallel, and the monocular camera only moves on one side of the two targets. At this time, only labels need to be made on the target surfaces facing the monocular camera on the first target and the second target. In one example, as Figure 7 shown, both the first target and the second target are single-sided targets, and there are at least two labels on the target surfaces of the first target and the second target facing the position of the monocular camera. The first type of absolute distance constraint value is the distance between the corner points of the labels on the target surface of the first target, and the first type of absolute distance constraint value can also be set as the distance between the corner points of the labels on the target surface of the second target. The second type of absolute distance constraint value can be set as the distance between the corner points of the labels on the target surface of the first target and the corner points of the labels on the target surface of the second target. The image frame captured by the monocular camera contains the image information of the target surface of the first target and the image information of the target surface of the second target. In one example, the processing unit can update the initial camera pose under the image frame according to the first type of absolute distance constraint value. In another example, the processing unit can update the initial camera pose under the image frame according to the second type of absolute distance constraint value.
[0054] According to the positional relationship between the moving trajectory of the monocular camera and the targets, the number, position, and label arrangement of the targets can be flexibly changed. This method for determining the camera pose has the characteristics of low cost, high precision, and strong robustness, and can be flexibly applied to different visual SLAM scenarios such as map mapping and autonomous driving.
[0055] The above has elaborated in detail a method for determining the pose of a camera in an embodiment of the present application. Next, a monocular camera vision system in an embodiment of the present application is provided. The monocular camera vision system includes a double-sided target, a monocular camera, and a processing unit.
[0056] Double-sided target, there is at least one double-sided target, the target surface of the double-sided target includes a first surface and a second surface, the first surface includes a plurality of labels, the second surface includes a plurality of labels, and the labels include a plurality of corner points;
[0057] Monocular camera, the monocular camera is used to capture an image sequence, the image sequence includes a plurality of image frames, wherein, one image frame includes the image information of at least one double-sided target;
[0058] A processing unit, which is configured to obtain the image coordinates of each corner point and the world coordinates corresponding to the image coordinates according to an image frame;
[0059] The processing unit is further configured to obtain an initial camera pose according to the image coordinates and the world coordinates corresponding to the image coordinates, wherein the initial camera pose corresponds to an image frame;
[0060] The processing unit is further configured to update the initial camera pose according to an absolute distance constraint value, wherein the absolute distance constraint value is set based on the distances between corner points.
[0061] In some embodiments, the double-sided target is any one or a combination of a rectangular plate, a cube, or a frustum. In one example, the target is a cube target plate, and the target surfaces are two adjacent surfaces on the cube target plate, and each target surface includes at least two labels. In one example, the target is a cube target plate, and all surfaces of the cube target plate are target surfaces, and each target surface includes at least two labels. In one example, the target is a frustum, and the frustum includes at least three target surfaces, and each target surface includes at least two labels. Double-sided targets of different sizes can be used as visual markers according to actual test requirements. Labels can also be made on multiple surfaces of a cube or a frustum to adapt to different SLAM application scenarios.
[0062] In some embodiments, the material of the target is one or a combination of paper, foam, plastic, or metal. In some embodiments, the labels on the target surface are made by common processing techniques such as printing, printing, or laser processing. The labels can be one or a combination of visual labels such as QR codes, checkerboards, AprilTags, or ArUcoTags. In one embodiment, the scale information of the target can be adjusted by controlling the processing parameters. The scale information of the target includes the target size, the size of the labels on the target surface, the distance between the labels, and the orientation relationship between the labels, etc. In one example, the first type of absolute distance constraint value is set according to the distances between the corner points of different labels on the same target surface. The second type of absolute distance constraint value is set according to the distances between the corner points of the labels on different target surfaces. In one example, an orientation constraint angle can also be set based on the orientation relationship between multiple labels to perform orientation constraints on the world coordinates corresponding to the image coordinates of the corner points. In some embodiments, after obtaining the world coordinates corresponding to the image coordinates of the corner points according to the image coordinates of the corner points, coplanarity constraints of the corner points and collinearity constraints of multiple corner points can also be set to update the world coordinates of the corner points and improve the accuracy of monocular camera visual SLAM.
[0063] In some embodiments, the double-sided target can also be extended and applied to the tests of other vision sensors, such as the parameter calibration and tests of sensors like lidar, millimeter-wave radar, and depth camera. The parameters include one or a combination of more than one of ranging accuracy, resolution angle, target detection rate, or detection distance. In one example, the surface of the double-sided target includes a checkerboard pattern, which can be used as a scale marker for testing the ranging accuracy of lidar. In one example, the target surface of the double-sided target is coated with a reflective film or a reflective coating, and the reflectivities of different target surfaces are different. The double-sided target can be used as a vision marker for testing the detection rate of lidar for targets with different reflectivities.
[0064] The processing unit can be a field-programmable gate array (FPGA), a system on chip (SoC), a central processing unit (CPU), a network processor (NP), a digital signal processing circuit, a microcontroller unit (MCU), a programmable logic device (PLD), an application-specific integrated circuit (ASIC), or any combination thereof for implementing related functions. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or other integrated chips.
[0065] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on some computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0066] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a means for implementing the functions specified in one or more of the following Figure 1 flows or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the following blocks.
[0067] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction means, and the instruction means implements the functions specified in one or more of the following Figure 1 flows or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the following blocks.
[0068] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the following Figure 1 flows or multiple flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more of the following blocks.
[0069] In the description of the present application, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, "at least two" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0070] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" used herein includes any and all combinations of some related listed items. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances. The above are only the preferred embodiments of this application and are not used to limit this application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.
Claims
1. A method for determining the pose of a camera, characterized in that Comprising: Based on a monocular camera to collect an image sequence, the image sequence includes a plurality of image frames, wherein one of the image frames includes image information of at least one double-sided target, the target surface of the double-sided target includes a first surface and a second surface, the first surface includes a plurality of labels, the second surface includes a plurality of the labels, and the label includes a plurality of corner points; According to the image frame, obtain the image coordinates of each of the corner points and the world coordinates corresponding to the image coordinates; According to the image coordinates and the world coordinates corresponding to the image coordinates, obtain an initial camera pose, wherein the initial camera pose corresponds to one of the image frames; Update the initial camera pose according to an absolute distance constraint value, wherein the absolute distance constraint value is set based on the distance between the corner points.
2. The method according to claim 1, characterized in that, The absolute distance constraint value is the distance between the corner points of the labels that are on the same first surface or the distance between the corner points of the labels that are on the same second surface.
3. The method according to claim 1 or 2, characterized in that, The updating the initial camera pose according to the absolute distance constraint value includes: According to the absolute distance constraint value, update the world coordinates corresponding to the image coordinates in the image frame to obtain the updated world coordinates corresponding to the image coordinates; According to the image coordinates and the updated world coordinates corresponding to the image coordinates, update the initial camera pose.
4. The method according to claim 1, characterized in that, The absolute distance constraint value is the distance between the corner points of the label on the first surface and the corner points of the label on the second surface.
5. The method according to claim 1 or 4, characterized in that, The updating the initial camera pose according to the absolute distance constraint value includes: According to the image sequence, obtain a first image frame and a second image frame collected in chronological order, wherein the first image frame contains the image information of the labels of one of the target surfaces, and the second image frame contains the image information of the labels of the other target surface; According to the absolute distance constraint value, update the world coordinates corresponding to the image coordinates in the second image frame to obtain the updated world coordinates corresponding to the image coordinates in the second image frame; According to the image coordinates in the second image frame and the updated world coordinates corresponding to the image coordinates in the second image frame, update the initial camera pose.
6. The method according to claim 1, wherein The double-sided target is one or a combination of a rectangular plate, a cube or a frustum of a cone.
7. The method according to claim 1, characterized in that, The double-sided target is a combination of one or more materials such as paper, foam, plastic or metal.
8. A monocular camera vision system, characterized in that, Comprising: A double-sided target, there is at least one double-sided target, the target surface of the double-sided target includes a first surface and a second surface, the first surface includes a plurality of labels, the second surface includes a plurality of the labels, and the label includes a plurality of corner points; A monocular camera, the monocular camera is used to collect an image sequence, the image sequence includes a plurality of image frames, wherein one of the image frames includes the image information of at least one of the double-sided targets; A processing unit, the processing unit is used to obtain the image coordinates of each of the corner points and the world coordinates corresponding to the image coordinates according to the image frame; The processing unit is further configured to obtain an initial camera pose according to the image coordinates and the world coordinates corresponding to the image coordinates, wherein the initial camera pose corresponds to one of the image frames; The processing unit is further configured to update the initial camera pose according to an absolute distance constraint value, wherein, the absolute distance constraint value is set based on the distances between the corner points.