Camera pose calculation device, calculation method, and information processing program

The computing device calculates camera pose by identifying lines in a single image, addressing the inefficiencies of existing methods by minimizing deviations, thus simplifying and speeding up the process.

JP7837193B2Active Publication Date: 2026-03-30KK TOYOTA CHUO KENKYUSHO +1
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-25
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing methods for determining camera posture require special image acquisition and separate measurement of feature points, which are time-consuming and labor-intensive.

Method used

A computing device that identifies parallel and angled lines in a single image to calculate camera pose by minimizing the sum of squared deviations, reducing the need for preliminary image acquisition and feature point measurement.

Benefits of technology

Enables efficient determination of camera orientation and position from a single image, significantly reducing preparation time and effort.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837193000004
    Figure 0007837193000004
  • Figure 0007837193000005
    Figure 0007837193000005
  • Figure 0007837193000006
    Figure 0007837193000006
Patent Text Reader

Abstract

To provide a technology that is able to calculate a posture of a camera from a photographed image of a single viewpoint of the camera.SOLUTION: A computing device that computes a posture of a camera includes a straight-line specifying unit that specifies a pair of parallel first straight line and second straight line on a photographed image captured using the camera and one third straight line having a known angle with respect to the first straight line and the second straight line. The computing device includes a variable generation unit that generates a plurality of variables representing positions of both ends of the first straight line, positions of both ends of the second straight line, positions of both ends of the third straight line, and a posture and a position of the camera. The computing device includes a computing unit that computes solutions of the plurality of variables so as to minimize a sum of squares of deviations between projection positions at which both the ends of the first straight line, both the ends of the second straight line, and both the ends of the third straight line are projected on a screen of the camera and positions on the photographed images at both the ends of the first straight line, both the ends of the second straight line, and both the ends of the third straight line. The computing device includes an acquisition unit that acquires a posture of the camera from the solutions of the plurality of variables.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed in this specification relates to an arithmetic device and an arithmetic method capable of calculating the posture of a camera, etc.

Background Art

[0002] When performing various controls using a camera, it may be necessary to obtain the three-axis posture of the camera. Conventionally, a technique for obtaining the posture of a camera by observing three or more feature points with known positions is known. Also, a technique for simultaneously obtaining the focal length and the position and posture of a camera by observing four or more feature points arranged to spread sufficiently in the depth direction is known. Incidentally, related techniques are disclosed in Patent Documents 1 and 2.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the prior art, it is necessary to specially acquire an image for estimating the camera posture or separately measure the positions of feature points in the image at the site where the camera is installed. This is a problem because it takes time and effort for preliminary preparation.

Means for Solving the Problems

[0005] One embodiment of the computing device disclosed herein is a computing device for calculating the pose of a camera. The computing device includes a line identification unit that identifies a pair of parallel first and second lines on an image captured using a camera, and a third line having a known angle with respect to the first and second lines. The computing device includes a variable generation unit that generates a plurality of variables representing the positions of both ends of the first line, the positions of both ends of the second line, the positions of both ends of the third line, the pose of the camera, and the position. The computing device includes a calculation unit that calculates a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the second line, and the third line onto the camera's screen, and the positions of the ends of the first line, the second line, and the third line on the captured image is minimized. The computing device includes an acquisition unit that obtains the pose of the camera from the solution for the plurality of variables.

[0006] The above configuration generates multiple variables representing the positions of both ends of the first line, the positions of both ends of the second line, the positions of both ends of the third line, and the camera's orientation and position. By calculating the solutions for these multiple variables, the camera's orientation can be obtained. The camera's orientation can be easily determined from only a single viewpoint image of the camera. Since there is no need to specially acquire images for camera orientation estimation or to separately measure the positions of feature points in the images, the amount of preparation required is reduced.

[0007] The first, second, and third lines may be lines that exist in the same plane. Details of the effects will be explained in the examples.

[0008] The first and second lines may be lines existing in the first plane. The third line may be a line existing in the second plane that intersects the first plane at a known angle. Details of the effects will be explained in the examples.

[0009] One of the first and second planes may be the ground, and the other may be a plane of an object placed on the ground. Details of the effects will be explained in the examples.

[0010] The known angle may be a right angle.

[0011] The calculation unit may further include a measuring unit that measures the actual distance between the first and second lines, or the height of the camera lens center from the ground. The acquisition unit may acquire the scales of the endpoints of the first line, the endpoints of the second line, the endpoints of the third line, and the camera position from among the solutions of multiple variables, based on the measurement results of the measuring unit. Details of the effects will be explained in the examples.

[0012] The variable generation unit may generate three variables representing the endpoints of the first line, three variables representing the endpoints of the second line, and three variables representing the endpoints of the third line by defining an XY plane in which the first and second lines exist, and in which the direction of the first and second lines is either the X-axis or the Y-axis. Details of the effects will be explained in the examples.

[0013] The variable generation unit may generate three variables representing the camera's three-dimensional attitude angles. Alternatively, the variable generation unit may generate one variable representing the camera's position by setting the camera's position to the origin of the XY plane. Details of the effects will be explained in the examples.

[0014] One embodiment of the calculation method disclosed herein is a method for calculating the pose of a camera. The calculation method comprises the step of identifying a pair of parallel first and second lines on an image captured using a camera, and a third line having a known angle with respect to the first and second lines. The calculation method comprises the step of generating a plurality of variables representing the positions of both ends of the first line, the positions of both ends of the second line, the positions of both ends of the third line, the pose of the camera, and the position. The calculation method comprises the step of calculating a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the ends of the second line, and the ends of the third line onto the camera's screen, and the positions of the ends of the first line, the ends of the second line, and the ends of the third line on the captured image is minimized. The calculation method comprises the step of obtaining the pose of the camera from the solution for the plurality of variables. Details of the effects will be described in the examples.

[0015] One embodiment of the information processing program disclosed herein is an information processing program that is loaded into the computer of the information processing device and calculates the camera's orientation. The information processing program causes the information processing device to function as a line identification means for identifying a pair of parallel first and second lines on an image captured using the camera, and a third line having a known angle with respect to the first and second lines. The information processing program causes the information processing device to function as a variable generation means for generating a plurality of variables representing the positions of both ends of the first line, the positions of both ends of the second line, the positions of both ends of the third line, the camera's orientation, and its position. The information processing program causes the information processing device to function as a calculation means for calculating a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the second line, and the third line onto the camera's screen and the positions of the ends of the first line, the second line, and the third line on the captured image is minimized. The information processing program causes the information processing device to function as an acquisition means for obtaining the camera's orientation from the solution for the plurality of variables. Details of the effects will be described in the examples. [Brief explanation of the drawing]

[0016] [Figure 1] This is a schematic perspective view of forklift 1. [Figure 2] This is a block diagram of forklift 1. [Figure 3] This is a flowchart that explains the specific details of the calculation process. [Figure 4] This is an example of an image captured by camera 30 in Example 1. [Figure 5] This figure shows the calculation results in Example 1. [Figure 6] This is an example of an image captured by camera 30 in Example 2. [Figure 7] This figure shows the calculation results in Example 2. [Figure 8] This is an example of an image taken when calculating the orientation and position of a truck's cargo bed. [Figure 9] This is an example of an image taken when performing camera calibration. [Modes for carrying out the invention] [Examples]

[0017] (Configuration of Forklift 1) The forklift 1 of this embodiment will be described below with reference to Figures 1 and 2. Figure 1 is a schematic perspective view of the forklift 1, and Figure 2 is a block diagram. The forklift 1 is an unmanned forklift. As shown in Figure 1, the forklift 1 comprises a body 2, a control unit 10 (shown in Figure 2), a mast 20, forks 22, a backrest 23, a camera mounting unit 24, and a camera 30.

[0018] The vehicle body 2 is equipped with front wheels 28 and rear wheels 29 on each of its sides. The front wheels 28 are connected to a drive wheel motor (not shown) via a drive mechanism and are rotationally driven by a movement control unit 17 (shown in Figure 2). The rear wheels 29 are connected to a steering device (not shown) and the direction of the wheels is adjusted by the movement control unit 17.

[0019] The mast 20 is a support column attached to the front of the vehicle body 2, and its axis extends vertically. The fork 22 is attached to the mast 20 so as to be movable vertically. The fork 22 has a pair of claws 22a and 22b. The camera fixing part 24 is a member that fixes the camera 30.

[0020] The block diagram in Figure 2 is explained below. The forklift 1 comprises a control unit 10, forks 22, front wheels 28, rear wheels 29, a camera 30, and an encoder 40. The control unit 10 is a device that controls the forklift 1. The control unit 10 can be configured, for example, by a computer equipped with a CPU, ROM, RAM, etc. By executing a program, the computer allows the control unit 10 to function as a measurement unit 11, a straight line identification unit 12, a variable generation unit 13, a calculation unit 14, a convergence determination unit 15, an acquisition unit 16, etc., as shown in Figure 2. In other words, the control unit 10 functions as a calculation device that calculates the position and orientation of the camera 30.

[0021] The encoder 40 is a component equipped with a sensor that detects the rotation angle of the wheels. The vehicle body position and attitude calculation unit 19 determines the direction and amount of movement of the vehicle body 2 based on the rotation angle of the front wheels 28 detected by the encoder 40. The forks 22 are raised and lowered by the cargo handling device control unit 18. The vertical position of the forks 22 can be determined by the amount of drive of the cargo handling device control unit 18.

[0022] (Calculation of camera 30's posture and position) When using camera 30 to control various aspects of forklift 1 (e.g., moving the vehicle body 2, operating the forks 22), it is necessary to determine the orientation and position of camera 30. However, determining the camera's position and orientation from a single image acquired by camera 30 without prior knowledge of the objects in the image is extremely difficult. The technology described herein, based on the Manhattan World Hypothesis (the hypothesis that artificial outdoor (city) and indoor (warehouse, factory, etc.) structures have strong parallel and right-angle characteristics), makes it possible to estimate the orientation and position of camera 30 from a single image. This is explained below.

[0023] In this specification, it is assumed that camera 30 can be approximated by a pinhole model, and position and orientation estimation is treated as a geometric problem. Also, horizontal surfaces, including the floor, may all be referred to as "ground."

[0024] The specific details of the calculation process for the attitude and position of camera 30 will be explained using the flowchart in Figure 3. In step S5, the measurement unit 11 acquires measured values ​​to determine the scale described later. The objects to be measured to obtain the measured values ​​can be various. For example, the actual distance between the first straight line SL and the second straight line SL2, described later, may be measured using a pocket laser rangefinder or the like. Alternatively, the height of the center of the camera 30's lens from the ground may be measured. The measurement may be performed automatically by the measurement unit 11, or the user may input the measured results to the control unit 10.

[0025] Furthermore, if the measured values ​​are stored in the memory of the control unit 10 and read out for use, step S5 can be omitted.

[0026] In step S10, an image is captured by camera 30. Figure 4 shows an example of an image captured by camera 30.

[0027] In step S20, the line identification unit 12 identifies the first line SL1, the second line SL2, and the third line SL3 on the captured image acquired by the camera 30. The first line SL1 to the third line SL3 are lines that exist on the same plane (ground) in the real world. The first line SL1 and the second line SL2 are a pair of parallel lines. The third line SL3 is a line that has a known angle with respect to the first line SL1 and the second line SL2. The selection of the first line SL1 to the third line SL3 may be performed automatically by the line identification unit 12 or by the user. Various methods such as RANSAC (RANDOM SAmple Consensus) can also be used.

[0028] Figure 4 shows specific examples of the first to third lines SL1 to SL3. Figure 5 shows the calculation results performed in step S50, which will be described later. Figure 5 corresponds to a vertical projection of Figure 4 (i.e., a view of the ground from above, looking vertically down). In the examples of Figures 4 and 5, the third line SL3 is a line perpendicular to the first line SL1 and the second line SL2.

[0029] In step S30, the variable generation unit 13 generates 13 variables representing the positions of both ends of the first straight line SL1, the positions of both ends of the second straight line SL2, the positions of both ends of the third straight line SL3, and the attitude and position of the camera 30. This will be explained below with reference to Figures 4 and 5.

[0030] The variable generation unit 13 defines the ground on which the first straight line SL1 to the third straight line SL3 exist as the XY plane. Also, the direction vertically upward with respect to the ground is defined as the Z-axis direction. That is, the plane where z = 0 is the ground. Further, a coordinate system is determined such that the directions of the first straight line SL1 and the second straight line SL2 are the Y-axis direction, and the direction of the third straight line SL3 is the X-axis. Thereby, the origin OC is determined. The origin OC is the reference point of the entire system. The origin OC may be at an arbitrary position because it can be coordinate-transformed later.

[0031] As shown in FIG. 5, the coordinate values of the end point P11 of the first straight line SL1 are (p 1x , p 1y , 0). The coordinate values of the end point P12 of the first straight line SL1 are (p 1x , p 2y , 0). The x-coordinates of the end points P11 and P12 are common. Also, the z-coordinate is 0. Therefore, the end points P11 and P12 can be represented using three variables p 1x , p 1y , p 2y .

[0032] Similarly, the coordinate values of the end point P21 of the second straight line SL2 are (q 1x , q 1y , 0). The coordinate values of the end point P22 of the second straight line SL2 are (q 1x , q 2y , 0). Therefore, the end points P21 and P22 can be represented using three variables q 1x , q 1y , q 2y .

[0033] The coordinate values of the end point P31 of the third straight line SL3 are (a 1x , a 1y , 0). The coordinate values of the end point P32 of the third straight line SL3 are (a 2x , a 1y , 0). The y-coordinates of the end points P31 and P32 are common. Also, the z-coordinate is 0. Therefore, the end points P31 and P32 can be represented using three variables a 1x , a 1y , a 2x .

[0034] Based on the above, the variables representing the coordinates of the six endpoints of the first, second, and third lines SL1, SL2, and SL3 can be reduced from 12 XY coordinate values ​​to 9. This is because, as mentioned above, conditions for parallelism and perpendicularity exist, and these conditions are incorporated into the method for determining the variables. In other words, the Manhattan world hypothesis is incorporated into the optimization objective function.

[0035] Note that the third line SL3 may not be perpendicular to the first line SL1 and the second line SL2, and may have a slope θ (a known value). In this case, the coordinate value of the endpoint P32 is (a 2x ,(a 2x -a 1y We can use )tanθ,0). In this case as well, we can have 9 variables representing the coordinates of the 6 endpoints.

[0036] The variable generation unit 13 also generates three variables, rx, ry, and rz, which represent the three-dimensional attitude angles of the camera 30. (rx, ry, rz) are variables obtained by multiplying a rotation axis vector, which is unitized to a length of 1, by the rotation angle Φ around that axis.

[0037] The variable generation unit 13 also generates a variable indicating the position of the camera 30. Specifically, the XY coordinate values ​​of the camera 30's position are set to 0. That is, the camera 30 is assumed to be located vertically above the origin OC in Figure 5. This allows the position of the camera 30 to be represented by a single variable h indicating its height in the Z direction. This makes it possible to avoid increasing the number of unnecessary variables.

[0038] Formulated as described above, the problem has 9 variables at the ends of the three straight lines, 3 variables for the orientation of camera 30, and 1 variable for the height (Z coordinate value of the position) of camera 30. Therefore, it can be reduced to an optimization problem with a total of 13 variables. Let's explain the effect of reducing the number of variables. The positions of the first straight line SL1 to the third straight line SL3 and the orientation and position of camera 30 have arbitrariness with respect to the direction of the overall coordinate axes, the origin position, and the scale mentioned above. Therefore, by reducing the number of variables as described above, the arbitrariness as a minimization problem can be reduced as much as possible. The only arbitrariness left is the scale.

[0039] In step S40, the calculation unit 14 converts the first line SL1, the second line SL2, and the third line SL3 into a camera image using camera parameters. That is, it determines the projected positions of the endpoints P11 and P12 of the first line SL1, P21 and P22 of the second line SL2, and P31 and P32 of the third line SL3 onto the camera screen. This will be explained in detail below. The position of a point in 3D is (x,y,z) T Let T represent the transpose of the matrix. Let (u,v) be the 2D X,Y position on the screen. The 3D position of the point (x,y,z) is given by the following equation. T This can be converted to a 2D X,Y position (u,v).

number

number

number

[0040] Based on the above formula, the two-dimensional X,Y positions (u,v) of the endpoints of the first line SL1, the second line SL2, and the third line SL3 are calculated. This gives the endpoint P11(p 1u ,p 1v ), end point P12(p 2u ,p 2v ), end point P21(q 1u ,q 1v ), end point P22(q 2u ,q 2v ), end point P31(a 1u ,a 1v ), end point P32(q 2u ,q 2v ), is calculated.

[0041] In step S50, the calculation unit 14 calculates the solution for 13 variables such that the sum of the squares of the deviations between the values ​​at both ends of the first line SL1, the second line SL2, and the third line SL3 on the captured image, and the two-dimensional X,Y positions (u,v) calculated above, is minimized (more precisely, local minimum). That is, p 1x ,p 1y ,p 2y ,q 1x ,q 1y ,q 2y ,a 1x ,a 1y ,a 2x The solution for each set of variables (variables representing the endpoints of the first line SL, the second line SL2, and the third line SL3), and rx, ry, rz (variables representing the camera's 3D attitude angles), and h (variable representing the camera's height) is calculated (simultaneously). When the model is holed in convergence, one set of solutions is obtained. This is not a unique solution but a particular solution, although the only remaining arbitrariness is in the scale.

[0042] This optimization problem is mechanically differentiable using algebraic equations because the objective function is derived from a combination of coordinate transformations and the transformation of the central projection of a pinhole camera. It can be solved using the Levenberg-Marquardt method (quasi-Newton's method) or Newton's method, and can achieve a computational speed sufficient for integration into onboard systems such as industrial vehicles.

[0043] In step S60, the convergence determination unit 15 determines whether the calculation has converged or not. This is because, since optimization is performed while retaining arbitrariness, convergence may be difficult under certain conditions. The convergence determination can be performed using the general convergence determination method of Newton's method.

[0044] Furthermore, since there is also arbitrariness in the front and back sides, it is necessary to check whether the solution has converged to the opposite side (e.g., the camera position below the ground). Note that by providing approximate initial values, it is possible to suppress convergence to the opposite side. The initial values ​​may be determined by self-position estimation (odometry, etc.) by the vehicle position and attitude calculation unit 19. If it is determined that convergence has occurred (S60: YES), proceed to step S70.

[0045] In step S70, the acquisition unit 16 determines the overall scale. As mentioned above, in the calculation in step S50, the overall scale is the only quantity that cannot be determined. Therefore, the scale is determined based on the measurement results measured by the measurement unit 11 in step S5. Then, by multiplying the entire system, excluding the attitude parameters, by a constant based on the determined scale, a general solution (real-world solution) can be obtained. As a result, it becomes possible to acquire the attitude and position of the camera 30.

[0046] (effect) In this embodiment, the camera's orientation can be easily determined from a single-viewpoint image captured by the camera 30. Since there is no need to specially acquire images for camera orientation estimation or to separately measure the positions of feature points within the image, the amount of preparation required is significantly reduced.

[0047] In this embodiment, the orientation and position of camera 30 can be determined from an image showing a pair of parallel lines (on the same plane) and a straight line whose angle with the parallel lines is known. These combinations of straight lines exist in many real-world locations, such as roads and crosswalks, or parts of the outer frame of a rectangular building. Therefore, this embodiment's technology has a wide range of applications and is suitable for practical use. [Examples]

[0048] Example 2 is a configuration in which the method for specifying the first straight line SL1, the second straight line SL2, and the third straight line SL3 differs from that of Example 1. Only the parts that differ from Example 1 will be described.

[0049] In step S20, the line identification unit 12 identifies a pair of parallel lines, a first line SL1 and a second line SL2, that exist in the first plane. It also identifies a second plane that intersects the first plane at a known angle, and identifies a third line SL3 that exists in that second plane.

[0050] Let's explain using a specific example. Figure 6 shows an example of an image captured by camera 30 used in Example 2. Figure 7 shows the calculation result calculated in step S50. Figure 7 corresponds to a vertical projection of Figure 6. The first line SL1 and the second line SL2 are a pair of parallel lines existing on the first plane (ground). The wall surface WF of the neighboring building corresponds to the second plane that intersects the first plane (ground) at a known angle (right angle). The third line SL3 exists on the second plane (wall surface WF). Specifically, the third line SL3 is a horizontal line located between the warehouse wall panel and the concrete portion below the warehouse wall.

[0051] In step S30, the variable generation unit 13 defines the ground on which the first line SL1 and the second line SL2 exist as the XY plane. It also defines the direction of the first line SL1 and the second line SL2 as the X-axis direction. Then, it determines the Y-axis direction using a right-handed system.

[0052] As shown in Figure 7, the coordinates of the endpoint P11 of the first straight line SL1 are (a 1x ,a 1y The coordinates of the endpoint P12 of the first line SL1 are (a 2x ,a 1y Similarly, the coordinates of the endpoint P21 of the second line SL2 are (b 1x ,b 1y The coordinates of the endpoint P22 of the second line SL2 are (b 2x ,b 1yThe coordinates of the endpoint P31 of the third line SL3 are (k 1x ,k 1y ,k 1z ) The coordinates of the endpoint P32 of the third line SL3 are (k 2x ,k 1y ,k 1z ) This is the result. Similar to Example 1, the number of variables can be reduced by using conditions for parallelism and perpendicularity.

[0053] The calculations performed by the calculation unit 14 in steps S40 and S50 are the same as in Example 1, so a detailed explanation will be omitted. However, when two planes are used as in Example 2, the dimension of arbitrariness in the local optimum is one dimension higher than when one plane is used as in Example 1. Therefore, in Example 2, the solution is obtained without knowing the positional relationship between the first plane (ground) and the second plane (wall WF) (where the line where both surfaces intersect is). Therefore, the scale of the ground and the scale of the wall WF are determined separately. In other words, there remains an independent arbitrariness in the height of the ground and the Y coordinate value of the wall WF. Thus, in the calculation, the Y coordinate value of the third line SL3 (k 1y ) can be arbitrarily determined. However, when determining the camera's attitude and position using the ground, the Y coordinate value k of the wall surface WF is used. 1y Since we do not use (i.e., the scale), any value is acceptable, and there is no need to specify an exact value.

[0054] In the real world (see Figure 6), the Y-axis position of the third line SL3 is located on the negative side compared to the second line SL2. However, in the calculation result of step S50 (see Figure 7), the Y-axis position of the third line SL3 is located on the positive side compared to the second line SL2. In other words, the Y-axis position of the third line SL3 is completely different in the real world and in the calculation result. This is because the calculation converges while ignoring the size (scale) of the wall WF. However, as mentioned above, this is not a problem because it is possible to accurately determine the camera's attitude and position in the ground coordinate system regardless of the location of the wall WF (third line SL3).

[0055] To determine the scale of the ground, information is needed regarding either the distance between two points on the ground or the height of camera 30 from the ground. The two points on the ground can be any two points that are visible in the image captured by camera 30 (Figure 6) and are known to be on the ground. The distance between the points only needs to be a value that allows for accurate determination of their position (u,v) on the screen.

[0056] (effect) Even when only one pair of parallel lines can be identified on the ground, the orientation and position of camera 30 can be calculated by using a straight line contained in a wall surface whose angle with the ground is known. This expands the range of applications of the technology described herein.

[0057] The surface of a building has a higher plane accuracy than the ground. This is because the ground has various slopes, taking into account factors such as drainage. In the technology of this embodiment, the attitude and position of the camera 30 can be calculated using the highly accurate surface of the building, thereby improving the accuracy of the calculation results.

[0058] The walls of buildings are generally perpendicular to the ground. Therefore, by using the straight lines contained within the walls, it becomes possible to easily identify the straight lines that follow the Manhattan World Hypothesis.

[0059] Although examples of the technology disclosed herein have been described in detail above, these are merely illustrative and do not limit the scope of the claims. The technology described in the claims includes various modifications and changes to the specific examples illustrated above.

[0060] (modified version) The scope of application of the technology described herein is not limited to estimating the attitude and position of cameras installed on moving objects or buildings. For example, it can also be applied to determining the attitude and position of objects captured by a camera (e.g., truck bed, shelves). Using the example image in Figure 8, we will explain how to calculate the attitude and position of a truck bed. In step S20, the straight line identification unit 12 identifies a pair of parallel straight lines (first straight line SL1, second straight line SL2) and a straight line (third straight line SL3) perpendicular to the pair of straight lines that exist within the same plane (side panel). Since the length of the side panel in the vehicle's longitudinal direction and its height in the vehicle's height direction are known, the attitude and position of the camera relative to the truck bed can be calculated using the technology described herein. As a result, it becomes possible to determine the attitude and position of the truck bed.

[0061] The technology described herein is also applicable to camera calibration, which determines the positional relationship between an object captured by the camera (e.g., the forks of a forklift) and the camera. This will be explained using an example image in Figure 9. Figure 9 is an example image when the camera 30 is mounted on the backrest 23 (see Figure 1) of the forklift 1. In step S20, the straight line identification unit 12 identifies a pair of parallel straight lines (first straight line SL1, second straight line SL2) that exist in the same plane, and a straight line (third straight line SL3) that is perpendicular to the pair of straight lines. Since the lengths and widths of the forks 22a and 22b are known, the attitude and position of the camera 30 relative to the forks 22 can be calculated using the technology described herein. Even when the forks are replaced, the attitude and position of the fork tips can be easily determined.

[0062] In this embodiment, the case in step S20 where the line identification unit 12 identifies three lines was described, but the embodiment is not limited to this. It is possible to identify four or more lines. The more lines that are known to lie on an accurate plane are increased, the higher the accuracy of the calculation results can be. Note that, following the Manhattan World Hypothesis, the number of variables increases by three for each additional line. However, if the number of variables is in the range of several tens, the burden of calculation processing can be suppressed, making it possible to incorporate it into the control cycle of industrial vehicles, etc. Furthermore, if the number of variables is in the range of several hundred, it is possible to solve it stably.

[0063] In this embodiment, we have described a case where a pair of parallel lines are identified on the ground, but the embodiment is not limited to this. For example, in the captured image in Figure 6, the lines SL4 and SL5 present on the wall surface WF may be identified as a pair of parallel lines. In reality, the wall surface WF has higher planar accuracy than the ground. Therefore, by using the wall surface WF to calculate the attitude and position of the camera 30, it is possible to improve the accuracy of the calculation results.

[0064] The rotation transformation (matrix) R can take various forms. For example, it is possible to use a three-way rotation matrix, commonly used in automobiles, which rotates θz around the Z axis, then θy around the Y axis, and then θx around the X axis. In this case, the parameters of the rotation transformation are θx, θy, and θz. The order of rotations must also be predetermined.

[0065] In this embodiment, a forklift 1 was described as a mobile body equipped with a camera 30, but the invention is not limited to this form. The technology described herein can be applied to various other mobile bodies (e.g., mobile robots, mobile manipulators, drones). Furthermore, the technology described herein can be applied not only to mobile bodies but also to cameras 30 mounted on stationary objects.

[0066] The control unit 10 is not limited to being mounted on the forklift 1. For example, it may be a server located on a network such as the internet.

[0067] The technical elements described herein or in the drawings demonstrate technical usefulness individually or in various combinations, and are not limited to the combinations described in the claims at the time of filing. Furthermore, the technologies illustrated herein or in the drawings achieve multiple objectives simultaneously, and achieving even one of these objectives constitutes technical usefulness in itself. [Explanation of Symbols]

[0068] 1: Forklift 10: Control Unit 11: Measurement Unit 12: Line Identification Unit 13: Variable Generation Unit 14: Calculation Unit 15: Convergence Judgment Unit 16: Acquisition Unit 30: Camera SL1: First Line SL2: Second Line SL3: Third Line

Claims

1. A computing device that calculates the camera's orientation, A line identification unit that identifies a pair of parallel first and second lines on an image captured using the camera, and a third line having a known angle with respect to the first and second lines, A variable generation unit that generates a plurality of variables representing the positions of both ends of the first straight line, the positions of both ends of the second straight line, the positions of both ends of the third straight line, the attitude and position of the camera, A calculation unit calculates a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the ends of the second line, and the ends of the third line onto the camera screen and the positions of the ends of the first line, the ends of the second line, and the ends of the third line on the captured image is minimized. An acquisition unit that acquires the camera's pose from the solutions of the aforementioned multiple variables, Equipped with, The variable generation unit defines an XY plane in which the first and second lines exist, where the direction of the first and second lines is either the X-axis or the Y-axis, and generates three variables representing the endpoints of the first line, three variables representing the endpoints of the second line, and three variables representing the endpoints of the third line. The variable generation unit generates three variables that represent the three-dimensional attitude angles of the camera, The variable generation unit generates a single variable indicating the position of the camera by setting the camera's position to the origin of the XY plane. Computing device.

2. The calculation device according to claim 1, wherein the first line, the second line, and the third line are lines that lie in the same plane.

3. The first line and the second line are lines that lie in the first plane. The calculation device according to claim 1, wherein the third line is a line that lies in a second plane that intersects the first plane at a known angle.

4. The computing device according to claim 3, wherein one of the first plane and the second plane is the ground, and the other is a plane of an object placed on the ground.

5. The calculation device according to any one of claims 1 to 4, wherein the known angle is a right angle.

6. The system further includes a measuring unit for measuring the actual distance between the first straight line and the second straight line, or the height of the center of the camera lens from the ground. The calculation device according to any one of claims 1 to 5, wherein the acquisition unit acquires the scales of the endpoints of the first line, the endpoints of the second line, the endpoints of the third line, and the position of the camera from among the solutions of the plurality of variables, based on the measurement results of the measurement unit.

7. A calculation method for determining the camera's orientation, The steps include identifying a pair of parallel first and second lines in the captured image taken using the camera, and a third line having a known angle with respect to the first and second lines, A generation step of generating a plurality of variables representing the positions of both ends of the first line, the positions of both ends of the second line, the positions of both ends of the third line, the pose and position of the camera, The steps include: calculating a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the ends of the second line, and the ends of the third line onto the camera screen and the positions of the ends of the first line, the ends of the second line, and the ends of the third line on the captured image is minimized; A step of obtaining the camera's pose from the solution of the aforementioned multiple variables, Equipped with, The generation step involves defining an XY plane in which the first and second lines exist, where the direction of the first and second lines is either the X-axis or the Y-axis, thereby generating three variables indicating the endpoints of the first line, three variables indicating the endpoints of the second line, and three variables indicating the endpoints of the third line. The generation step generates three variables that represent the three-dimensional attitude angles of the camera, The generation step generates a single variable representing the camera's position by setting the camera's position to the origin of the XY plane. Calculation method.

8. An information processing program that is loaded into the computer of an information processing device and calculates the camera's pose, A line identification means for identifying a pair of parallel first and second lines in an image captured using the aforementioned camera, and a third line having a known angle with respect to the first and second lines, A variable generation means that generates a plurality of variables representing the positions of both ends of the first straight line, the positions of both ends of the second straight line, the positions of both ends of the third straight line, the attitude and position of the camera, A calculation means for calculating a solution for the plurality of variables such that the sum of the squares of the deviations between the projected positions obtained by projecting the ends of the first line, the ends of the second line, and the ends of the third line onto the camera screen and the positions of the ends of the first line, the ends of the second line, and the ends of the third line on the captured image is minimized. An acquisition means for acquiring the camera's pose from the solutions of the aforementioned multiple variables, The information processing device is made to function, The variable generation means defines an XY plane in which the first and second lines exist, where the direction of the first and second lines is either the X-axis or the Y-axis, thereby generating three variables representing the endpoints of the first line, three variables representing the endpoints of the second line, and three variables representing the endpoints of the third line. The variable generation means generates three variables that represent the three-dimensional attitude angles of the camera, The variable generation means generates a single variable indicating the position of the camera by setting the position of the camera to the origin of the XY plane. Information processing program.

Citation Information

Patent Citations

  • Single camera calibration

    EP3901913A1

  • Camera attitude parameter estimation device

    JP2011215063A

  • On-vehicle camera system, and calibration method and program for same

    JP2013115540A

  • Image processing device, image processing program and image processing method

    JP2019105992A

  • Direction detection system, direction detection method, and direction detection program

    JP2019174243A