Image processing device and image processing method
By dividing images into areas based on camera position, setting thresholds for uniform feature point distribution, and matching points across images, the image processing device improves calibration accuracy for three-dimensional position recognition.
Patent Information
- Application Number
- JP2024531786
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-12-04
- Estimated Expiration
- 2042-07-05
AI Technical Summary
Existing image processing devices face reduced calibration accuracy due to unsuitable distribution of feature points in captured images, leading to discarded calibration opportunities.
The image processing device employs an area division unit to divide captured images into multiple areas based on camera mounting position, a feature point detection unit to detect feature points exceeding a threshold in each area, and a threshold setting unit to ensure uniform distribution, followed by a feature point matching unit to associate and match points across images, and a relative parameter calculation unit to calculate accurate parameters based on these points.
This approach enhances the accuracy of image processing device calibration, enabling more precise recognition of three-dimensional positions in the external world.
Smart Images

Figure 0007780648000001 
Figure 0007780648000002 
Figure 0007780648000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device and an image processing method. [Background technology]
[0002] In order to recognize the three-dimensional position of the external world captured in an image using an image processing device, the image processing device must be calibrated. Calibration is performed by calculating camera parameters and relative parameters. Camera parameters are parameters that indicate the lens distortion and focal length of the camera used in the image processing device. Relative parameters are parameters that indicate the relative position and orientation of two cameras or a single moving camera before and after movement.
[0003] When the camera that captures the images is a stereo camera consisting of two cameras, one of the cameras is set as the reference, and the position and orientation of the other camera relative to the reference camera are calculated as relative parameters. Also, using the position of the right camera as the reference, the image processing device calculates the relative three-dimensional position ((x, y, z) axis position) and relative orientation ((x, y, z) axis angle) of the left camera to the reference, which is called calibration.
[0004] A known technique for calibrating an image processing device is described in Patent Document 1. Patent Document 1 states that "an image in which feature points are sufficiently distributed at a distance according to the intended use of the two cameras is selected, and the extrinsic parameters of the two cameras are calculated using the image coordinates of the feature points and corresponding points distributed at a distance according to the intended use of the two cameras in the image." [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Laid-Open No. 2009-276233 Summary of the Invention [Problem to be solved by the invention]
[0006] The distribution of feature points detected from a captured image is important for highly accurate calibration of an image processing device. Captured images with feature point distributions that are not suitable for calibration result in reduced calibration accuracy. In the method described in Patent Document 1, which judges captured images based on the distance distribution of feature points and extracts only captured images that are expected to produce highly accurate calibration results, and then calibrates the image capture device, only captured images that are expected to produce highly accurate calibration results are extracted. However, with the technology disclosed in Patent Document 1, captured images with a calculated number of feature points below a threshold are not suitable for calibration. Captured images that are not suitable for this configuration are discarded, which reduces the opportunities for calibration.
[0007] The present invention has been made in view of the above circumstances, and has as its object to improve the accuracy of calibration of an image processing device. [Means for solving the problem]
[0008] The image processing device according to the present invention comprises: The position and size of the area into which the multiple images captured by the in-vehicle camera are divided are changed according to the mounting position of the in-vehicle camera, an area dividing unit that divides each of a plurality of captured images into a plurality of divided areas; a feature point detecting unit that detects portions of each divided area that exceed a threshold as feature points; a threshold setting unit that sets a threshold for each divided area so that the distribution of feature points detected in each divided area is uniform; and a feature point matching unit that matches the feature points detected in the plurality of captured images by associating them with each other among the plurality of captured images; When the on-board camera is a stereo camera, the axial position and axial angle of one of the two cameras constituting the stereo camera when the one camera captures an image are used as a reference, and the relative axial position and axial angle of the other camera when the other camera captures another image are used as relative parameters that indicate the relationship between the relative positions and attitudes of the two cameras, The image processing device further includes a relative parameter calculation unit that calculates relative parameters based on the feature points associated with each other among the plurality of captured images. The image processing device of the present invention also includes an area division unit that changes the position and size of the areas into which multiple captured images taken by the on-board camera are divided depending on the installation position of the on-board camera, and divides each of the multiple captured images into multiple divided areas; a feature point detection unit that detects parts that exceed a threshold for each divided area as feature points; a threshold setting unit that sets a threshold for each divided area so that the distribution of feature points detected in each divided area is even; a feature point matching unit that matches and matches the feature points detected in the multiple captured images between the multiple captured images; and, if the on-board camera is a monocular camera, a relative parameter calculation unit that calculates relative parameters based on the feature points matched between the multiple captured images, using the axial position and axial angle of the monocular camera when one of the multiple captured images was taken as a reference and the relative axial position and axial angle of the monocular camera when the other captured images were taken as relative parameters, when the on-board camera is a monocular camera. [Effects of the Invention]
[0009] According to the present invention, by improving the accuracy of calibration of the image processing device, the image processing device can recognize three-dimensional positions in the external world more accurately. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing an example of the internal configuration of an image processing device according to an embodiment of the present invention; [Figure 2] 10 is a flowchart illustrating an example of overall processing of an image processing device according to an embodiment of the present invention. [Figure 3] 10 is a flowchart showing an example of an equalization feature point detection process performed by an equalization feature point detection unit according to an embodiment of the present invention. [Figure 4] 1 is a diagram illustrating an example of an image divided into a uniform grid according to an embodiment of the present invention; [Figure 5] 10A and 10B are diagrams illustrating an example of an image divided into a grid by increasing the number of divisions according to an embodiment of the present invention. [Figure 6] 10A and 10B are diagrams illustrating an example of a captured image divided according to the situation of a vehicle according to an embodiment of the present invention. [Figure 7] FIG. 10 is a diagram showing a state in which a photographed image photographed at a time before A in the time series is divided according to one embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing a state in which a photographed image photographed at time B in the time series has been divided according to one embodiment of the present invention. [Figure 9] FIG. 4 is a diagram showing an example of feature points detected from a first captured image according to an embodiment of the present invention. [Figure 10] FIG. 10 is a diagram showing an example of feature points detected from a second captured image according to an embodiment of the present invention. [Figure 11] FIG. 10 is a diagram showing an example of a rectified image in which the distribution of feature points is biased, to which a conventional technique is applied. [Figure 12] 10 is an example of a rectified image in which the distribution of feature points is equalized by applying a technique according to an embodiment of the present invention. [Figure 13] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a computer according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functions or configurations are designated by the same reference numerals, and redundant description will be omitted. The present invention is applicable to, for example, a computing device for vehicle control capable of communicating with an on-board ECU (Electronic Control Unit) for an Advanced Driver Assistance System (ADAS) or Autonomous Driving (AD).
[0012] [One embodiment] FIG. 1 is a block diagram showing an example of the internal configuration of an image processing device 100 according to an embodiment. The image processing device 100 is one of the devices mounted in an image processing ECU (not shown) mounted on a vehicle. The image processing device 100 also has a highly accurate calibration function.
[0013] The on-board camera 101 is an example of a camera mounted on a vehicle so that the image processing device 100 can recognize the outside world of the vehicle, and captures images of the outside world at predetermined times. The on-board camera 101 outputs captured images 10 of the surroundings of the vehicle to the equalized feature point detection unit 102 of the image processing device 100 in order to recognize the three-dimensional distance of the outside world. The on-board camera 101 may be two stereo cameras or one monocular camera.
[0014] The captured images 10 output by the in-vehicle camera 101 to the equalized feature point detection unit 102 may be two captured images 10 captured simultaneously by two cameras (stereo cameras), or may be two captured images 10 captured at different times by one camera (monocular camera). When a monocular camera is used as the in-vehicle camera 101, for example, one captured image 10 captured by the monocular camera of an object while the vehicle is stationary, and another captured image 10 captured of the same object after the vehicle has moved slightly, are output.
[0015] The image processing device 100 calibrates the relative parameters of the on-board camera 101 based on two captured images 10, and performs processing to recognize the three-dimensional distance of the external world from the captured images 10. Here, the calibration by the image processing device 100 is performed each time the on-board camera 101 is started up. For example, calibration is performed immediately after the vehicle is turned on and the on-board camera 101 is started up, or at predetermined time intervals while the vehicle is traveling. The image processing device 100 includes an equalized feature point detection unit 102, a feature point matching unit 103, a relative parameter calculation unit 104, a parallax image generation unit 105, and a three-dimensional distance recognition unit 106. The processing of each functional unit will be described below with reference to FIG. 2.
[0016] 2 is a flowchart showing an example of the overall processing of the image processing device 100. An example of an image processing method of the image processing device 100 is shown in FIG.
[0017] First, the equalized feature point detection unit 102 acquires a captured image 10 from the vehicle-mounted camera 101 (S1). The received captured image 10 is stored in a storage device such as a RAM 53 or nonvolatile storage 55 shown in Fig. 13 (described later), and is read from the storage device and written again as appropriate in the processing of each functional unit of the image processing device 100. In the following explanation, the processing of inputting and outputting the captured image 10 to and from the storage device will be omitted.
[0018] Then, in step S1, the equalized feature point detection unit 102 analyzes the captured image 10 received from the in-vehicle camera 101 and detects the contours of objects (roadways, sidewalks, other vehicles, signs, etc.) that appear in the captured image 10. After that, the equalized feature point detection unit 102 performs the equalized feature point detection process shown in Fig. 3 (S2) and detects characteristic points such as corners of the contour lines and intersections of the contour lines as feature points.
[0019] The equalized feature point detection unit 102 applies a unique method in the equalized feature point detection process shown in step S2, and performs operations to detect feature points that will increase the accuracy of the relative parameters calculated later by the relative parameter calculation unit 104. Operations by the equalized feature point detection unit 102 include dividing the captured image 10, detecting and counting the number of feature points in the areas divided by the division process, and setting a threshold value (individual threshold value) used for detecting feature points. Through the operations of the equalized feature point detection unit 102, feature points are detected so that the number of feature points in each divided area into which the captured image 10 is divided by the division process is equal.
[0020] As shown in FIG. 1, the equalized feature point detection unit 102 includes an area division unit 1021, a feature point detection unit 1022, and a threshold setting unit 1023. The area dividing unit (area dividing unit 1021) divides each of a plurality of captured images 10 taken by the vehicle-mounted camera (vehicle-mounted camera 101) into a plurality of divided areas. The feature point detection unit (feature point detection unit 1022) detects, for each divided region, a portion that exceeds a threshold (individual threshold) as a feature point. The threshold setting unit (threshold setting unit 1023) sets a threshold (individual threshold) for each divided region so that the distribution of feature points detected for each divided region is uniform. The equalized feature point detection process performed by each functional unit of the equalized feature point detection unit 102 will be described below with reference to FIG. 3 and subsequent figures.
[0021] 3 is a flowchart showing an example of equalization feature point detection processing performed by the equalization feature point detection unit 102. The equalization feature point detection processing is part of an image processing method performed by each functional unit of the image processing device 100 shown in FIG.
[0022] First, the area dividing unit 1021 divides the captured image 10 received from the vehicle-mounted camera 101 into a predetermined size and number of divisions (S11). Here, an example of the processing performed by the area dividing unit 1021 and examples of divided areas will be described with reference to FIGS.
[0023] For example, the area dividing unit 1021 takes in one of two captured images 10 acquired from the in-vehicle camera 101. The area dividing unit (area dividing unit 1021) divides the captured image 10 by increasing the number of divisions and reducing the size of the divided areas so that at least one feature point is detected in each divided area. Examples of dividing the captured image 10 include those shown in Figures 4 and 5.
[0024] FIG. 4 is a diagram showing an example of an image (captured image 10) divided into an equal grid pattern. The captured image 10 captured by the in-vehicle camera 101 shown in FIG. 4 is divided into four equally spaced grid-shaped divided regions of 2 rows and 2 columns by the region dividing unit 1021. In this case, the number of divisions of the captured image 10 is "4". Furthermore, the size of each divided region is the same. By dividing the captured image 10 into four, the size of each divided region becomes one-fourth the size of the original captured image 10. Even when dividing the captured image 10 into four in this way, it is expected that feature points in the captured image 10 will be evenly distributed in each divided region, depending on the object.
[0025] FIG. 5 is a diagram showing an example of an image divided into a grid by increasing the number of divisions. The captured image 10 captured by the in-vehicle camera 101 shown in FIG. 5 is divided by the region dividing unit 1021 into 16 equally spaced grid-shaped divided regions of 4 rows and 4 columns. In this case, the number of divisions of the captured image 10 is "16." The divided regions are all the same size. By increasing the number of divisions, the size of the divided regions shown in FIG. 5 becomes one-fourth the size of the divided regions shown in FIG. 4. By increasing the number of divisions, the region dividing unit 1021 can evenly distribute feature points included in the captured image 10 in each divided region.
[0026] The reason why the region dividing unit 1021 divides the captured image 10 into more parts and thereby distributes feature points evenly within the captured image 10 will be explained below. For example, uneven distribution of feature points within one region cannot be corrected later. For this reason, if signs and vehicles appear in the upper half of the captured image 10 before division and a roadway appears in the lower half, it is expected that more feature points will be detected unevenly in the upper half of the captured image 10 than in the lower half.
[0027] Therefore, by increasing the number of divisions of the captured image 10 and dividing the captured image 10 into smaller divided regions, even if there is a bias in the feature points within each divided region, the feature points are evenly distributed across the entire captured image 10. For example, if 100 feature points are randomly distributed across the entire captured image 10, the feature points will be unevenly distributed across the divided regions. This is because, if 100 feature points are randomly distributed across the entire captured image 10 and then the entire captured image 10 is divided into 100 equal parts, it is expected that some divided regions will contain one or more feature points and some will not contain any feature points. On the other hand, if the entire captured image 10 is divided into 100 equal parts and one feature point is distributed across each divided region, the feature points can be evenly distributed across the entire captured image 10. Therefore, by equally dividing the captured image 10 and then detecting feature points in each divided region, the image processing device 100 can perform calibration using the feature points detected across the entire captured image 10.
[0028] FIG. 6 is a diagram showing an example of a captured image 10 divided in accordance with the vehicle situation. Considering that the on-board camera 101 is mounted on a vehicle, it is highly likely that objects at relatively close range will appear large in the lower part of the captured image 10, and objects at a long distance will appear small in the upper part of the captured image 10. Furthermore, if there are no other vehicles traveling ahead of the vehicle, the lower part of the captured image 10 will show a large area of the roadway with flat density.
[0029] Therefore, the region dividing unit (region dividing unit 1021) changes the positions and sizes of the regions into which the captured image 10 is divided, depending on the mounting position of the on-board camera (on-board camera 101). For example, the region dividing unit 1021 divides the captured image 10 into a small number of regions and large divided regions in areas where close-range objects appear large, taking into account the vehicle situation and the characteristics of the captured image 10. On the other hand, the region dividing unit 10 divides the captured image 10 into a large number of regions and small divided regions in areas where long-range objects appear small.
[0030] In an example of the segmentation process shown in FIG. 6, the distance from the vehicle to the object in the far-distance portion (the upper half of the captured image 10) in the direction of travel of the vehicle must be recognized in detail by the 3D distance recognition unit 106 (see FIG. 1). To recognize feature points from the captured image 10, the region segmentation unit 1021 of the equalized feature point detection unit 102 determines the segmentation pattern (the number of segments and the segmentation size of each segment) of the captured image 10, taking into consideration the driving scene, such as the average size of objects captured in the captured image and the direction and extent of movement of feature points in that portion over time. Therefore, the region segmentation unit (region segmentation unit 1021) segments the captured image 10 in a portion predicted to have many feature points into smaller segments than portions predicted to have few feature points. For example, in FIG. 6, the portion near the center of the far-distance portion (the portion showing the roadway) is segmented into smaller segments, and the portions near both ends of the far-distance portion (the portion showing the sidewalk) are segmented into larger segments.
[0031] In this way, the area division unit 1021 can divide the captured image 10 into areas that are expected to have few feature points (close-up roadways, sidewalks, etc.) with larger sizes than areas that are expected to have many feature points (distant roadways, sidewalks, etc.).
[0032] The area dividing unit 1021 changes the number of divisions and the division size according to the characteristics of the object shown in the captured image 10, so that the number of feature points detected from the divided area showing a close-up object and the number of feature points detected from the divided area showing a far-up object can be made closer to equal. As a result, it is possible to improve the accuracy of calibration of the captured image 10 in the depth direction by the image processing device 100.
[0033] Next, an example will be described in which the region dividing unit 1021 performs different divisions from previous and next time-series images of a traveling vehicle (pre-time series A and post-time series B). 7 and 8 are diagrams showing examples of captured images 10 captured by the vehicle-mounted camera 101 at different times. FIG. 7 is a diagram showing how a captured image 10 captured at a time A in the time series is divided. Fig. 8 is a diagram showing how the captured image 10 captured at time B in the time series is divided. The captured image 10 shown in Fig. 8 was captured a time t after the captured image 10 captured in Fig. 7. In other words, the captured image 10 shown in Fig. 8 is an image captured by the in-vehicle camera 101 with the vehicle moving forward for the time t.
[0034] The area division unit (area division unit 1021) divides the multiple captured images 10 into different sizes based on prediction results of feature points corresponding to three-dimensional positions in the outside world detected from the multiple captured images 10 captured by the on-board camera (on-board camera 101) as the vehicle moves. For example, the area division unit 1021 predicts the locations of feature points corresponding to three-dimensional positions detected from each of the captured image 10 captured at an earlier time point A in the time series and the captured image 10 captured at a later time point B in the time series. Then, the area division unit 1021 divides the two captured images 10 into different division sizes so that the predicted positions of the feature points are included.
[0035] Comparing the captured image 10 shown in FIG. 7 and FIG. 8, the height H1 of the upper region divided into eight divided regions in the image shown in FIG. 7 is height H2 in the captured image 10 shown in FIG. 8, making the upper region of the captured image 10 wider. The left side of the image shown in FIG. 7 is divided into width W11, which includes the roadway in the distance, and width W12, which includes the sidewalk in the distance. Meanwhile, when the captured image 10 shown in FIG. 8 was captured, a vehicle was moving forward and approaching the object. Therefore, the left side of the image shown in FIG. 8 is wider than width W11 shown in FIG. 7, as indicated by width W21, which includes the roadway in the distance. Meanwhile, width W22, which includes the sidewalk in the distance, of the left side of the image shown in FIG. 8 is narrower than width W12 shown in FIG. 7.
[0036] In order to match the divided areas of the photographed image 10 taken before time series A and the photographed image 10 taken after time series B, it is necessary that there are divided areas of the photographed image 10 taken before time series A and divided areas of the photographed image 10 taken after time series B. For this reason, the two photographed images 10 are assumed to have the same number of divisions.
[0037] In this way, the area dividing unit 1021 changes the division size of the captured image 10 before and after the time series, so that the feature point detecting unit 1022 can reliably detect feature points of an area (for example, a roadway) where there is an object to note in the traveling direction of the vehicle. Furthermore, the feature point matching unit 103, which will be described later, can reduce malfunctions of matching the wrong divided areas when matching divided areas with each other.
[0038] Returning to FIG. 3 again, the explanation will be continued. After the process of step S11, the feature point detection unit 1022 selects one of the divided areas divided by the area division unit 1021 (S12). Then, the feature point detection unit 1022 performs feature point detection for the divided area selected in step S12 using an individual threshold value (S13).
[0039] In the feature point detection process, a detection threshold is used to increase or decrease the number of feature points that the feature point detection unit 1022 detects in a divided area. Typically, one feature point detection threshold value is set for each captured image 10. However, the feature point detection unit 1022 according to this embodiment can set a detection threshold value individually for each divided area. Therefore, the detection threshold set for each divided area is called an "individual threshold." The individual threshold value is a value set for each divided area by the threshold setting unit 1023, which will be described later.
[0040] The individual thresholds set by the threshold setting unit 1023 have a settable range of values. Here, a luminance threshold is used as the individual threshold. Using the luminance threshold enables binarization processing for each pixel, such that areas above a certain level of luminance are white and areas below a certain level of luminance are black. For example, increasing the luminance threshold increases the number of black areas in the captured image 10, while decreasing the luminance threshold increases the number of white areas in the captured image 10. If the luminance threshold is increased to its maximum value, the captured image 10 is completely filled with black, and if the luminance threshold is decreased to its minimum value, the captured image 10 is completely filled with white. In either case, feature points cannot be detected. For this reason, a range of luminance thresholds effective for feature point detection is determined for each driving scene. The threshold setting unit 10 then increases or decreases the individual threshold within this range.
[0041] After step S13, the feature point detection unit 1022 counts the number of feature points in the divided area detected in step S13 (S14). The number of feature points in the divided area counted by the feature point detection unit 1022 is temporarily stored in the RAM 53 (see FIG. 13) or the like.
[0042] After the processing by the feature point detection unit 1022 (S12 to S14), the processing by the threshold setting unit 1023 is started. The threshold setting unit 1023 sets a threshold to equalize the feature points detected by the feature point detection unit 1022. Therefore, for each divided area in which the number of feature points was counted in step S14, the threshold setting unit 1023 compares the number of feature points with a target value and determines whether the number of feature points is greater than the target value (S15). In steps S15 to S17, the threshold setting unit (threshold setting unit 1023) sets a higher threshold for divided areas in which the number of feature points is greater than the target value, and sets a lower threshold for divided areas in which the number of feature points is less than the target value. Therefore, individual thresholds are set so that the number of feature points detected by the feature point detection unit 1022 is close to the target value. The number of feature points will now be described with reference to FIGS. 9 and 10.
[0043] FIG. 9 is a diagram showing an example of feature points detected from the first captured image 10. In FIG. FIG. 10 is a diagram showing an example of feature points detected from the second photographed image 10. In FIG. 9 and 10 represent feature points. However, the solid circle represents a feature point that indicates the same part in two captured images 10 and is detected and matched, while the dashed circle represents a feature point that is not detected in the same part in the two captured images 10.
[0044] Typically, feature point matching is performed by comparing the feature amount of one feature point in one captured image 10 (base image) with the feature amounts of all feature points in the other captured image 10 (reference image) to find the feature point with the closest feature amount. A feature amount is a numerical value that indicates the shape of a feature point, and is obtained by converting the arrangement of pixels (brightness values) around the feature point into a numerical value using a certain procedure. Feature points can take on a variety of shapes, but if they represent the same object, they are expected to be the same feature points. However, because feature point matching is performed on divided areas at the same position in the two captured images 10, the load is lighter than when matching feature points detected across the entire captured image 10, and the time required to complete matching can also be shortened.
[0045] Returning to the explanation of Figure 3. Assume that the upper left side (back side) of the captured image 10 shown in FIG. 9 is a divided area divided by the division pattern shown in FIG. 6. In this case, many feature points are detected in the divided area on the upper left side (back side) of the captured image 10, so the threshold setting unit 1023 determines that the number of feature points in this divided area is greater than the target value (NO in S15). Therefore, the threshold setting unit 1023 performs a process of increasing the individual threshold (S16). As a result, the detection threshold in this divided area becomes the maximum value. In other words, it becomes the maximum detection threshold within the settable range. However, by dividing the current individual threshold up to the maximum value at predetermined intervals, the individual threshold may be set at the divided intervals in step S16. Therefore, the individual threshold may be set to a value smaller than the maximum value.
[0046] On the other hand, suppose that the lower side (near side) of the photographed image 10 shown in FIG. 9 is a divided area divided by the division pattern shown in FIG. 6. In this case, since almost no feature points are detected in the divided area on the lower side (near side) of the photographed image 10, the threshold setting unit 1023 determines that the number of feature points in this divided area is equal to or less than the target value (NO in S15). Therefore, the threshold setting unit 1023 performs a process of lowering the feature point detection threshold (S17). As a result, the detection threshold in this divided area becomes the minimum value. In other words, it becomes the smallest detection threshold within the settable range. However, by dividing the current individual threshold value to the minimum value at predetermined intervals, the individual threshold value may be set at the divided intervals in step S17. For this reason, the individual threshold value may be set to a value greater than the minimum value.
[0047] After step S16 or S17, the equalization feature point detection unit 102 determines whether the processing from step S12 onward has been performed on all divided regions divided by the region division unit 1021 (S18). If the equalization feature point detection unit 102 determines that there are divided regions that have not been processed (NO in S18), it performs the processing from step S12 on the unprocessed divided regions again. Here, the feature point detection unit (feature point detection unit 1022) detects feature points using a threshold value set by a threshold setting unit (threshold setting unit 1023). Therefore, the number of feature points detected by the feature point detection unit 1022 matches the target value. On the other hand, if the equalization feature point detection unit 102 determines that the processing has been performed on all divided regions (YES in S18), it ends the equalization feature point detection processing shown in FIG. 3.
[0048] 3, after individual thresholds are set for a certain divided area by the processes of steps S16 and S17, the feature point detection unit 1022 performs feature point detection processing for another divided area using the individual thresholds set in the previous process. The reason for performing such processing is to shorten the processing time for all divided areas when the equalization feature point detection processing shown in FIG. 3 is implemented.
[0049] However, if the image processing device 100 has sufficient processing power, it is more effective for the feature point detection unit 1022 to perform feature point detection again in the same divided region. In this case, the feature point detection unit 1022 detects feature points again in the same divided region using the individual thresholds set by the threshold setting unit 1023 for the divided region. Thereafter, the processing from step S12 can be performed on another divided region without the threshold setting unit 1023 setting the individual thresholds.
[0050] Returning to Figures 1 and 2, the explanation continues. 3, the feature point matching unit (feature point matching unit 103) matches the feature points detected in the multiple captured images 10 between the multiple captured images 10 (S3). At this time, the feature point matching unit 103 matches the feature points detected as the same part in the two captured images 10 output from the equalized feature point detection unit 102. Therefore, the feature point matching unit 103 can match the feature points detected in the two captured images 10 before division, as in the conventional case.
[0051] When two captured images 10 are divided to generate divided areas, the feature point matching unit (feature point matching unit 103) performs matching by associating feature points for each divided area that is located at the same position in the multiple captured images 10. Therefore, compared to conventional processing that associates and matches feature points across the entire captured image 10, the feature point matching unit 103 can associate and match feature points that are uniformly detected across the entire captured image 10.
[0052] The feature point matching process is as described above with reference to Figures 9 and 10. For example, the feature point matching unit 103 utilizes the divided areas divided by the area dividing unit 1021, and makes changes so that one feature point in a divided area in one captured image 10 is compared with a feature point in a divided area corresponding to the other captured image 10. This process limits the range of the divided areas to be matched, so the feature point matching unit 103 can avoid erroneous matching with feature points detected in other divided areas, and the time required for the matching operation can be shortened.
[0053] The relative parameter calculation unit (relative parameter calculation unit 104) calculates relative parameters based on the feature points associated between the multiple captured images 10 (S4). Then, the relative parameter calculation unit 104 calculates relative parameters using the camera parameters and the matched feature points. The relative parameters are parameters that indicate the relative position and orientation relationship of one camera as seen from the other camera before and after the movement of two vehicle-mounted cameras 101 or one moving vehicle-mounted camera 101. When calculating the relative parameters, the relative parameter calculation unit 104 performs calculations to minimize the sum of reprojection errors of all feature points matched in the two captured images 10.
[0054] Here, when the on-board camera (on-board camera 101) is a stereo camera, the relative parameter calculation unit (relative parameter calculation unit 104) uses the axial position and axial angle of one of the two cameras constituting the stereo camera when it captures a captured image 10 as a reference, and uses the relative axial position and axial angle of the other camera when it captures another captured image 10 as a relative parameter. For example, using the position of the right camera as a reference, the relative three-dimensional position ((x, y, z) axis position) and relative orientation ((x, y, z) axis angles) of the left camera with respect to the reference are used as relative parameters.
[0055] If the on-board camera 101 is a monocular camera, the captured images 10 taken by the single moving on-board camera 101 before and after the movement are used. However, the posture of a single camera is unlikely to remain constant before and after the movement. For example, under ideal conditions where the vehicle is traveling straight on a straight line, the posture of the on-board camera 101 remains constant. However, in reality, the posture of the on-board camera 101 is expected to change because the angle of the on-board camera 101 changes left and right with slight steering wheel operations, and the angle of the on-board camera 101 changes up and down when the brakes or accelerator are applied.
[0056] For this reason, when the on-board camera (on-board camera 101) is a monocular camera, the relative parameter calculation unit (relative parameter calculation unit 104) uses the axial position and axial angle of the monocular camera when one of multiple captured images 10 taken by the monocular camera as the vehicle moves as a reference, and sets the relative axial positions and axial angles of the monocular camera when other captured images 10 are taken as relative parameters. For example, using the position of the monocular camera taken at a certain moment as a reference, the relative parameters are the three-dimensional position ((x, y, z) axis position) and relative orientation ((x, y, z) axis angle) of the monocular camera taken after time t relative to the reference. As a result, the relative parameters are calculated correctly even when a monocular camera is used as on-board camera 101.
[0057] Next, the parallax image generating unit (parallax image generating unit 105) generates parallax images of the multiple captured images 10 based on the relative parameters (S5). Then, the parallax image generating unit 105 creates a parallelized image by parallelizing the two captured images 10 based on the relative parameters. In the parallelized image, the epipolar line is horizontal and has the same Y coordinate. Then, the parallax image generating unit 105 performs stereo matching on the parallelized image to generate a parallax image.
[0058] Next, the three-dimensional distance recognition unit (three-dimensional distance recognition unit 106) recognizes the three-dimensional distance of the external world based on the parallax image (S6). For example, the three-dimensional distance recognition unit 106 calculates the three-dimensional distance of the external world based on the parallax value of the parallax image generated by the parallax image generation unit 105, the focal length of the camera included in the camera parameters, the pixel size of the light receiving element configured in the vehicle-mounted camera 101, and the like, and recognizes objects in the external world. Thereafter, the processing of the image processing device 100 ends.
[0059] As described above, even if the on-board camera 101 is configured with a single monocular camera, the 3D distance recognition unit 106 can calculate the 3D distance in the external world and recognize objects in the external world. In this case, it is necessary for the on-board camera 101 to capture images of the external world before and after the vehicle moves in a chronological order, and to accurately know how the on-board camera 101 moved between the two captured images 10. If the on-board camera 101 is configured with two stereo cameras, the camera positions of the two captured images 10 taken by the stereo cameras are fixed and accurately known. On the other hand, since the camera position before and after the movement of a single monocular camera that moves with the vehicle is likely to have large errors, the recognition accuracy of the 3D distance in the external world is lower than when a stereo camera is used, but there is an advantage in that the 3D distance can be recognized with a single camera.
[0060] Up to this point, we have explained the functional units configured in the image processing device 100. Next, we will explain the effect when the technology according to this embodiment is applied and the equalized feature point detection unit 102 equalizes feature points. Here, we will explain an example of a rectified image generated by the parallax image generation unit 105 after the relative parameter calculation unit 104 calculates relative parameters, with reference to FIGS. 11 and 12.
[0061] 11 is a diagram showing an example of a rectified image in which the distribution of feature points is biased, to which a conventional technique is applied. The horizontal axis of FIG. 11 represents the X axis of the captured image 10, and the vertical axis represents the Y axis of the captured image 10.
[0062] The black circles and dashed lines shown in the figure indicate the feature points detected from the captured image 10 (base image) taken by one camera and the tilt of the rectified image. The white circles and solid lines shown in the figure indicate the feature points detected from the captured image 10 (reference image) taken by the other camera and the tilt of the rectified image. Ideally, it is expected that the two rectified images will have the same tilt. Here, the two captured images 10 are divided into 2-row x 2-column regions 21 to 24.
[0063] In the example of a rectified image shown in Fig. 11, when calculating relative parameters to minimize the reprojection error, weights are calculated for feature points. A weight is a value that is accumulated for the error of a feature point detected from a two-dimensional captured image 10 (errors indicated by white and black circles in the figure). In an area with many feature points, even if the error of one feature point is small, the errors of many feature points are accumulated, so the weight for the evaluation value of the feature point becomes large. On the other hand, in an area with few feature points, only the errors of a small number of feature points are accumulated, so even if the error of one feature point is large, the weight for the evaluation value of the feature point becomes small.
[0064] Considering the relationship between feature points and weights, the weight in the relative parameter calculation for region 22 where feature points are densely located as shown in Fig. 11 is large, and the weight for region 23 where feature points are sparsely located is small. For this reason, in the relative parameter calculation, relative parameters optimized for the region with a large weight and densely located feature points are calculated, resulting in a large deviation in region 23 where feature points are sparsely located. As a result, the captured image 10 (reference image) indicated by the solid line is significantly tilted and not parallelized relative to the captured image 10 (standard image) indicated by the dashed line.
[0065] FIG. 12 shows an example of a rectified image in which the distribution of feature points is equalized by applying the technology according to this embodiment. In FIG. 12, the black circles, white circles, dashed lines and solid lines shown in the figure are the same as those shown in FIG. In the example of the parallelized images shown in Fig. 12, feature points are distributed evenly across regions 21 to 24 (entire images) of the two captured images 10. Therefore, the weights in the calculation of the relative parameters are also equalized within the two captured images 10. As a result, calculation results with less distortion can be obtained. Furthermore, compared to the captured image 10 (standard image) shown by the dashed line, the captured image 10 (reference image) shown by the solid line has reduced tilt and is parallelized.
[0066] Next, the hardware configuration of the computer 50 that constitutes the image processing device 100 will be described. Fig. 13 is a block diagram showing an example of the hardware configuration of the calculator 50. The calculator 50 is an example of hardware used as a computer that can operate as the image processing device 100 according to this embodiment. The image processing device 100 according to this embodiment realizes an image processing method in which the functional blocks shown in Figs. 1 and 3 cooperate with each other by causing the calculator 50 (computer) to execute a program.
[0067] The computer 50 includes a CPU (Central Processing Unit) 51, a ROM (Read Only Memory) 52, and a RAM (Random Access Memory) 53, each connected to a bus 54. The computer 50 further includes a non-volatile storage 55 and a network interface 56.
[0068] The CPU 51 reads out program code of software that realizes each function according to this embodiment from the ROM 52, loads it into the RAM 53, and executes it. Variables, parameters, etc. that are generated during the calculation processing of the CPU 51 are temporarily written to the RAM 53, and these variables, parameters, etc. are read out as appropriate by the CPU 51. However, an MPU (Micro Processing Unit) may be used instead of the CPU 51.
[0069] The nonvolatile storage 55 may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flexible disk, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a magnetic tape, or a nonvolatile memory. In addition to an operating system (OS) and various parameters, programs for operating the computer 50 are recorded in the nonvolatile storage 55. The ROM 52 and the nonvolatile storage 55 record programs, data, and the like required for the CPU 51 to operate, and are used as an example of a computer-readable non-transitory storage medium that stores programs executed by the computer 50.
[0070] The network interface 56 may be, for example, a NIC (Network Interface Card), and various data can be transmitted and received between devices via a LAN (Local Area Network), dedicated line, etc. connected to the terminal of the NIC.
[0071] In the image processing device 100 according to the embodiment described above, when detecting feature points from the captured image 10, the detected feature points are operated so as to have a distribution suitable for calibration within the captured image 10, thereby making it possible to improve the accuracy of calibration (calculation of relative parameters). As a result, the three-dimensional distance recognition unit 106 included in the image processing device 100 can recognize three-dimensional positions expressed by three-dimensional distances in the external world more accurately than ever before.
[0072] Here, the two captured images 10 used for matching are divided into the same size and number of divisions by the region dividing unit 1021. The captured image 10 is divided into divided regions so that at least one feature point is detected. Therefore, by distributing feature points evenly in each divided region, the feature points are distributed evenly across the entire captured image 10. Furthermore, it becomes easier to match feature points between divided regions at the same position in each captured image 10.
[0073] Furthermore, the size of the divided regions is changed according to the number of feature points predicted for each part of the captured image 10. The captured image 10 is divided so that the size of parts predicted to have many feature points is small, and the size of parts predicted to have few feature points is large. As a result, at least one feature point is detected even in parts predicted to have few feature points.
[0074] Furthermore, the individual thresholds for detecting feature points can be changed for each divided region. Therefore, by setting the individual thresholds high for parts where the number of feature points is predicted to be large, the number of detected feature points can be reduced. Conversely, by setting the individual thresholds low for parts where the number of feature points is predicted to be small, the number of detected feature points can be increased. As a result, the number of detected feature points for each divided region is equalized.
[0075] Furthermore, even captured images 10 that would have been discarded in the past due to uneven placement of feature points can be used for calibration by the image processing device 100. This increases the opportunities for calibration in the image processing device 100, allowing for more frequent and accurate calibration than before.
[0076] [Variations] If one of the cameras constituting the stereo camera fails, the other camera (for example, the right camera) may be used to perform calibration using the captured images 10 taken before and after the vehicle moves. In this case, one of the cameras can be calibrated using the same process as when a monocular camera is used for the in-vehicle camera 101.
[0077] Alternatively, two captured images 10 captured by a stereo camera may be acquired at different times, and the same divided areas of the two captured images 10 captured at different times may be matched. For example, the equalized feature point detection unit 102 detects feature points from the divided areas into which the two captured images 10 acquired at the first time are divided, and then sets an individual threshold. Then, the equalized feature point detection unit 102 detects feature points for the divided areas into which the two captured images 10 acquired at the next time are divided, using the individual threshold set at the first time. By using the individual threshold set at the first time to detect feature points next for the captured images 10 captured twice in a short period of time in this way, the three-dimensional distance recognition unit 106 can more accurately recognize three-dimensional distances in the external world.
[0078] The present invention is not limited to the above-described embodiment, and it goes without saying that various other applications and modifications are possible without departing from the gist of the present invention as set forth in the claims. For example, the above-described embodiment has described the system configuration in detail and specifically to clearly explain the present invention, and is not necessarily limited to a system including all of the described configurations. Furthermore, it is also possible to add, delete, or replace part of the configuration of the present embodiment with other configurations. In addition, the control lines and information lines shown are those that are considered necessary for the explanation, and do not necessarily show all the control lines and information lines in the product. In reality, it can be assumed that almost all components are interconnected. [Explanation of symbols]
[0079] 10...Captured image, 100...Image processing device, 101...In-vehicle camera, 102...Equalized feature point detection unit, 103...Feature point matching unit, 104...Relative parameter calculation unit, 105...Disparity image generation unit, 106...3D distance recognition unit, 1021...Area division unit, 1022...Feature point detection unit, 1023...Threshold setting unit
Claims
1. An area dividing unit that changes the positions and sizes of areas into which a plurality of images captured by an on-board camera are divided according to the mounting position of the on-board camera, and divides each of the plurality of captured images into a plurality of divided areas; a feature point detection unit that detects a portion exceeding a threshold value for each divided region as a feature point; a threshold setting unit that sets the threshold for each divided region so that the distribution of the feature points detected for each divided region is uniform; a feature point matching unit that performs matching by associating the feature points detected in the plurality of captured images among the plurality of captured images; and a relative parameter calculation unit that, when the on-board camera is a stereo camera, uses an axial position and an axial angle of one of the two cameras constituting the stereo camera when the one camera captured the captured image as a reference, and uses a relative axial position and an axial angle of the other camera when the other camera captured another captured image as a relative parameter indicating a relationship between the relative positions and attitudes of the two cameras, and calculates the relative parameter based on the feature points associated between the multiple captured images. Image processing device.
2. An area dividing unit that changes the position and size of areas into which multiple images captured by the on-board camera are divided according to the mounting position of the on-board camera, and divides each of the multiple captured images into multiple divided areas; a feature point detection unit that detects a portion exceeding a threshold value for each divided region as a feature point; a threshold setting unit that sets the threshold for each divided region so that the distribution of the feature points detected for each divided region is uniform; a feature point matching unit that performs matching by associating the feature points detected in the plurality of captured images among the plurality of captured images; and a relative parameter calculation unit that, when the on-board camera is a monocular camera, calculates, as a reference, an axial position and an axial angle of the monocular camera when one of the plurality of captured images is captured by the monocular camera as the vehicle on which the monocular camera is mounted moves, and calculates, as relative parameters, relative axial positions and axial angles of the monocular camera when another of the captured images is captured, based on the feature points associated between the plurality of captured images. Image processing device.
3. The threshold setting unit increases the threshold set for the divided area in which the number of feature points is greater than a target value, and decreases the threshold set for the divided area in which the number of feature points is less than the target value. The feature point detection unit detects the feature points using the threshold value set by the threshold value setting unit.
3. The image processing device according to claim 1 or 2.
4. The area dividing unit divides the captured image by increasing the number of divisions and reducing the size of the divided areas so that at least one feature point is detected for each divided area. The image processing device according to claim 3 .
5. The region dividing unit divides the captured image in a portion where it is predicted that there are many feature points into a size smaller than that of a portion where it is predicted that there are few feature points.
3. The image processing device according to claim 1 or 2.
6. The region dividing unit divides the plurality of captured images into regions of different sizes based on prediction results of feature points corresponding to three-dimensional positions of the outside world detected from the plurality of captured images captured by the on-board camera as the vehicle moves.
3. The image processing device according to claim 1 or 2.
7. The feature point matching unit performs matching by associating the feature points for each of the divided areas that are located at the same position in the plurality of captured images. The image processing device according to claim 5 .
8. a parallax image generating unit that generates parallax images of the plurality of captured images based on the relative parameters; a three-dimensional distance recognition unit that recognizes three-dimensional distances in the outside world based on the parallax images; 3. The image processing device according to claim 1 or 2.
9. A process of dividing a plurality of captured images into a plurality of divided areas by changing the position and size of the areas into which the captured images are divided according to the mounting position of the in-vehicle camera; a process of detecting a portion exceeding a threshold value for each divided region as a feature point; a process of setting the threshold value for each divided region so that the distribution of the feature points detected for each divided region is uniform; a process of matching the feature points detected in the plurality of captured images by associating them among the plurality of captured images; and if the on-board camera is a stereo camera, a process of setting an axial position and an axial angle of one of the two cameras constituting the stereo camera when the one camera captured the captured image as a reference, and a relative axial position and an axial angle of the other camera when the other camera captured another captured image as relative parameters indicating a relationship between the relative positions and attitudes of the two cameras, and calculating the relative parameters based on the feature points associated among the multiple captured images. Image processing methods.
10. A process of dividing a plurality of captured images into a plurality of divided areas by changing the positions and sizes of the areas into which the captured images are divided according to the mounting position of the on-board camera; a process of detecting a portion exceeding a threshold value for each divided region as a feature point; a process of setting the threshold value for each divided region so that the distribution of the feature points detected for each divided region is uniform; a process of matching the feature points detected in the plurality of captured images by associating them among the plurality of captured images; and if the on-board camera is a monocular camera, a relative parameter calculation process is performed in which, among a plurality of images taken by the monocular camera as the vehicle on which the monocular camera is mounted moves, an axial position and an axial angle of the monocular camera when one of the images is taken is used as a reference, and the relative axial positions and axial angles of the monocular camera when other images are taken are used as relative parameters, and the relative parameters are calculated based on the feature points associated between the plurality of images. Image processing methods.
Citation Information
Patent Citations
Parameter calculating apparatus, parameter calculating system and program
JP2009276233A
Image processing device, image processing method, and program
JP2012234258A
Stereo image generating apparatus, stereo image generating method, and computer program for stereo image generation
JP2013114505A
Stereo image generation device, stereo image generation method and computer program for stereo image generation
JP2013123123A
Moving image processor, moving image processing method and program for moving image processing
JP2013186816A