Image processing apparatus and image processing method

The image processing apparatus addresses inconsistencies in texture-based and disparity analysis-based detection systems by using multiple detection units and analysis units to enhance object detection accuracy in autonomous driving environments.

JP7871204B2Active Publication Date: 2026-06-08ASTEMO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
ASTEMO LTD
Filing Date
2023-01-13
Publication Date
2026-06-08

AI Technical Summary

Technical Problem

Texture-based detection systems, despite their high performance, often output unintended results due to mismatches between learning images and current driving scenes, leading to inconsistencies when combined with disparity analysis-based systems, resulting in missed detections and ineffective utilization of high detection performance.

Method used

An image processing apparatus comprising a first detection unit for texture-based detection, a second detection unit for disparity analysis-based detection, an inconsistent region extraction unit to identify regions where one unit detects an object while the other does not, and an inconsistent region analysis unit to determine the presence or absence of an object in these regions using a different process.

Benefits of technology

Enables accurate determination of object presence or absence by utilizing the high detection performance of texture-based systems, reducing system malfunctions in autonomous driving by confirming object existence from both texture and 3D structure perspectives.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007871204000004
    Figure 0007871204000004
  • Figure 0007871204000005
    Figure 0007871204000005
  • Figure 0007871204000006
    Figure 0007871204000006
Patent Text Reader

Abstract

To provide an image processing device and an image processing method which are capable of enhancing the accuracy of determining consistency between two types of detection.SOLUTION: An image processing device includes: a first detection unit; a second detection unit; a non-consistency area extraction unit; and a non-consistency area analysis unit. The first detection unit detects an object from an image output by a sensor unit. The second detection unit detects an object from an image output by the sensor unit in a process different from the first detection unit. The non-consistency area extraction unit compares the detection results from the first detection unit and the second detection unit, and extracts a non-consistency area in which the first detection unit detects an object, and the second detection unit does not detect an object. The non-consistency area analysis unit performs detection differently from the first detection unit and the second detection unit, and determines presence / absence of an object in the non-consistency area.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus and an image processing method.

Background Art

[0002] In recent years, with the development of image recognition technologies represented by deep learning, the performance of texture-based methods for detecting objects from images has been significantly improved. However, texture-based methods may output unintended results due to the mismatch between the learning images and the current driving scenes. Therefore, in order to suppress the malfunction of the system, it is necessary to determine the consistency of the texture-based output results.

[0003] In order to determine the consistency of the texture-based output results, a 3D-based detection method is used in combination to confirm the presence of an object from both viewpoints of texture and 3D structure. For example, a stereo camera can acquire information on a luminance image and a disparity image (3D information) with a single sensor. Therefore, by comparing the detection results based on texture and disparity analysis (3D-based), it is possible to determine the consistency by confirming that the same object can be detected by both methods.

[0004] A method for determining the presence of an object based on the outputs of different methods is disclosed in, for example, Patent Document 1. Patent Document 1 describes a sensor fusion system that detects the same detection target by a plurality of types of detection means (sensors, etc.) and comprehensively determines their detection results, and a vehicle control device using the same. The sensor fusion system described in Patent Document 1 takes the product of the probability distributions output from each of the plurality of output means.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

[0006] However, when comparing the output of a texture-based system with high detection performance, such as deep learning, with that of a conventional disparity analysis-based system, there are cases where the texture-based output detects something while the disparity analysis-based output fails to detect it. If one of the two methods fails to detect something, then, as in the sensor fusion system described in Patent Document 1, the result will be that it is not detected when judged by the product of probability distributions. Therefore, the high detection performance of the texture-based system cannot be effectively utilized.

[0007] The objective of this invention is to provide an image processing device and an image processing method that can accurately determine the presence or absence of an object using two types of detection results, taking into consideration the above-mentioned problems. [Means for solving the problem]

[0008] To solve the above problems and achieve the objective, an image processing apparatus according to one aspect of the present invention comprises a first detection unit, a second detection unit, an inconsistent region extraction unit, and an inconsistent region analysis unit. The first detection unit detects an object from an image output by a sensor unit. The second detection unit detects an object from an image output by a sensor unit using a different process than that of the first detection unit. The inconsistent region extraction unit compares the detection results of the first detection unit and the second detection unit and extracts inconsistent regions where the first detection unit detects an object and the second detection unit does not detect an object. The inconsistent region analysis unit performs detection different from that of the first and second detection units to determine the presence or absence of an object in the inconsistent region. [Effects of the Invention]

[0009] According to one aspect of the present invention, the presence or absence of an object can be determined with high accuracy using two types of detection results. Furthermore, issues, configurations, and effects other than those mentioned above will be clarified by the following description of the embodiments. [Brief explanation of the drawing]

[0010] [Figure 1]This is a schematic diagram showing the configuration of a vehicle equipped with an image processing device according to the first embodiment. [Figure 2] This is a diagram illustrating the principle of triangulation. [Figure 3] This diagram illustrates the process of identifying the same distance-measuring target from images captured by the left and right cameras according to the first embodiment. [Figure 4] This diagram illustrates the detection of an object by a texture-based detection unit according to the first embodiment. [Figure 5] This diagram illustrates the detection of an object by the disparity analysis-based detection unit according to the first embodiment. [Figure 6] This figure shows the factors that cause the object to fail to be detected by the disparity analysis-based detection unit according to the first embodiment. [Figure 7] This figure illustrates the processing of the inconsistent region extraction unit according to the first embodiment. [Figure 8] This is a flowchart showing the inconsistent region extraction process according to the first embodiment. [Figure 9] This figure illustrates the processing of the inconsistent region analysis unit according to the first embodiment. [Figure 10] This figure illustrates another example of size selection for the matching window according to the first embodiment. [Figure 11] This is a flowchart showing the inconsistent region analysis process according to the first embodiment. [Figure 12] This is a flowchart showing the inconsistent region analysis process according to the second embodiment. [Modes for carrying out the invention]

[0011] 1. First Embodiment The first embodiment of the image processing apparatus and image processing method will be described below with reference to Figures 1 to 11. Common parts in each figure are denoted by the same reference numerals.

[0012] [Example of vehicle configuration] First, the configuration of a vehicle equipped with the image processing device according to the first embodiment will be described with reference to Figure 1. FIG. 1 is a schematic configuration diagram showing the configuration of a vehicle equipped with an image processing apparatus.

[0013] As shown in FIG. 1, a vehicle (automobile) 1 is equipped with an in-vehicle camera 100, an image processing apparatus 200, and a vehicle control apparatus 300. The in-vehicle camera 100 shows a specific example of the sensor unit according to the present invention.

[0014] The in-vehicle camera 100 is composed of a plurality of imaging devices arranged at a predetermined interval. The plurality of imaging devices are installed in a predetermined direction around the vehicle 1, for example, facing the front of the vehicle 1. Hereinafter, the in-vehicle camera 100 will be described as a stereo camera system composed of two cameras, a left camera 110 and a right camera 120.

[0015] Note that the in-vehicle camera may be a camera system composed of three or more cameras, or may be configured by combining a plurality of single cameras. Further, the in-vehicle camera may be configured by combining sensors that output a luminance image and a disparity image, respectively.

[0016] The left camera 110 and the right camera 120 of the in-vehicle camera 100 image the front of the vehicle 1. The left camera 110 and the right camera 120 are arranged side by side in the left-right direction orthogonal to the front-rear direction of the vehicle 1. Using the images captured by the left camera 110 and the right camera 120, the distance to the target can be measured based on the principle of triangulation. The in-vehicle camera 100 outputs a luminance image and a disparity image. The images (luminance image and disparity image) captured by the in-vehicle camera 100 are supplied to the image processing apparatus 200.

[0017] The image processing apparatus 200 performs image correction processing such as camera calibration and distortion correction on the captured image. Then, objects such as objects and people are recognized from the image on which the image correction processing has been performed. The image processing apparatus 200 includes a texture-based detection unit 210, a disparity analysis-based detection unit 220, a mismatch region extraction unit 230, and a mismatch region analysis unit 240.

[0018] The object recognition result from the image processing device 200 is sent to the vehicle control device 300. The vehicle control device 300 acquires the object recognition result as information necessary for autonomous driving and controls the driving state of vehicle 1. For example, the vehicle control device 300 controls braking to avoid contact with objects or people in front of vehicle 1, driving at a speed that follows another vehicle in front of vehicle 1, and steering to maintain the lane of the road while driving.

[0019] [Principle of triangulation] Next, the principle of triangulation, which is used to measure the distance to an object, will be explained with reference to Figures 2 and 3. Figure 2 is a diagram illustrating the principle of triangulation. Figure 3 is a diagram illustrating the process of identifying the same distance-measuring object from images captured by left and right cameras 110 and 120.

[0020] The parallax d, which is the difference in the position of the same distance-measuring object in the images captured by the left camera 110 and the right camera 120 that constitute the stereo camera, is calculated by equation (1). In this case, the position in which the distance-measuring object is captured by the left camera 110 is X. L The position where the distance measurement target is captured by the right camera 120 is X R Let's assume that.

number

[0021] The distance Z to the object to be measured is calculated by equation (2). In this equation, B is the baseline length and f is the focal length.

number

[0022] In order to calculate the parallax d using the above equation (1), it is necessary to associate the areas in the left image captured by the left camera 110 and the right image captured by the right camera 120 where the same distance measurement target is captured. As shown in Figure 3, the association of areas where the same distance measurement target is captured is performed by setting a matching window W of a predetermined size in the left image and identifying the area with the same texture as the area within the matching window from the right image.

[0023] Specifically, W(x,y) is the pixel value of the matching window in the left image, and I is the pixel value of the right image. R Let (x,y) be the coordinates, and calculate the SAD (Sum of Absolute Difference) using equation (3) while searching the right image. At this time, the scan positions are i and j. Also, let w be the horizontal length of the matching window W, and h be the vertical length of the matching window W. Note that the vertical direction is perpendicular to the horizontal and front-to-back directions.

number

[0024] Then, by finding the region where the SAD value is minimized, identical regions can be identified. Note that identifying identical regions is not limited to using SAD; any other method such as SSD (Sum of Squared Difference) or NCC (Normalized Cross Correlation) may also be used.

[0025] [Texture-based detection unit] Next, the texture-based detection unit 210 will be described with reference to Figure 4. Figure 4 illustrates the detection of an object by the texture-based detection unit 210.

[0026] The texture-based detection unit 210 detects objects (other vehicles, pedestrians, motorcycles, and other obstacles) from the brightness image captured by the in-vehicle camera 100. As shown in Figure 4, the texture-based detection unit 210 inputs the image acquired by the in-vehicle camera 100 into a neural network and estimates a rectangular frame (top-left coordinate, width, height) that encloses the area to be detected.

[0027] Furthermore, the texture-based detection unit 210 may simultaneously estimate the type, distance, and movement speed of the object to be detected, in addition to the rectangular frame surrounding the area to be detected. Also, the detection of objects by the texture-based detection unit 210 is not limited to using a neural network, but any method capable of detecting objects may be used.

[0028] As shown in Figure 4, the output of the texture-based detection unit 210 includes not only correct detections where the target was detected correctly, but also false detections where the target was mistakenly estimated to be a target. False detections can occur due to differences between the training image and the current scene. False detections can also occur when a texture similar to the target learned for detection is coincidentally recognized. False detections must be eliminated because they indicate an inability to correctly recognize the surrounding environment and can lead to system malfunctions.

[0029] [Disparity analysis-based detection unit] Next, the disparity analysis-based detection unit 220 will be explained with reference to Figure 5. Figure 5 illustrates the detection of an object by the disparity analysis-based detection unit 220.

[0030] The parallax analysis-based detection unit 220 detects objects (other vehicles, pedestrians, motorcycles, and other obstacles) from the parallax image captured by the in-vehicle camera 100. In the driving scene shown in Figure 5A, the parallax image shown in Figure 5B is output from the in-vehicle camera 100 to the parallax analysis-based detection unit 220. Each pixel in the parallax image stores a parallax value. In the parallax image shown in Figure 5B, objects with closer distances to vehicle 1 are displayed in darker colors.

[0031] In the road surface region of the disparity image, the disparity value stored in each pixel changes smoothly as the vertical position of the image changes. On the other hand, in the vehicle region, the same disparity value is clustered together. When a v-disparity map is created from this disparity image, with the horizontal axis d (disparity value) and the vertical axis v (vertical position of the image) as shown in Figure 5C, three-dimensional objects such as vehicles become straight lines in the vertical direction, and the road surface region becomes a straight line sloping downwards to the right. By determining the parameter of this downward-sloping straight line, a road surface model can be estimated.

[0032] Next, by removing the parallax that matches the estimated road surface model, we can eliminate parallax values ​​other than those of vehicles, as shown in Figure 5D. Then, among the parallax values ​​remaining in the parallax image after removing the parallax that matches the road surface model, we set a rectangular frame to enclose nearby identical parallax values. As a result, we can determine the vehicle area, as shown in Figure 5E.

[0033] The disparity analysis-based detection unit 220 may simultaneously estimate the type, distance, and movement speed of the object to be detected when estimating the rectangular frame (top-left coordinate, width, height) of the target area. Furthermore, object detection by the disparity analysis-based detection unit 220 is not limited to using a method for creating a v-disparity map, but may use any method for estimating the area of ​​the object to be detected.

[0034] [Failure to detect by the disparity analysis-based detection unit] Next, the factors that cause the disparity analysis-based detection unit 220 to fail to detect an object will be explained with reference to Figure 6. Figure 6 shows the factors that cause the disparity analysis-based detection unit 220 to fail to detect an object.

[0035] Figure 6A shows a case where the matching window for disparity calculation is large relative to the object being detected. When the matching window is large relative to the object being detected, other objects surrounding the object for which disparity calculation is to be performed may be included within the matching window.

[0036] In Figure 6A, the parallax for a distant vehicle was calculated, but because the matching window is large, surrounding street trees and trucks are included in the same matching window. The parallax calculation method described above (see Figure 3) cannot accurately calculate parallax when objects at different distances are included within the matching window, as shown in Figure 6A.

[0037] If the parallax cannot be calculated accurately, the presence of a vehicle cannot be detected even using the method for creating the v-disparity map described above (see Figure 5). Consequently, the object detection by the parallax analysis-based detection unit 220 will fail.

[0038] Figure 6B shows a case where the matching window is too small for the target to be detected. When searching for the same region as the matching window set in the left image captured by the left camera 110 in the right image captured by the right camera 120, if the matching window is too small, locally identical patterns may appear.

[0039] As shown in Figure 6B, the texture within the matching window set in the left image and the regions a, b, and c in the right image have exactly the same texture. As a result, the disparity analysis-based detection unit 220 cannot calculate the disparity. Consequently, the object detection by the disparity analysis-based detection unit 220 fails.

[0040] Thus, when the disparity analysis-based detection unit 220 fails to detect an object, an inconsistency occurs where the texture-based detection unit 210 detects an object, but the disparity analysis-based detection unit 220 does not detect the corresponding object.

[0041] [Inconsistent area extraction part] Next, the mismatch region extraction unit 230 will be explained with reference to Figure 7. Figure 7 illustrates the processing of the inconsistent region extraction unit.

[0042] The inconsistent region extraction unit 230 receives the output of the texture-based detection unit 210 and the output of the parallax analysis-based detection unit 220. The information of the objects detected by both detection units 210 and 220 is then output directly to the vehicle control device 300. On the other hand, if the parallax analysis-based detection unit 220 does not detect an object that the texture-based detection unit 210 has detected (i.e., no detection), it is possible that the texture-based detection unit 210 has made a false detection, or that the parallax analysis-based detection unit 220 was unable to detect the object (it was not detected).

[0043] If the texture-based detection unit 210 detects an object that the disparity analysis-based detection unit 220 does not detect, the inconsistent region extraction unit 230 extracts the region detected by the texture-based detection unit 210 as an inconsistent region. The inconsistent region extraction unit 230 outputs the extracted inconsistent region to the inconsistent region analysis unit 240.

[0044] Figure 7A shows the detection results of the texture-based detection unit 210. In Figure 7A, detected objects are displayed with dashed rectangular frames. High-precision recognition technologies such as deep learning can accurately detect detection targets (objects) such as vehicles and trucks shown in Figure 7A. On the other hand, due to discrepancies between the training image and the current driving scene, false detections may occur where a rectangular frame is output even though no detection target exists.

[0045] Figure 7B shows the detection results of the parallax analysis-based detection unit 220. In Figure 7B, detected objects are shown with dashed rectangular frames. The parallax analysis-based detection unit 220 can detect three-dimensional objects that are at a different height from the road surface, but it may not be able to detect certain objects (resulting in non-detection). As mentioned above, one reason for non-detection is that the size of the matching window used to calculate parallax is not appropriate for the size of the object on the image, and therefore the parallax cannot be calculated accurately.

[0046] [Inconsistent Region Extraction Process] Next, the inconsistent region extraction process by the inconsistent region extraction unit 230 will be explained with reference to Figure 8. Figure 8 is a flowchart showing the process for extracting inconsistent regions.

[0047] First, the inconsistent region extraction unit 230 acquires the detection result from the texture-based detection unit 210 (hereinafter referred to as the "texture-based detection result") (S101). The inconsistent region extraction unit 230 also acquires the detection result from the disparity analysis-based detection unit 220 (hereinafter referred to as the "disparity analysis-based detection result") (S102).

[0048] Next, the inconsistent region extraction unit 230 calculates IoU (Intersection over Union) for the texture-based detection result and the disparity analysis-based detection result (S103). Subsequently, the inconsistent region extraction unit 230 determines whether the value of IoU is greater than or equal to a predetermined threshold (S104). The processing in steps S103 and S104 is performed for each disparity analysis-based detection result for which the associated flag has not been set.

[0049] If the texture-based detection result and the parallax analysis-based detection result detect the same object, most of the rectangular frames output by the texture-based detection unit 210 and the parallax analysis-based detection unit 220 will overlap. Therefore, if the IoU value is greater than or equal to a predetermined threshold, it can be determined that the texture-based detection result and the parallax analysis-based detection result have detected the same object.

[0050] On the other hand, if there is no combination of texture-based detection results and parallax analysis-based detection results in which the IoU value is equal to or greater than a predetermined threshold, it can be determined that the texture-based detection unit 210 has detected an object, but the parallax analysis-based detection unit 220 has not detected the corresponding object.

[0051] The texture-based detection unit 210 and the disparity analysis-based detection unit 220 may output the type of object, distance, movement speed, etc. In this case, in step S104, it may be determined whether or not they are the same object, including not only the IoU value but also information such as the type of object, distance, and movement speed.

[0052] In step S104, if it is determined that the IoU value is greater than or equal to a predetermined threshold (if S104 is determined to be YES), the inconsistent region extraction unit 230 sets a positive detection flag for the texture-based detection result to be determined (S105). Then, the inconsistent region extraction unit 230 sets a flag that has been associated with the disparity analysis-based detection result to be determined (S106).

[0053] In step S104, if it is determined that the value of IoU is not equal to or greater than a predetermined threshold (if S104 is determined to be NO), the inconsistent region extraction unit 230 sets an inconsistent region flag for the texture-based detection result to be determined (S107). The processing in steps S103 to 107 is performed for each texture-based detection result.

[0054] Subsequently, the inconsistent region extraction unit 230 outputs the determination result (S108). Specifically, the inconsistent region extraction unit 230 outputs the texture-based detection result with the positive detection flag set to the vehicle control device 300. The inconsistent region extraction unit 230 also outputs the texture-based detection result with the inconsistent region flag set to the inconsistent region analysis unit 240. After processing in S108, the inconsistent region extraction unit 230 terminates the inconsistent region extraction process.

[0055] In this embodiment, inconsistency region determination is performed for all texture-based detection results. However, in image processing devices mounted on mobile devices such as those in the automotive field, computational resources are limited, so it may not be possible to perform inconsistency region determination for all texture-based detection results. In such cases, inconsistency region determination may be limited to only high-priority detection results (for example, detection results with a priority higher than a predetermined rank). This reduces the processing performed by the inconsistency region extraction unit 230.

[0056] Priority can be determined based on at least one of the following factors in relation to the texture-based detection results: distance to vehicle 1, relative speed, time to collision, distance to the predicted path of vehicle 1, and past tracking results.

[0057] [Inconsistency area analysis department] Next, the mismatch region analysis unit 240 will be explained with reference to Figure 9. Figure 9 is a diagram illustrating the processing of the inconsistent region analysis unit.

[0058] The inconsistent region analysis unit 240 determines whether the inconsistent region output by the inconsistent region extraction unit 230 is a false detection by the texture-based detection unit 210 or an undetected region by the disparity analysis-based detection unit 220.

[0059] Figure 9A shows the inconsistent region output by the inconsistent region extraction unit 230. The inconsistent region analysis unit 240 obtains the width w, which is the horizontal length of the inconsistent region, and the height h, which is the vertical length of the inconsistent region. The vertical direction is defined as the direction perpendicular to the horizontal and front-to-back directions.

[0060] Figure 9B shows the size of the matching window, which is set based on the width w and height h of the mismatched area. As shown in Figure 9B, the matching window is provided with sizes 1, 2, and 3. The width w2 and height h2 of size 2 are larger than the width w1 and height h1 of size 1. Also, the width w3 and height h3 of size 3 are larger than the width w2 and height h2 of size 2.

[0061] Of sizes 1 to 3, the largest size whose vertical and horizontal dimensions do not exceed the vertical dimensions h and horizontal dimensions w of the mismatched area is selected as the matching window. This allows for the selection of an appropriate size as the matching window according to the size of the mismatched area. Note that the number of pre-prepared matching windows according to the present invention is not limited to three (three types), but any number (types) of sizes can be set.

[0062] Figure 9C shows an example of calculating the disparity of an inconsistent region. The inconsistent region analysis unit 240 uses the set matching window size to search for the same region as the inconsistent region in the right image and calculate the disparity. At this time, the disparity is calculated only for the inconsistent region, rather than for the entire reference left image. This reduces the processing performed by the inconsistent region analysis unit 240.

[0063] Furthermore, when the texture-based detection unit 210 outputs distance information to the object, the inconsistent region analysis unit 240 can obtain distance information to the object in the inconsistent region. Therefore, the inconsistent region analysis unit 240 can predict the position of the object captured in the right image by converting the distance to the object detected by the texture-based detection unit 210 into parallax. The inconsistent region analysis unit 240 then scans the matching window only within a range of ±α from the predicted object position. This limits the search range and reduces the processing performed by the inconsistent region analysis unit 240.

[0064] Figure 9D shows an example of determining whether an inconsistent region is a target for detection or a false positive. The inconsistent region analysis unit 240 determines whether an inconsistent region is a target for detection or a false positive based on the calculated parallax. Specifically, the created parallax image is plotted on a ud voting space with the horizontal axis representing the horizontal direction of the image (u) and the vertical axis representing the parallax (d). In this case, three-dimensional objects with the same parallax continuing vertically are compared with areas where no three-dimensional objects exist (such as a road surface), and a large number of votes are concentrated at a specific parallax d location.

[0065] Therefore, the inconsistent region analysis unit 240 checks for each u-column in the ud voting space whether there are pixels that have received more than a predetermined voting threshold, and counts the number of u-columns that have received more than the voting threshold. Then, if the proportion of u-columns that have received more than the voting threshold out of all u-columns is greater than or equal to a predetermined threshold, the inconsistent region analysis unit 240 determines that an inconsistent region is the target of detection. In other words, the inconsistent region analysis unit 240 determines that the detection result of the texture-based detection unit 210 is correct and that there is a target object.

[0066] On the other hand, if the proportion of all u-columns that have received more than or equal to the voting threshold is less than a predetermined threshold, then the inconsistent region is not detected. Therefore, the inconsistent region analysis unit 240 determines that it is a false detection by the texture-based detection unit 210. Note that the existence of the target 3D object is not limited to being determined by the method described above, but may be determined using any method.

[0067] [Other examples of matching window size selection] Next, another example of selecting the size of the matching window will be explained with reference to Figure 10. Figure 10 illustrates another example of selecting the size of the matching window.

[0068] Figure 10A shows an example where the size of one of the divisions of the mismatched region, which is divided into three parts horizontally and vertically (both horizontally and vertically), is set to the size of the matching window. Note that the number of divisions of the mismatched region is not limited to three; it may be two, four or more, or different numbers. Furthermore, the number of divisions of the mismatched region may differ horizontally and vertically. In addition, upper and lower limits may be set for the sizes of the divisions of the mismatched region.

[0069] Figure 10B shows an example of setting the size of the matching window based on the type of mismatched area and the distance to vehicle 1. If the texture-based detection unit 210 has detected the type of object and estimated the distance to vehicle 1, the mismatched area analysis unit 240 can obtain information on the type of object in the mismatched area and the distance to vehicle 1.

[0070] If the type of object and its distance from vehicle 1 are known, the size of the object captured in the image can be predicted. Therefore, the size of the matching window is predetermined for each combination of the type of object and its distance from vehicle 1, and this table or map is stored in the memory unit. The mismatched area analysis unit 240 refers to the table or map stored in the memory unit and determines the size of the matching window based on the type of object and its distance from vehicle 1.

[0071] [Inconsistent Region Analysis Processing] Next, the inconsistency region analysis process performed by the inconsistency region analysis unit 240 will be explained with reference to Figure 11. Figure 11 is a flowchart showing the inconsistent region analysis process.

[0072] First, the inconsistent region analysis unit 240 acquires information about the inconsistent region from the inconsistent region extraction unit 230 (S201). Next, the inconsistent region analysis unit 240 sets the size of the matching window that scans the inconsistent region (S202). Subsequently, it calculates the disparity of the inconsistent region (S203).

[0073] Next, the inconsistent region analysis unit 240 casts the disparity calculated in step S203 into the ud voting space (S204). Next, the inconsistent region analysis unit 240 determines whether the proportion of pixels in column u that have received more than or equal to the voting threshold is greater than or equal to a predetermined threshold (S205).

[0074] In step S205, if the percentage of pixels in column u that have received more than or equal to the voting threshold is determined to be greater than or equal to a predetermined threshold (if S205 is determined to be YES), the inconsistent region analysis unit 240 determines that the inconsistent region is the target of detection and sets a positive detection flag in the texture-based detection result (S206).

[0075] On the other hand, in step S205, if it is determined that the proportion of pixels in column u that have received more than or equal to the voting threshold is not equal to or equal to a predetermined threshold (if S205 is determined to be NO), the inconsistent region analysis unit 240 determines that the inconsistent region is a false detection and sets a false detection flag in the texture-based detection result (S207). The processing in steps S202 to 207 is performed for each inconsistent region.

[0076] Subsequently, the inconsistent region analysis unit 240 outputs the determination result (S208). That is, the inconsistent region analysis unit 240 outputs the texture-based detection result with the positive detection flag set to the vehicle control device 300 (see Figure 1). After processing in S208, the inconsistent region analysis unit 240 terminates the inconsistent region analysis process.

[0077] As explained above, the inconsistent region extraction unit 230 outputs highly reliable detection results that were detected by both the texture-based detection unit 210 and the disparity analysis-based detection unit 220. The inconsistent region analysis unit 240 analyzes the inconsistent regions that were not detected by the disparity analysis-based detection unit 220 but were detected by the texture-based detection unit 210, and outputs detection results that determine that these regions are targets for detection in three dimensions.

[0078] The output results from the inconsistent region extraction unit 230 and the inconsistent region analysis unit 240 are consistent because they both confirm the existence of the object from the perspectives of both texture and three-dimensional structure. Therefore, the presence or absence of an object can be determined with high accuracy using the two types of detection results. The vehicle control device 300 controls the driving state of vehicle 1 according to these output results. As a result, malfunctions of the vehicle control system can be suppressed.

[0079] Furthermore, the inconsistent region analysis unit 240 determines the presence or absence of an object in the inconsistent region where the texture-based detection unit 210 has detected an object, using a different approach. This makes it possible to effectively utilize the high detection performance of the texture-based detection system to improve the accuracy of determining the presence or absence of an object.

[0080] 2. Second Embodiment Next, a second embodiment of the image processing apparatus and image processing method will be described with reference to Figure 12. Figure 12 is a flowchart showing the inconsistent region analysis process according to the second embodiment.

[0081] The image processing apparatus according to the second embodiment has the same configuration as the image processing apparatus 200 according to the first embodiment. The difference between the image processing apparatus according to the second embodiment and the image processing apparatus 200 is the consistent region analysis processing performed by the inconsistent region analysis unit 240. Therefore, the consistent region analysis processing according to the second embodiment will be described here, and the description of the configuration and processing common to the first embodiment will be omitted.

[0082] The mismatch region analysis unit 240 according to the second embodiment selects a method for calculating parallax with higher accuracy than the parallax image output by the in-vehicle camera 100, and analyzes the mismatch region in detail. Examples of methods for outputting highly accurate disparity include machine learning-based methods such as SGM (Semi Global Matching) and CNN (Convolutional Neural Network). Furthermore, the inconsistent region analysis unit according to the present invention may use any other method that outputs highly accurate disparity.

[0083] Methods for outputting high-precision parallax generally involve more processing (higher processing cost) than the parallax calculation method explained using Figure 3. However, methods for outputting high-precision parallax can calculate the parallax of the region including the detection target more accurately. As a result, the mismatch region analysis unit 240 according to the second embodiment can more accurately confirm the presence of the target.

[0084] [Inconsistent Region Analysis Processing] First, the inconsistent region analysis unit 240 obtains information about the inconsistent region from the inconsistent region extraction unit 230 (S301). Next, the inconsistent region analysis unit 240 calculates highly accurate disparity of the inconsistent region using machine learning-based methods such as SGM or CNN (S302).

[0085] Next, the inconsistent region analysis unit 240 casts the disparity calculated in step S302 into the ud voting space (S303). Next, the inconsistent region analysis unit 240 determines whether the proportion of u-columns containing pixels with votes exceeding a voting threshold is above a predetermined threshold (S304).

[0086] In step S304, if the percentage of pixels in column u that have received more than or equal to the voting threshold is determined to be greater than or equal to a predetermined threshold (if S304 is determined to be YES), the inconsistent region analysis unit 240 determines that the inconsistent region is the target of detection and sets a positive detection flag in the texture-based detection result (S305).

[0087] On the other hand, in step S304, if it is determined that the proportion of pixels in column u that have received more than or equal to the voting threshold is not equal to or equal to a predetermined threshold (if S304 is determined to be NO), the inconsistent region analysis unit 240 determines that the inconsistent region is a false detection and sets a false detection flag in the texture-based detection result (S306). The processing in steps S302 to 306 is performed for each inconsistent region.

[0088] Subsequently, the inconsistent region analysis unit 240 outputs the determination result (S307). That is, the inconsistent region analysis unit 240 outputs the texture-based detection result with the positive detection flag set to the vehicle control device 300 (see Figure 1). After processing in S307, the inconsistent region analysis unit 240 terminates the inconsistent region analysis process.

[0089] 3. Summary (1) The image processing apparatus 200 according to the above embodiment includes a texture-based detection unit 210 (first detection unit), a parallax analysis-based detection unit 220 (second detection unit), an inconsistent region extraction unit 230, and an inconsistent region analysis unit 240. The texture-based detection unit 210 detects objects from the image output by the in-vehicle camera 100 (sensor unit). The parallax analysis-based detection unit 220 detects objects from the image output by the in-vehicle camera 100 using a different process than that of the texture-based detection unit 210. The inconsistent region extraction unit 230 compares the detection results of the texture-based detection unit 210 and the parallax analysis-based detection unit 220, and extracts inconsistent regions where the texture-based detection unit 210 detects objects and the parallax analysis-based detection unit 220 does not detect objects. The inconsistent region analysis unit 240 performs detection different from that of the texture-based detection unit 210 and the parallax analysis-based detection unit 220 to determine the presence or absence of objects in the inconsistent region. This allows for highly accurate determination of the presence or absence of an object using the detection results of the texture-based detection unit 210 and the disparity analysis-based detection unit 220. Furthermore, since the presence or absence of an object is determined using a different process than that of the texture-based detection unit 210 and the disparity analysis-based detection unit 220 for inconsistent regions where the texture-based detection unit 210 has detected an object, the accuracy of object presence or absence determination can be improved. In other words, the high detection performance of the texture-based system can be effectively utilized to improve the accuracy of object presence or absence determination.

[0090] (2) Furthermore, the in-vehicle camera 100 (sensor unit) according to the above embodiment outputs both a luminance image and a parallax image. This allows for the detection of objects using both luminance images and disparity images.

[0091] (3) Furthermore, the texture-based detection unit 210 (first detection unit) according to the above embodiment detects an object from the brightness image output by the in-vehicle camera 100 (sensor unit). This allows for the determination of the presence or absence of an object using high texture-based detection capabilities.

[0092] (4) Furthermore, the parallax analysis-based detection unit 220 (second detection unit) according to the above embodiment detects an object from the parallax image output by the in-vehicle camera 100 (sensor unit). This allows for the detection of objects in three dimensions, and enables the detection of the size of the detected object and the distance to the object.

[0093] (5) Furthermore, when the texture-based detection unit 210 (first detection unit) according to the above embodiment outputs multiple detection results, the inconsistent region extraction unit 230 determines the priority of the multiple detection results according to at least one of the following: the distance to the vehicle 1, the predicted path of the vehicle 1, the time until collision with the vehicle 1, and the tracking information of the object. The inconsistent region extraction unit 230 then extracts inconsistent regions from the detection results whose priority is higher than a predetermined rank. This reduces the processing performed by the inconsistent region extraction unit 230.

[0094] (6) Furthermore, the mismatch region analysis unit 240 according to the first embodiment described above sets a matching window size according to the mismatch region and calculates the parallax, and determines the presence or absence of an object from the calculation result. This allows for the accurate calculation of parallax in mismatched areas, improving the accuracy of determining the presence or absence of an object.

[0095] (7) The inconsistent region analysis unit 240 according to the first embodiment described above selects a matching window from a plurality of pre-prepared matching windows that does not exceed the size of the inconsistent region. Alternatively, the inconsistent region analysis unit 240 sets a matching window that is the size of a division of the inconsistent region. Alternatively, the inconsistent region analysis unit 240 sets a matching window that does not exceed the size of the inconsistent region based on at least one of the type of inconsistent region and the distance to the inconsistent region. This makes it easy to set a matching window sized according to the mismatch area.

[0096] (8) Furthermore, the mismatch region analysis unit 240 according to the second embodiment described above calculates a parallax with higher accuracy than the parallax analysis-based detection unit 220 (second detection unit) in the mismatch region, and determines the presence or absence of an object from the calculation result. This improves the accuracy of determining the presence or absence of an object.

[0097] (9) In the image processing method according to the above embodiment, the texture-based detection unit 210 (first detection unit) detects an object from the image output by the in-vehicle camera 100 (sensor unit). The parallax analysis-based detection unit 220 (second detection unit) detects an object from the image output by the in-vehicle camera 100 using a different process than that of the texture-based detection unit 210. Next, the inconsistent region extraction unit 230 compares the detection results of the texture-based detection unit 210 and the parallax analysis-based detection unit 220, and extracts inconsistent regions where the texture-based detection unit 210 detects an object and the parallax analysis-based detection unit 220 does not detect an object. Then, the inconsistent region analysis unit 240 performs a different detection than that of the texture-based detection unit 210 and the parallax analysis-based detection unit 220 to determine the presence or absence of an object in the inconsistent region. This allows for highly accurate determination of the presence or absence of an object using the detection results of the texture-based detection unit 210 and the disparity analysis-based detection unit 220. Furthermore, since the presence or absence of an object is determined using a different process than that of the texture-based detection unit 210 and the disparity analysis-based detection unit 220 for inconsistent regions where the texture-based detection unit 210 has detected an object, the accuracy of object presence or absence determination can be improved. In other words, the high detection performance of the texture-based system can be effectively utilized to improve the accuracy of object presence or absence determination.

[0098] The present invention is not limited to the embodiments described above and shown in the drawings, and various modifications can be made without departing from the gist of the invention as described in the claims.

[0099] Furthermore, the embodiments described above are explained in detail for the purpose of clearly illustrating the present invention, and are not necessarily limited to those comprising all the described configurations. It is also possible to replace parts of the configuration of one embodiment with those of another embodiment, and to add configurations from other embodiments to the configuration of one embodiment. Additionally, it is possible to add, delete, or replace parts of the configuration of each embodiment with those of other embodiments.

[0100] In this specification, although terms such as "parallel" and "orthogonal" are used, these do not mean only strictly "parallel" and "orthogonal," but may also refer to states that are "approximately parallel" or "approximately orthogonal," which include "parallel" and "orthogonal" and are within a range in which they can perform their functions. [Explanation of Symbols]

[0101] 1…Vehicle, 100…In-vehicle camera (sensor unit), 110…Left camera, 120…Right camera, 200…Image processing device, 210…Texture-based detection unit (first detection unit), 220…Disparity analysis-based detection unit (second detection unit), 230…Inconsistent region extraction unit, 240…Inconsistent region analysis unit, 300…Vehicle control device

Claims

1. A first detection unit detects an object from the image output by the sensor unit, A second detection unit detects an object from the image output by the sensor unit using a different process than that of the first detection unit, An inconsistency region extraction unit compares the detection results of the first detection unit and the second detection unit, and extracts inconsistency regions where the first detection unit detects an object and the second detection unit does not detect an object. The system includes an inconsistent region analysis unit that performs detection different from that of the first detection unit and the second detection unit to determine the presence or absence of an object in the inconsistent region. Image processing device.

2. The sensor unit outputs both a luminance image and a disparity image. The image processing apparatus according to claim 1.

3. The first detection unit is a texture-based detection unit that detects an object from the brightness image output by the sensor unit. The image processing apparatus according to claim 2.

4. The second detection unit is a disparity analysis-based detection unit that detects an object from the disparity image output by the sensor unit. The image processing apparatus according to claim 2.

5. The aforementioned sensor unit is an in-vehicle camera mounted on the vehicle. When the first detection unit outputs multiple detection results, the inconsistent region extraction unit determines the priority of the multiple detection results according to at least one of the following: the distance to the vehicle, the predicted path of the vehicle, the time until collision with the vehicle, and the tracking information of the object, and extracts the inconsistent region from the detection results with a priority higher than a predetermined rank. The image processing apparatus according to claim 1.

6. The aforementioned mismatch region analysis unit sets a matching window size according to the mismatch region, calculates the parallax, and determines the presence or absence of an object from the calculation result. The image processing apparatus according to claim 1.

7. The inconsistency region analysis unit selects a matching window from a plurality of pre-prepared matching windows that does not exceed the size of the inconsistency region, sets a matching window with a size equal to the division of the inconsistency region, or sets a matching window that does not exceed the size of the inconsistency region based on at least one of the type of the inconsistency region and the distance to the inconsistency region. The image processing apparatus according to claim 6.

8. The aforementioned mismatch region analysis unit calculates a parallax with higher accuracy than the second detection unit in the mismatch region and determines the presence or absence of an object from the calculation result. The image processing apparatus according to claim 4.

9. The first detection unit detects the object from the image output by the sensor unit. The second detection unit detects an object from the image output by the sensor unit using a different process than the first detection unit. The inconsistent region extraction unit compares the detection results of the first detection unit and the second detection unit, extracts inconsistent regions where the first detection unit detects an object and the second detection unit does not detect an object. The mismatch region analysis unit performs detections different from those of the first and second detection units to determine the presence or absence of an object in the mismatch region. Image processing methods.