Image processing device, image processing method, and obstacle detection system

The image processing apparatus accurately estimates three-dimensional coordinates from two-dimensional images by utilizing parallax information, effectively addressing the challenge of obstacle detection in rail transit systems.

JP2025086696APending Publication Date: 2025-06-09HITACHI LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023200888
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-09

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately estimate three-dimensional coordinates of objects in images captured by cameras, especially when the object's feet are not visible, leading to incorrect obstacle detection in rail transit systems.

Method used

An image processing apparatus that calculates three-dimensional coordinates by acquiring parallax information, detecting objects, extracting a representative value of parallax information at locations with low variation, and converting two-dimensional coordinates into three-dimensional coordinates using this representative value.

Benefits of technology

Enables accurate estimation of three-dimensional coordinates from two-dimensional coordinates regardless of the object's position, improving obstacle detection accuracy and safety in rail transit systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide an image processing device capable of accurately estimating three-dimensional coordinates from two-dimensional coordinates regardless of the position of an object within the captured image or the like.SOLUTION: An image processing device 110 for calculating the three-dimensional coordinates of an object in images captured by a plurality of cameras for capturing the area in front of a train 100 comprises: a disparity information acquisition unit 111 which acquires disparity information on the basis of the images; an object detection unit 112 which detects an object using the images and sets a first region where the object has been detected; a disparity information extraction unit 113 which extracts a representative value of the disparity information at a location where the disparity is small in the first region; and a coordinate conversion unit 114 which converts the two-dimensional coordinates of the first region into three-dimensional coordinates on the basis of the representative value extracted by the disparity information extraction unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and an obstacle detection system. In particular, the present invention relates to an image processing apparatus and the like suitable for calculating the three-dimensional coordinates of an object in an image captured by a camera mounted on a moving body.

Background Art

[0002] In recent years, due to concerns about a shortage of personnel associated with the aging of drivers and the reduction of operation costs, research has been conducted on automatic driving in existing rail transit systems. In a rail transit system in which a transport vehicle travels on a track, when there is an obstacle on the track, it is impossible to avoid it by steering. Therefore, detecting obstacles on the track is important for improving the safety and operability of the rail transit system. Currently, the driver visually detects obstacles on the track and on the route. On the other hand, for autonomous driving, a mechanism for automatically detecting obstacles on the route is required, and methods using external sensors such as millimeter-wave radars, lidar, and cameras have been studied.

[0003] Patent Document 1 describes that an object detection unit predicts the position of a current object on the road surface based on the position of the object on the road surface obtained from a past captured image and vehicle motion information, and sets a detection frame at the corresponding position on the current captured image to detect the object. And it is described that in the first position estimation unit, when it is determined that the lower region of the detection frame is within the shooting range by the detection frame range determination unit, the current position of the object is estimated based on the foot position of the image within the detection frame. Further, it is described that in the second position estimation unit, when it is determined by the detection size determination unit that the size of the detection frame is equal to or larger than a predetermined size, the current position of the object on the road surface is estimated based on the magnification rate of the object in the current captured image with respect to the object in the past captured image.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] In order to detect an obstacle on an orbit using an image captured by a camera, it is necessary to detect an object in the image and determine whether the three-dimensional coordinates of the detected object exist around the orbit. In this case, it is necessary to estimate the three-dimensional coordinates from the two-dimensional coordinates of the object in the image. However, conventionally, on the premise that the feet of the object to be detected are visible, if the feet are not visible, the three-dimensional coordinates may not be correctly estimated. An object of the present invention is to provide an image processing apparatus and an image processing method capable of accurately estimating three-dimensional coordinates from two-dimensional coordinates regardless of the position of an object in a captured image. Another object of the present invention is to provide an obstacle detection system capable of accurately detecting an obstacle.

Means for Solving the Problems

[0006] In order to solve the above problems, the present invention is an image processing apparatus that calculates the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body, and includes a parallax information acquisition unit that acquires parallax information based on the image, an object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image, a parallax information extraction unit that extracts a representative value of the parallax information at a location where the parallax information is small within the first region, and a coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the parallax information extraction unit. In this case, an image processing apparatus capable of accurately estimating three-dimensional coordinates from two-dimensional coordinates regardless of the position of an object in a captured image can be provided.

[0007] Here, for example, the parallax information extraction unit sets a plurality of second regions by dividing the first region according to the variation of the parallax information within the first region, and extracts a representative value of the parallax information at a location where the parallax information becomes small among the plurality of second regions. In this case, the location where the object actually exists within the first region can be specified. Also, for example, the second region is set by obtaining the number of divisions in the vertical and horizontal directions of the image based on the size of the first region and dividing the first region according to the number of divisions. In this case, the sizes of the regions included in the second region can be made approximately the same. Furthermore, for example, the number of divisions is obtained such that the number of pixels included in the second region is close to a predetermined number of pixels. In this case, the sizes of the regions included in the second region can be made approximately the same. Moreover, for example, the moving body is a train, and the object detection unit detects an object around the track on which the train runs. In this case, an object that poses an obstacle to the train can be detected. And, for example, the object detection unit detects at least one of a traffic signal and a security guard as an object around the track. In this case, an object with a high necessity for detection can be detected. Also, for example, the image is an image captured by a stereo camera incorporating a plurality of cameras. In this case, even a distant object can be detected. Furthermore, for example, the representative value is at least one of the mode value, average value, median value, and k - mean of the parallax information. In this case, a value suitable as the representative value can be calculated. Moreover, for example, the moving body is an automobile, and the object detection unit detects an object that may cause a collision. In this case, an object that poses an obstacle to the automobile can be detected.

[0008] The present invention is also an image processing method for calculating the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body. By a processor executing a program recorded in a memory, parallax information is obtained based on the image, an object is detected using the image, a first region is set as a location where the object is detected in the image, a representative value of the parallax information at a location where the parallax information is small within the first region is extracted, and based on the extracted representative value, the two-dimensional coordinates of the first region are converted into three-dimensional coordinates. In this case, it is possible to provide an image processing method capable of accurately estimating three-dimensional coordinates from two-dimensional coordinates regardless of the position of the object within the captured image.

[0009] Here, for example, the moving body is a train, and an object around the track on which the train travels is detected. In this case, an object that poses an obstacle to the train can be detected. Also, for example, at least one of a signal and a security guard is detected as an object around the track. In this case, an object that is highly necessary to be detected can be detected. Furthermore, for example, the image is an image captured by a stereo camera having a plurality of built-in cameras. In this case, even a distant object can be detected. Still further, for example, the representative value is at least one of the mode value, average value, median value, and k-mean of the parallax information. In this case, a value suitable as the representative value can be calculated.

[0010] Furthermore, the present invention is an obstacle detection system for detecting obstacles, comprising an image processing device that calculates the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body, and a detection device that detects an obstacle according to the three-dimensional coordinates of the object calculated by the image processing device. The image processing device includes a parallax information acquisition unit that acquires parallax information based on the image, an object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image, a parallax information extraction unit that extracts a representative value of the parallax information at a location where the parallax information is small within the first region, and a coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the parallax information extraction unit. In this case, it is possible to provide an obstacle detection system that can accurately detect obstacles.

Advantages of the Invention

[0011] According to the present invention, it is possible to provide an image processing device and an image processing method that can accurately estimate three-dimensional coordinates from two-dimensional coordinates regardless of the position of an object in a captured image. In addition, it is possible to provide an obstacle detection system that can accurately detect obstacles.

Brief Description of the Drawings

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Embodiments for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. <Overall Description of Train Control System 1> FIG. 1 is a block diagram showing the overall configuration of the train control system 1 of the present embodiment. The train control system 1 according to the present embodiment is mounted on a train 100 running on a track. The train 100 is an example of a moving body. The train 100 is not particularly limited, and may be composed of a normal railway vehicle using rails as the track, or may be composed of vehicles such as a monorail, a light rail, or a cable car.

[0014] The train control system 1 includes an image processing device 110, an imaging device 120, and an upper control device 130. Among these, the imaging device 120 includes a plurality of cameras that image the front of the train 100. The imaging device 120 is, for example, a stereo camera incorporating a plurality of cameras. When the train control system 1 of the present embodiment is mounted on the train 100, it is preferable that objects at a longer distance can be detected. By using the imaging device 120 as a stereo camera, it becomes possible to detect objects at a longer distance. In the present embodiment, by arranging a plurality of cameras at a predetermined distance apart, the same object can be imaged from different directions.

[0015] The upper control device 130 controls the train 100 based on the result of image processing by the image processing device 110. The upper control device 130 is, for example, a driver's cab monitor device or a transmission device. The image processing device 110 is connected to the imaging device 120 and calculates the three-dimensional coordinates of an object in the image captured by the imaging device 120. The image processing device 110 is a computer device. Therefore, the image processing device 110 includes a processor such as a CPU (Central Processing Unit) which is an arithmetic means, and a main memory which is a storage means. Here, the processor executes various software such as an OS (basic software) and an app (application software). Also, the main memory is a storage area for storing various software and data used for its execution. Further, the image processing device 110 includes, as an auxiliary storage device, a storage such as an HDD (Hard Disk Drive) or an SSD (Solid State Drive), and a communication interface for communicating with the outside.

[0016] The image processing device 110 includes a parallax information acquisition unit 111, an object detection unit 112, a parallax information extraction unit 113, and a coordinate conversion unit 114. The parallax information acquisition unit 111 acquires the image captured by the imaging device 120. Then, the parallax information acquisition unit 111 acquires parallax information based on the images captured using a plurality of cameras of the imaging device 120. And the acquired parallax information is transmitted to the parallax information extraction unit 113. At this time, the parallax information is what represents the depth recorded for the image by stereo matching using at least two cameras as a parallax value. Stereo matching is a technique for estimating the depth in an image using two images of the same static scene taken from different viewpoints.

[0017] The object detection unit 112 acquires the image captured by the imaging device 120. Then, the object detection unit 112 detects an object using the acquired image, and sets a first region as a location where the object is detected in the image. The object detection unit 112 detects an object in an image using, for example, a DNN (Deep Neural Network). The DNN is one of the means used in machine learning. It is a technology that extracts features of an object and learns, enabling detection of various objects and improvement of object detection accuracy. Note that the method for detecting an object in the object detection unit 112 may be other than DNN, for example, a method using pattern matching or SSD (Single Shot MultiBox Detector). That is, any method may be used as long as the target object can be detected from the image.

[0018] Then, the object detection unit 112 estimates the two-dimensional coordinate position of the detected object as the object detection result. FIG. 2 is a diagram showing the result of estimating the two-dimensional coordinate position of the detected object. The object detection result is shown in a uv coordinate system with the upper left of the image G captured by the imaging device 120 as the origin (0, 0). When an object is detected, the object detection unit 112 estimates the center (uobs, vobs) coordinates of the object and the width and height of the object on the image G, and indicates the position of the object in the image G by a detection frame W1 shown in FIG. 2. The region indicated by this detection frame W1 is an example of a first region where an object is detected in the image G.

[0019] The disparity information extraction unit 113 extracts a representative value of the disparity information at a location where the disparity information is small within the region indicated by the detection frame W1. Specifically, the disparity information extraction unit 113 obtains the variation in the disparity information at the corresponding location where the object is detected by the object detection unit 112 based on the disparity information obtained from the disparity information acquisition unit 111 and the object detection result obtained from the object detection unit 112. Then, the disparity information extraction unit 113 identifies a location where the variation in the disparity information is small and extracts a representative value of the disparity information at that location. At this time, the representative value of the disparity information is at least one of the mode value, average value, median value, and k-mean of the disparity information in the identified region. Here, k-mean is one of the clustering methods and is an algorithm that classifies values by obtaining the centroid of data.

[0020] Based on the disparity information extracted by the disparity information extraction unit 113, the coordinate conversion unit 114 converts the two-dimensional coordinates of the location where an object is detected in the image G into three-dimensional coordinates.

[0021] <Detailed description of the operation of the disparity information extraction unit 113> Next, the operation of the disparity information extraction unit 113 will be described in detail. FIG. 3 is a flowchart showing the processing of the disparity information extraction unit 113 in the present embodiment. Hereinafter, with reference to FIG. 3, the processing of the disparity information extraction unit 113 of the present embodiment will be described. First, the disparity information extraction unit 113 acquires the disparity information of the corresponding location where an object is detected, and calculates the variation of the disparity information (S301). FIGS. 4(A) to (C) are diagrams for explaining the process in which the disparity information extraction unit 113 calculates the variation of the disparity information in S301 of FIG. 3. Among these, FIG. 4(A) shows the result of detecting an object in the image G. As described with reference to FIG. 2, the object detection unit 112 detects an object in the image G, and the detection frame W1 indicates the position of the object in the image G. FIG. 4(B) is a diagram schematically showing the disparity information acquired by the disparity information acquisition unit 111. In this case, the locations indicated by +, 〇, □, and △ indicate that substantially the same level of disparity information is obtained in the image G. FIG. 4(C) shows the result of superimposing the disparity information on the image G. Then, the disparity information extraction unit 113 calculates the variation of the disparity information within the detection frame W1 using the following Equation (1). In Equation (1), s is the standard deviation, n is the number of pixels within the detection frame W1, xi is the disparity value of each pixel, and x bar is the average of the disparity values within the detection frame W1. That is, in this case, the standard deviation s is calculated as the variation.

[0022]

Equation

[0023] Note that the variation may be calculated by an equation other than the above Equation 1, or by the following Equation 2. In Equation 2, s is the standard deviation, ave is the average value of the parallax, and num is the threshold value. That is, in this case, the left side of Equation 2 is calculated as the variation.

[0024] [Equation]

[0025] Returning to FIG. 3, the parallax information extraction unit 113 determines whether or not the variation of the parallax information calculated in S301 is equal to or less than a predetermined threshold value (S302). Here, as the threshold value, a plurality of debug data are prepared in advance, and a calculated predetermined value (for example, 15) is set. Note that the method for determining the variation in S302 may be other than the standard deviation s. For example, a histogram may be created, and if there are a plurality of bins exceeding a certain frequency, it may be determined that the variation is large. Further, for example, an effective parallax ratio may be calculated, and if the effective parallax ratio is equal to or less than a certain value, it may be determined that the variation is large.

[0026] As a result of S302, if the variation of the parallax information is equal to or less than the threshold value (YES in S302), the parallax information extraction unit 113 calculates a representative value of the parallax information at the corresponding location (S303). In this case, the parallax information extraction unit 113 calculates a representative value of the parallax information for the entire detection frame W1. That is, as described above, the parallax information extraction unit 113 calculates at least one of the mode value, average value, median value, and k-mean of the parallax information in the detection frame W1 as the representative value. Here, as a method for calculating the representative value, a plurality of debug data are prepared in advance, and a method (for example, k-mean) that optimizes the ranging accuracy is applied. On the other hand, if the variation is larger than the threshold value (NO in S302), the parallax information extraction unit 113 divides the detection frame W1 into a plurality of small regions and calculates the variation of the parallax information (S304: small region division process). Then, the parallax information extraction unit 113 acquires the representative value of the parallax information calculated in S303 or S304 (S305).

[0027] FIG. 5 is a flowchart that details the small region division process of S304 in FIG. 3. First, the disparity information extraction unit 113 divides the region indicated by the detection frame W1 into a specified number of small regions (S501). This small region is an example of the second region. Next, the disparity information extraction unit 113 obtains disparity information for each of the divided small regions, calculates the variation by the method of S301, and selects the small region with the minimum variation (S502). Then, at least one of the average value, the mode value, the median value, and k - mean is calculated as a representative value of the disparity information corresponding to the object detection result from the disparity information of the region selected in S502 (S503).

[0028] In S302 of FIG. 3, when the variation of the disparity information is below the threshold, for example, the inside of the detection frame W1 is almost occupied by the object to be detected and other images such as the background are hardly included. In this case, the representative value of the disparity information of the region of the detection frame W1 is considered to be for the object to be detected. On the other hand, in S302, when the variation of the disparity information exceeds the threshold, for example, it is considered that the inside of the detection frame W1 includes not only the object to be detected but also the background. In this case, small regions are set as in S304 of FIG. 3, and the small region with the minimum variation is selected. It is considered that the inside of the small region with the minimum variation is almost occupied by the object to be detected and hardly includes other images such as the background. Therefore, the location where the object actually exists within the region of the detection frame W1 can be specified. And the representative value of the disparity information of this location is set as being for the object to be detected. By doing so, the disparity information of the object to be detected can be accurately extracted.

[0029] FIGS. 6(A) to (B) are diagrams schematically showing the process performed in S501 of FIG. 5, which is the process of dividing the region indicated by the detection frame W1 into small regions W2. FIG. 6(A) shows the area of the detection frame W1. FIG. 6(B) is a diagram showing the case where the area of the detection frame W1 is divided into small areas W2. Here, it shows that the area of the detection frame W1 is divided into 16 small areas W2 by dividing it into 4 parts in both the vertical and horizontal directions of the image G. When dividing the detection frame W1, a predetermined value (for example, the minimum value is 2 and the maximum value is 10) is given to the maximum and minimum values, and the height and width of the detection frame W1 are equally divided according to the number of divisions specified based on the size of the detection frame W1. That is, the detection frame W1 is equally divided in both the vertical and horizontal directions of the image G based on the number of divisions. The number of divisions is calculated, for example, by the following Equation 3 based on the ratio of the number of pixels included in the detection frame W1 of the object detection result to the number of pixels of the image G input to the object detection unit 112. In Equation 3, n is the number of divisions, min is the minimum value, Area obs is the number of pixels of the detection frame W1, Area img is the number of pixels of the image G. At this time, the predetermined values given to the maximum and minimum values are variable.

[0030]

Equation

[0031] In this case, it can be said that the parallax information extraction unit 113 sets a plurality of small areas W2 by dividing this area according to the variation in the parallax information within the area indicated by the detection frame W1, and extracts the representative value of the parallax information at the location where the parallax information becomes small among the plurality of small areas W2. At this time, the small area W2 is set by obtaining the number of divisions in the vertical and horizontal directions of the image G based on the size of the area indicated by the detection frame W1, and dividing the area indicated by the detection frame W1 by this number of divisions. And it can also be said that the number of divisions is obtained so that the number of pixels included in the small area W2 is close to a predetermined number of pixels. As a result, the larger the area indicated by the detection frame W1, the larger the number of divisions. And the number of pixels included in the small area W2 becomes approximately the same.

[0032] According to the image processing apparatus 110 described in detail above, even if the feet of the object in the captured image G are not visible, it is possible to accurately convert two-dimensional coordinates into three-dimensional coordinates. Therefore, regardless of the position of the object in the captured image G, the three-dimensional coordinates can be accurately estimated from the two-dimensional coordinates.

[0033] <Explanation of obstacle detection system> Also, the upper control device 130 that can be regarded as an obstacle detection system for detecting obstacles together with the image processing device 110 detects obstacles according to the three-dimensional coordinates of the object calculated by the image processing device 110. In this case, the upper control device 130 is an example of a detection device. That is, when the image processing device 110 calculates the three-dimensional coordinates of the object and the object becomes an obstacle on the track, the upper control device 130 can detect this. Thereby, the upper control device 130 can accurately detect obstacles. And in this case, the upper control device 130 can perform control such as stopping the train 100, for example.

[0034] Note that the image processing device 110 and the upper control device 130 can also be made to detect objects around the track on which the train travels as objects and not detect objects that do not exist around the track. This can be realized, for example, by limiting the image G captured by the imaging device 120 to the periphery of the track. Also, the upper control device 130 may detect the track in the image G and detect objects that fall within a predetermined range in the image G based on this track. Also, the image processing device 110 and the upper control device 130 can be made to detect at least one of a signal and a security guard as an object around the track. When the image processing device 110 or the upper control device 130 detects a signal, it may perform recognition of the indication according to the three-dimensional position of the signal. Also at this time, the upper control device 130 can perform control of the train 100 according to the indication of the signal, for example.

[0035] Also, in the above-described form, a train was cited as the moving body and described, but the moving body is not limited to a train. For example, the moving body may be an automobile, and the object detection unit 112 may detect an object that may collide with itself.

[0036] <Explanation of the image processing method> The processing performed by the image processing apparatus 110 in the present embodiment described above is realized by the cooperation of software and hardware resources. That is, a processor such as a CPU provided in the image processing apparatus 110 loads a program that realizes each function of the image processing apparatus 110 into the main memory and executes it to realize these functions.

[0037] Therefore, the processing performed by the above-described image processing apparatus 110 is an image processing method for calculating the three-dimensional coordinates of an object in an image G captured by a plurality of cameras that image the front of the moving body. By the processor executing a program recorded in the memory, parallax information is acquired based on the image G, an object is detected using the image G, a first region is set as the location where the object is detected in the image G, and a representative value of the parallax information at a location where the parallax information is small within the first region is extracted. Based on the extracted representative value, it can be regarded as an image processing method for converting the two-dimensional coordinates of the first region into three-dimensional coordinates.

[0038] Note that the program for realizing the present embodiment can be provided not only by communication means but also by storing it in a recording medium such as a CD-ROM and providing it.

[0039] Also, the form described above will include at least the following technical matters. <Technical matter 1> An image processing apparatus that calculates the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body, comprising: a disparity information acquisition unit that acquires disparity information based on the image; an object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image; a disparity information extraction unit that extracts a representative value of the disparity information at a location where the disparity information is small within the first region; and a coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the disparity information extraction unit. <Technical matter 2> In the image processing apparatus described in Technical matter 1 above, the disparity information extraction unit sets a plurality of second regions by dividing the first region according to the variation in the disparity information within the first region, and extracts a representative value of the disparity information at a location where the disparity information becomes small among the plurality of second regions. <Technical matter 3> In the image processing apparatus described in Technical matter 2 above, the second region is set by obtaining the number of divisions in the vertical and horizontal directions of the image based on the size of the first region, and dividing the first region by the number of divisions. <Technical matter 4> In the image processing apparatus described in Technical matter 3 above, the number of divisions is obtained so that the number of pixels included in the second region is close to a predetermined number of pixels. <Technical matter 5> In the image processing apparatus described in any one of Technical matters 1 to 4 above, the moving body is a train, and the object detection unit detects an object around the track on which the train travels. <Technical matter 6> In the image processing apparatus described in Technical matter 5 above, the object detection unit detects at least one of a traffic signal and a security guard as an object around the track. <Technical matter 7> In the image processing apparatus described in any one of Technical matters 1 to 6 above, the image is an image captured by a stereo camera incorporating a plurality of cameras. <Technical matter 8> In the image processing apparatus shown by any one of the above Technical Matters 1 to 7, the representative value is at least one of the mode value, average value, median value, and k-mean of the disparity information. <Technical Matter 9> In the image processing apparatus shown by any one of the above Technical Matters 1 to 4, the moving body is an automobile, and the object detection unit detects an object that may collide.

[0040] <Technical Matter 10> An image processing method for calculating the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body. By a processor executing a program recorded in a memory, based on the image, disparity information is acquired, the image is used to detect an object, and a first region is set as a location where the object is detected in the image. A representative value of the disparity information at a location where the disparity information is small within the first region is extracted, and based on the extracted representative value, the two-dimensional coordinates of the first region are converted into three-dimensional coordinates. <Technical Matter 11> In the image processing method shown by the above Technical Matter 10, the moving body is a train, and an object around the track on which the train travels is detected. <Technical Matter 12> In the image processing method shown by the above Technical Matter 11, at least one of a traffic signal and a security guard is detected as an object around the track. <Technical Matter 13> In the image processing method shown by any one of the above Technical Matters 10 to 12, the image is an image captured by a stereo camera incorporating a plurality of cameras. <Technical Matter 14> In the image processing method shown by any one of the above Technical Matters 10 to 13, the representative value is at least one of the mode value, average value, median value, and k-mean of the disparity information.

[0041] <Technical Matter 15> An obstacle detection system for detecting obstacles, comprising: an image processing device that calculates the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body; and a detection device that detects an obstacle according to the three-dimensional coordinates of the object calculated by the image processing device. The image processing device includes: a parallax information acquisition unit that acquires parallax information based on the image; an object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image; a parallax information extraction unit that extracts a representative value of the parallax information at a location where the parallax information is small within the first region; and a coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the parallax information extraction unit.

[0042] As described above, the present embodiment has been described. However, the technical scope of the present invention is not limited to the scope described in the above embodiment. It is clear from the description of the claims that various modifications or improvements added to the above embodiment are also included in the technical scope of the present invention.

Explanation of reference numerals

[0043] 1…Train control system, 100…Train, 110…Image processing device, 111…Parallax information acquisition unit, 112…Object detection unit, 113…Parallax information extraction unit, 114…Coordinate conversion unit, 120…Imaging device

Claims

1. An image processing apparatus for calculating three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body, comprising: a disparity information acquisition unit that acquires disparity information based on the image; an object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image; a disparity information extraction unit that extracts a representative value of the disparity information at a location where the disparity information is small within the first region; a coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the disparity information extraction unit. The image processing apparatus is characterized by comprising the above.

2. The image processing apparatus according to claim 1, wherein the disparity information extraction unit sets a plurality of second regions by dividing the first region according to the variation of the disparity information within the first region, and extracts a representative value of the disparity information at a location where the disparity information becomes small among the plurality of second regions.

3. The image processing apparatus according to claim 2, wherein the second region is set by obtaining the number of divisions in the vertical and horizontal directions of the image based on the size of the first region and dividing the first region by the number of divisions.

4. The image processing apparatus according to claim 3, wherein the number of divisions is obtained such that the number of pixels included in the second region is close to a predetermined number of pixels.

5. The image processing apparatus according to claim 1, wherein the moving body is a train, and the object detection unit detects an object around the track on which the train travels.

6. The image processing apparatus according to claim 5, wherein the object detection unit detects at least one of a traffic signal and a security guard as an object around the track.

7. The image processing apparatus according to claim 1, wherein the image is an image captured by a stereo camera incorporating a plurality of cameras.

8. The image processing apparatus according to claim 1, wherein the representative value is at least one of the mode value, average value, median value, and k-mean of the disparity information.

9. The image processing apparatus according to claim 1, wherein the moving body is an automobile, and the object detection unit detects an object that may cause a collision.

10. An image processing method for calculating three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a moving body, comprising: a processor executes a program recorded in a memory to acquire disparity information based on the image. Detect an object using the image, and set a first region as a location where the object is detected in the image. Extract a representative value of the disparity information at a location where the disparity information is small within the first region. Convert the two-dimensional coordinates of the first region into three-dimensional coordinates based on the extracted representative value. Image processing method.

11. The mobile body is a train, and the image processing method according to claim 10, which detects an object around a track on which the train runs.

12. The image processing method according to claim 11, which detects at least one of a signal and a security guard as an object around the track.

13. The image processing method according to claim 10, wherein the image is an image captured by a stereo camera incorporating a plurality of cameras.

14. The image processing method according to claim 10, wherein the representative value is at least one of the mode value, average value, median value, and k-mean of the disparity information.

15. An obstacle detection system for detecting an obstacle, comprising: An image processing device that calculates the three-dimensional coordinates of an object in an image captured by a plurality of cameras that image the front of a mobile body; A detection device that detects an obstacle according to the three-dimensional coordinates of the object calculated by the image processing device; And comprising The image processing device A disparity information acquisition unit that acquires disparity information based on the image; An object detection unit that detects an object using the image and sets a first region as a location where the object is detected in the image; A disparity information extraction unit that extracts a representative value of the disparity information at a location where the disparity information is small within the first region; A coordinate conversion unit that converts the two-dimensional coordinates of the first region into three-dimensional coordinates based on the representative value extracted by the disparity information extraction unit; An obstacle detection system comprising the above.

Citation Information

Patent Citations

  • Object tracking device and program

    JP2011065338A