Object detection device, object detection system, moving object, and object detection method
The parallax map is generated by a stereo camera and the parallax processing of road surface and object are solved, which solves the accuracy of object detection in complex road environments and realizes high-performance object detection.
Patent Information
- Application Number
- CN202080067162.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-24
- Filing Date
- 2020-09-16
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2040-09-16
AI Technical Summary
It is difficult for existing object detection devices to achieve high-performance detection under factors such as road shape and structure.
A first parallax map is generated by a stereo camera, the road surface position is estimated through the road surface detection process, the object parallax determination process is used to determine the object parallax, and it is converted into a point group of world coordinate systems for detection.
Improve the accuracy and performance of object detection and reduce the error detection rate, especially in complex road environments.
Smart Images

Figure CN114450714B_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims the benefit of Japanese Patent Application No. 2019-173341 filed on September 24, 2019, and the entire disclosure of that application is incorporated herein by reference. Technical Field
[0003] The present disclosure relates to an object detection device, an object detection system, a mobile object, and an object detection method. Background Art
[0004] In recent years, object detection devices using stereo cameras have been installed in mobile vehicles such as automobiles to detect objects and measure distances. These object detection devices obtain a distance distribution (or parallax distribution) across multiple images captured by multiple cameras, and identify obstacles based on this distance distribution information (for example, see Patent Document 1).
[0005] Prior art literature
[0006] Patent Literature
[0007] Patent Document 1: Japanese Patent Application Laid-Open No. 5-265547 Summary of the Invention
[0008] The object detection device disclosed herein includes a processor configured to execute road surface detection processing, object disparity determination processing, and object detection processing. The road surface detection processing estimates the position of a road surface in real space based on a first disparity map. The first disparity map is generated based on the output of a stereo camera that captures an image containing the road surface. The first disparity map is a map that associates disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to the horizontal direction of the image captured by the stereo camera. The object disparity determination processing determines the disparity as an object disparity when the number of occurrences of disparity within each coordinate range in the first direction of the first disparity map exceeds a predetermined threshold corresponding to the disparity. The object detection processing transforms the object disparity information into a point cluster in the xz coordinate space of a world coordinate system, where each point indicates the presence or absence of an object. The object is detected by extracting a set of the point clusters.
[0009] The object detection system disclosed herein includes a stereo camera that captures multiple images having parallaxes, and an object detection device including at least one processor. The processor is configured to generate a first disparity map based on the output of the stereo camera that captures images including a road surface. The first disparity map is a map that associates disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to the horizontal direction of the images captured by the stereo camera. The processor is configured to further perform the following road surface detection processing, object disparity determination processing, and object detection processing. The road surface detection processing estimates the position of the road surface in real space based on the first disparity map. The object disparity determination processing determines the disparity as an object disparity when the number of occurrences of disparity within each coordinate range in the first direction of the first disparity map exceeds a predetermined threshold corresponding to the disparity. The object detection processing transforms the object disparity information into a point cluster in the xz coordinate space of a world coordinate system, where each point indicates the presence or absence of an object. Objects are detected by extracting a set of the point clusters.
[0010] The mobile object disclosed herein includes an object detection system. The object detection system includes a stereo camera that captures multiple images having parallaxes, and an object detection device having at least one processor. The processor is configured to generate a first disparity map based on the output of the stereo camera that captures an image containing a road surface. The first disparity map is a map that associates disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to the horizontal direction of the image captured by the stereo camera. The processor is configured to further perform road surface detection processing, object disparity determination processing, and object detection processing. The road surface detection processing estimates the position of the road surface in real space based on the first disparity map. The object disparity determination processing determines the disparity as an object disparity when the number of occurrences of disparity within each coordinate range in the first direction of the first disparity map exceeds a predetermined threshold corresponding to the disparity. The object detection processing transforms the object disparity information into a point cluster in the xz coordinate space of a world coordinate system, where each point indicates the presence or absence of an object. Objects are detected by extracting a set of the point clusters.
[0011] The object detection method disclosed herein includes acquiring or generating a first disparity map. The first disparity map is generated based on the output of a stereo camera that captures an image containing a road surface. The first disparity map is a map that associates disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to the horizontal direction of the image captured by the stereo camera. The object detection method includes performing road surface detection processing, object disparity determination processing, and object detection processing. The road surface detection processing includes estimating the position of a road surface in real space based on the first disparity map. The object disparity determination processing includes determining the disparity as an object disparity when the number of occurrences of disparity within each coordinate range in the first direction of the first disparity map exceeds a predetermined threshold corresponding to the disparity. The object detection processing includes transforming the object disparity information into a point cluster in the xz coordinate space of a world coordinate system, where each point indicates the presence or absence of an object, and detecting the object by extracting a set of the point clusters. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 1 is a block diagram showing a schematic structure of an object detection system according to an embodiment of the present disclosure.
[0013] Figure 2 It schematically shows the Figure 1 Side view of a moving body of an object detection system.
[0014] Figure 3 It schematically shows the Figure 1 Front view of a moving body of an object detection system.
[0015] Figure 4 is a block diagram showing a schematic structure of an object detection system according to another embodiment of the present disclosure.
[0016] Figure 5 It shows Figure 1 A flowchart of an example of processing performed by an object detection device.
[0017] Figure 6 A diagram for explaining an example of a first parallax image acquired or generated by the object detection device.
[0018] Figure 7 This is a flowchart showing an example of a process for estimating a road surface position.
[0019] Figure 8 This is a flowchart showing an example of a process of extracting a road surface candidate parallax from a first parallax image.
[0020] Figure 9 A diagram for explaining the positional relationship between a road surface and a stereo camera.
[0021] Figure 10 3 is a diagram for explaining the procedure for extracting road surface candidate parallaxes.
[0022] Figure 11 : is a diagram showing the road surface range used to generate a disparity histogram.
[0023] Figure 12 is the road parallax d r dv correlation diagram of the relationship between the vertical axis (v coordinate).
[0024] Figure 13 This is a diagram for explaining a method of detecting whether or not there is an object other than road surface parallax.
[0025] Figure 14 is the road parallax d r Flowchart for processing the relationship between the vertical left side of the image (v coordinate) using a straight line approximation.
[0026] Figure 15 is used to explain the road parallax d using the first straight line. r An approximate diagram of .
[0027] Figure 16 This is a diagram for explaining a method of determining the second straight line.
[0028] Figure 17 is a graph showing the road parallax d r A diagram showing an example of the result of linearly approximating the relationship between the vertical coordinate (v coordinate) of the image.
[0029] Figure 18 A diagram for explaining an example of a second parallax image from which unnecessary parallax is removed.
[0030] Figure 19 A diagram showing an example of a second parallax image.
[0031] Figure 20 is a diagram showing one of the areas on the second parallax image where a histogram is created.
[0032] Figure 21 This is a diagram for explaining a method of determining object parallax using a histogram.
[0033] Figure 22 This is a diagram for explaining a method of acquiring height information of object parallax.
[0034] Figure 23 This is a diagram for explaining a method of acquiring height information of object parallax.
[0035] Figure 24: is a diagram showing an example of the distribution of a point group representing object parallax in the ud space.
[0036] Figure 25 This is a diagram of the road surface viewed from the height direction (y direction).
[0037] Figure 26 This is a flowchart showing an example of a process for detecting an object.
[0038] Figure 27 It is a graph that transforms the object parallax into a point group on the xz plane in real space.
[0039] Figure 28 A diagram showing an example of a point group continuously arranged in the z direction extracted from a point group on the xz plane.
[0040] Figure 29 3 is a diagram showing an example of a point group on the xz plane after deleting point groups that are continuously arranged in the z direction.
[0041] Figure 30 3 is a diagram showing an example of a joint region generated from a point group on the xz plane.
[0042] Figure 31A This is a diagram for explaining a first example of extracting a point group continuously arranged in the z direction.
[0043] Figure 31B This is a diagram for explaining a second example of extracting a point group continuously arranged in the z direction.
[0044] Figure 32 This is a diagram showing an example of other vehicles traveling on the expressway.
[0045] Figure 33 This is a diagram showing an example of a method for outputting object detection results. DETAILED DESCRIPTION
[0046] Due to various factors such as the shape of the road and structures and objects on the road, an object detection device may not be able to achieve high performance. Therefore, it is desirable for an object detection device to have high performance. According to one aspect of the present disclosure, the performance of detecting an object can be improved.
[0047] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. It should be noted that the figures used in the following description are schematic. The dimensions, proportions, etc. in the drawings do not necessarily correspond to the actual ones. The figures showing images taken by a camera and parallax images, etc. include figures made for illustration. These images are different from the images actually taken or processed. In addition, in the following description, the "subject" is an object photographed by a camera. The "subject" includes objects, roads, the sky, etc. An "object" has a specific position and size in space. An "object" may also be referred to as a "three-dimensional object."
[0048] like Figure 1 As shown, the object detection system 1 includes a stereo camera 10 and an object detection device 20. The stereo camera 10 and the object detection device 20 can communicate via wired or wireless communication. The stereo camera 10 and the object detection device 20 can communicate via a network. The network can include, for example, a wired or wireless LAN (Local Area Network), a CAN (Controller Area Network), etc. The stereo camera 10 and the object detection device 20 can be housed in the same housing and constructed integrally. The stereo camera 10 and the object detection device 20 can be located in a mobile body 30 described later and can communicate with an ECU (Electronic Control Unit) in the mobile body 30.
[0049] A "stereo camera" is a plurality of cameras that have parallax with each other and cooperate with each other. A stereo camera includes at least two or more cameras. By making the multiple cameras cooperate, a stereo camera can shoot an object from multiple directions. A stereo camera can be a device that includes multiple cameras in one housing. A stereo camera can be a device that includes two or more cameras that are independent of each other and located at positions separated from each other. A stereo camera is not limited to multiple cameras that are independent of each other. In the present disclosure, for example, a camera having an optical mechanism that guides light incident on two separate points to one light receiving element can be used as a stereo camera. In the present disclosure, multiple images of the same subject shot from different viewpoints are referred to as "stereo images."
[0050] The stereo camera 10 includes a first camera 11 and a second camera 12. The first camera 11 and the second camera 12 each include an optical system defining an optical axis OX and an imaging device. The first camera 11 and the second camera 12 each have a different optical axis OX. In this specification, only one reference numeral OX is used to represent the optical axis OX of both the first camera 11 and the second camera 12. The imaging device includes a CCD image sensor (Charge-Coupled Device Image Sensor) and a CMOS image sensor (Complementary MOS Image Sensor). The imaging devices included in the first camera 11 and the second camera 12 can be located in the same plane perpendicular to the optical axis OX of each camera. The first camera 11 and the second camera 12 generate image signals representing images formed by the imaging devices. In addition, the first camera 11 and the second camera 12 can perform any processing on the captured images, such as distortion correction, brightness adjustment, contrast adjustment, and gamma correction.
[0051] The optical axes OX of the first and second cameras 11, 12 are oriented in a direction that allows them to capture the same subject. The optical axes OX and positions of the first and second cameras 11, 12 are determined so that the captured images include at least the same subject. In one embodiment, the optical axes OX of the first and second cameras 11, 12 are oriented parallel to each other. This parallelism is not strictly required; assembly and installation variations, as well as variations caused by time, are permitted. In another embodiment, the optical axes OX of the first and second cameras 11, 12 are not necessarily parallel and may face different directions. Even if the optical axes OX of the first and second cameras 11, 12 are not parallel, stereoscopic images can be generated by converting the images within the stereo camera 10 or the object detection device 20. The distance between the optical centers of the first and second cameras 11, 12 of the stereo camera 10 is referred to as the baseline length. The baseline length is equivalent to the distance between the lens centers of the first and second cameras 11, 12. The direction connecting the optical centers of the first and second cameras 11, 12 of the stereo camera 10 is referred to as the baseline length direction.
[0052] The first camera 11 and the second camera 12 are separated from each other in a direction intersecting the optical axis OX. In one embodiment of the multiple embodiments, the first camera 11 and the second camera 12 are arranged in the left-right direction. When facing forward, the first camera 11 is located on the left side of the second camera 12. When facing forward, the second camera 12 is located on the right side of the first camera 11. Since the positions of the first camera 11 and the second camera 12 are different, the positions of the subjects corresponding to each other in the two images captured by each camera are different. The first image output from the first camera 11 and the second image output from the second camera 12 are stereo images captured from different viewpoints. The first camera 11 and the second camera 12 capture the subject at a predetermined frame rate (for example, 30fps).
[0053] like Figure 2 and Figure 3 As shown, Figure 1 The object detection system 1 is mounted on a moving object 30. Figure 2 As shown in the side view of FIG, the first camera 11 and the second camera 12 are arranged to capture images in front of the mobile object 30. The first camera 11 and the second camera 12 are arranged to capture images that include the road surface. In other words, the stereo camera 10 is arranged to capture images that include the road surface. In one embodiment, the optical axis OX of the optical system of each of the first camera 11 and the second camera 12 is arranged to be substantially parallel to the front of the mobile object 30.
[0054] In this application, the direction in which the moving body 30 is traveling forward is referred to as the front or z-direction. The direction opposite to the front is referred to as the rear. The left and right directions are defined based on the state facing the front of the moving body 30. The direction perpendicular to the z-direction and from left to right is referred to as the x-direction. The x-direction may coincide with the baseline length direction. The direction perpendicular to and upward from the road surface near the moving body 30 is referred to as the height direction or the y-direction. The y-direction may be perpendicular to the x-direction and the z-direction. The x-direction is also referred to as the horizontal direction. The y-direction is also referred to as the vertical direction. The z-direction is also referred to as the depth direction.
[0055] "Mobile objects" in this disclosure may include, for example, vehicles and aircraft. Vehicles may include, for example, automobiles, industrial vehicles, rail vehicles, household vehicles, and fixed-wing aircraft taxiing on runways. Automobiles may include, for example, cars, trucks, buses, motorcycles, and trolleybuses. Industrial vehicles may include, for example, those used in agriculture and construction. Industrial vehicles may include, for example, forklifts and golf carts. Agricultural industrial vehicles may include, for example, tractors, tillers, rice transplanters, harvesters, combines, and lawn mowers. Construction industrial vehicles may include, for example, bulldozers, scrapers, excavators, truck cranes, dump trucks, and road rollers. Vehicles may include human-powered vehicles. The classification of vehicles is not limited to the above examples. For example, automobiles may include industrial vehicles that can travel on roads. The same vehicle may be included in multiple classifications. Aircraft may include, for example, fixed-wing aircraft and rotary-wing aircraft.
[0056] The vehicle 30 of the present disclosure travels on a travel path including a road, a track, etc. The surface of the travel path on which the vehicle 30 travels is called a road surface.
[0057] The first camera 11 and the second camera 12 are mounted at different locations on the mobile object 30. In one embodiment, the first camera 11 and the second camera 12 are mounted inside the mobile object 30, which is a vehicle, and can capture images of the exterior of the mobile object 30 through the windshield. For example, the first camera 11 and the second camera 12 are positioned in front of the rearview mirror or on the dashboard. In one embodiment, the first camera 11 and the second camera 12 can be attached to the front bumper, fender grille, side fenders, headlight module, hood, etc. of the vehicle.
[0058] Object detection device 20 includes an acquisition unit 21, an image processing unit 22 (processor), a memory 23, and an output unit 24. Object detection device 20 can be placed anywhere in mobile object 30. For example, object detection device 20 can be placed in the instrument panel of mobile object 30.
[0059] The acquisition unit 21 is an input interface of the object detection device 20, which is used to receive information input from the stereo camera 10 and other devices. A physical connector and a wireless communication device can be used as the acquisition unit 21. Physical connectors include electrical connectors corresponding to the transmission of electrical signals, optical connectors corresponding to the transmission of optical signals, and electromagnetic connectors corresponding to the transmission of electromagnetic waves. Electrical connectors include connectors that comply with IEC60603, connectors that comply with USB standards, connectors corresponding to RCA terminals, connectors corresponding to S terminals specified in EIAJ CP-1211A, connectors corresponding to D terminals specified in EIAJ RC-5237, connectors that comply with HDMI (registered trademark) standards, and connectors corresponding to coaxial cables including BNCs. Optical connectors include various connectors that comply with IEC 61754. Wireless communication devices include wireless communication devices that comply with various standards of Bluetooth (registered trademark) and IEEE802.11. Wireless communication devices include at least one antenna.
[0060] Image data corresponding to images captured by the first camera 11 and the second camera 12 may be input to the acquisition unit 21. The acquisition unit 21 transmits the input image data to the image processing unit 22. The acquisition unit 21 may correspond to a transmission method of a shooting signal of the stereo camera 10. The acquisition unit 21 may be connected to an output interface of the stereo camera 10 via a network.
[0061] The image processing unit 22 includes one or more processors. Processors include general-purpose processors that load specific programs and execute specific functions, and special-purpose processors that specialize in specific processing. Special-purpose processors include application-specific integrated circuits (ASICs). Processors include programmable logic devices (PLDs). PLDs include FPGAs (Field-Programmable Gate Arrays). The image processing unit 22 can be any of an SoC (System-on-a-Chip) and a SiP (System In a Package) in which one or more processors cooperate. The processing performed by the image processing unit 22 can also be converted into processing performed by a processor.
[0062] The image processing unit 22 includes various functional blocks, including a parallax image generation unit 25, a road surface detection unit 26, an unnecessary parallax removal unit 27, a clustering unit 28, and a grouping unit 29. The parallax image generation unit 25 generates a first parallax image based on the first and second images output from the stereo camera 10. The first parallax image is an image in which pixels representing parallax are arranged on a two-dimensional plane formed by a horizontal direction corresponding to the horizontal direction of the image captured by the stereo camera 10 and a vertical direction intersecting the horizontal direction. The horizontal direction is a first direction, and the vertical direction is a second direction. The horizontal and vertical directions may be orthogonal to each other. The horizontal direction corresponds to the width of the road surface. If the image captured by the stereo camera 10 includes a horizontal line, the horizontal direction is parallel to the horizontal line. The vertical direction corresponds to the direction of gravity in real space. The road surface detection unit 26, the unnecessary parallax removal unit 27, the clustering unit 28, and the grouping unit 29 perform a series of processes for detecting objects based on the first parallax image.
[0063] In the information processing in the image processing unit 22, the first parallax image undergoes various operations as a first parallax map. In this first parallax map, parallax information obtained from the output of the stereo camera 10 is associated with two-dimensional coordinates consisting of horizontal and vertical directions. These operations include calculations and writing to and reading from the memory 23. The first parallax image can also be considered a first parallax map. In the following description, processing of the first parallax image can also be considered processing of the first parallax map.
[0064] Each functional block of the image processing unit 22 may be a hardware module or a software module. The processing performed by each functional block can also be considered as processing performed by the image processing unit 22. The image processing unit 22 can perform all actions of each functional block. The processing that the image processing unit 22 causes any functional block to perform can also be performed by the image processing unit 22 itself.
[0065] The memory 23 stores programs used for various processes and information used in operations. The memory 23 includes a volatile memory and a non-volatile memory. The memory 23 includes a memory independent of the processor and a memory built into the processor.
[0066] The output unit 24 is the output interface of the object detection device 20. It can output the processing results of the object detection device 20 to other devices within the mobile object 30 or to devices external to the mobile object 30, such as other vehicles and roadside equipment. Other devices that can appropriately use the information received from the object detection device 20 include driving assistance systems such as automatic cruise control and safety equipment such as automatic braking systems. Similar to the acquisition unit 21, the output unit 24 includes various interfaces for both wired and wireless communication. For example, the output unit 24 includes a CAN interface to communicate with other devices within the mobile object 30.
[0067] The object detection device 20 can be configured to implement the processing performed by the image processing unit 22 described below by reading a program recorded in a non-temporary computer-readable medium. Non-temporary computer-readable media include magnetic storage media, optical storage media, magneto-optical storage media, and semiconductor storage media, but are not limited to these. Magnetic storage media include magnetic disks, hard disks, and magnetic tapes. Optical storage media include CDs (Compact Discs), DVDs, Blu-ray Discs (Blu-ray (registered trademark) Discs), and other optical disks. Semiconductor storage media include ROMs (Read Only Memory), EEPROMs (Electrically Erasable Programmable Read-Only Memory), and flash memories.
[0068] like Figure 4 As shown, in an object detection system 1A according to another embodiment of the present disclosure, the parallax image generation unit 25 may be mounted on hardware independent of the object detection device 20. Figure 4 In, with Figure 1 The same or similar components are used Figure 1 The same symbols are used to represent them. Figure 4 The parallax image generating unit 25 in FIG. 2 can also be regarded as a parallax image generating device. Figure 4 The parallax image generating unit 25 includes a processor. Figure 4 The processor included in the parallax image generating section 25 generates a first parallax image based on the first image and the second image outputted from the first camera 11 and the second camera 12 of the stereo camera 10 , respectively. Figure 4 The acquisition unit 21 included in the object detection device 20 acquires the first parallax image from the parallax image generation unit 25 . Figure 4 The object detection device 20 and the parallax image generation unit 25 in the embodiment can be regarded as one object detection device 20. Figure 4 In the object detection device 20 shown in FIG. 1 , the functions of the road surface detection unit 26, the unnecessary parallax removal unit 27, the clustering unit 28, and the grouping unit 29 of the image processing unit 22 are similar to those of the Figure 1 The corresponding functional blocks are the same.
[0069] The following will refer to Figure 5 The flowchart further illustrates the processing performed by each functional block of the image processing unit 22. Figure 5 1 is a flowchart for explaining the overall processing of the object detection method executed by the object detection device 20 .
[0070] First, in detail Figure 5Before describing the processing performed in each step of the flowchart, the outline and purpose of the processing in each step will be briefly described.
[0071] Step S101 is a step of acquiring or generating a first parallax image as a target for object detection at a stage prior to the first process described later. Step S101 is executed by the acquiring unit 21 or the parallax image generating unit 25 .
[0072] Step S102 is a step for estimating the road surface position. The processing performed in step S102 may also be referred to as the first processing (road surface detection processing). Step S102 is performed by the road surface detection unit 26. By estimating the road surface position, the parallax representing the road surface relative to the vertical coordinate can be found in the first parallax image. The road surface position is necessary for removing unnecessary parallax in subsequent processing and / or determining the height position of the road surface in real space.
[0073] Step S103 is a step of generating a second parallax image from which unnecessary parallax is removed. The processing performed in step S103 may also be referred to as a second processing (unnecessary parallax removal processing). The second processing is performed by the unnecessary parallax removal unit 27. Unnecessary parallax is the parallax represented by the pixels corresponding to the subject whose height from the road surface in actual space is included in a predetermined range. For example, unnecessary parallax includes the parallax of the road surface contained in the first parallax image and the parallax of the structure contained in the air portion, etc. By removing unnecessary parallax from the first parallax image, the possibility of erroneously detecting white lines on the road surface and structures existing above the road surface as objects on the road surface is reduced. Thus, the accuracy of object detection is improved. Step S103 may be omitted. When step S103 is omitted, the image processing unit 22 may not include the unnecessary parallax removal unit 27. The same applies to the following description.
[0074] In the information processing performed by the image processing unit 22, the second parallax image is subjected to various operations as a second parallax image obtained by removing unnecessary parallax information from the parallax information included in the first parallax image. The second parallax image can also be considered the second parallax map. In the following description, processing of the second parallax image can also be considered processing of the second parallax map.
[0075] Step S104 is a step for determining object disparity at each horizontal coordinate in the second disparity image. Step S104 is performed by the clustering unit 28. Object disparity refers to disparity in an area determined to be recognizable as an object in real space under predetermined conditions. If step S103 is omitted, step S104 uses the first disparity image instead of the second disparity image to determine object disparity at each horizontal coordinate in the first disparity image. This applies similarly to the following description.
[0076] Step S105 is a step for calculating height information related to the object parallax based on the object parallax distribution in the second parallax image and the road surface position estimated in step S102. Step S105 is performed by the clustering unit 28. The processing performed in steps S104 and S105 can also be referred to as the third processing (object parallax determination processing). If height information is not required, step S105 can be omitted.
[0077] Step S106 is a step of detecting an object by converting object parallax information into real-space coordinates and extracting a set (group) of object parallaxes. The process performed in step S106 may also be referred to as a fourth process (object detection process). The fourth process is performed by the grouping unit 29.
[0078] Step S107 is the step of outputting the detected object information from the output unit 24. Based on the results of step S106, the detected object position and width information at the angle of the stereo camera 10 can be obtained. The information obtained in step S105 includes the height information of the detected object. This information can be provided to other devices within the mobile object 30.
[0079] Next, the details of each step will be explained.
[0080] First, the image processing unit 22 acquires or generates a first parallax image (step S101). Figure 1 In the object detection system 1 shown in FIG. 1 , the image processing unit 22 generates a first parallax image based on the first image and the second image acquired by the acquisition unit 21. The generation of the first parallax image is performed by the parallax image generation unit 25. Figure 4 In the object detection system 1A shown, the image processing unit 22 acquires the first parallax image generated by the parallax image generating unit 25 via the acquiring unit 21. The image processing unit 22 may store the first parallax image in the memory 23 for subsequent processing.
[0081] The parallax image generating unit 25 generates a first parallax image by calculating the parallax distribution between the first image acquired from the first camera 11 and the second image acquired from the second camera 12. Since the method of generating the first parallax image is well known, only a brief description will be given.
[0082] The parallax image generating unit 25 divides the image of one of the first image and the second image (for example, the first image) into a plurality of small areas. The small area may be a rectangular area in which a plurality of pixels are arranged in the vertical and horizontal directions. For example, the small area may be composed of three pixels in the vertical direction and three pixels in the horizontal direction. The number of pixels contained in the vertical and horizontal directions of the small area is not limited to three. The number of pixels contained in the vertical and horizontal directions of the small area may be different. The parallax image generating unit 25 transforms each image divided into a plurality of small areas and the image of the other party (for example, the second image) in the horizontal direction and matches them. For image matching, a method using the SAD (Sum of Absolute Difference) function is known. This represents the sum of the absolute values of the differences between the brightness values within the small areas. When the SAD function is minimum, the two images are judged to be most similar. The matching of stereo images is not limited to the method using the SAD function, and other methods may also be used.
[0083] The parallax image generation unit 25 calculates the parallax of each small area based on the difference between the horizontal pixel positions of the two areas when matching the first image and the second image. The size of the parallax can be expressed in units of the horizontal width of the pixel. By performing interpolation processing, the size of the parallax can be calculated with an accuracy of less than one pixel. The size of the parallax corresponds to the distance between the subject photographed by the stereo camera 10 and the stereo camera 10 in real space. If the parallax is large, it means the distance is close; if the parallax is small, it means the distance is far. The parallax image generation unit 25 generates a first parallax image representing the calculated parallax distribution. Hereinafter, the pixels representing the parallax constituting the first parallax image will be referred to as parallax pixels. The parallax image generation unit 25 can generate a parallax image with the same clarity as the pixels of the original first and second images.
[0084] Figure 6 is a diagram for explaining the first parallax image. Figure 6 In FIG. 4 , another vehicle 42 is traveling on the road surface 41 in front of the moving object 30 .
[0085] like Figure 6 As shown, in the first parallax image, the pixels representing the parallax are located on a two-dimensional plane formed by the horizontal direction (first direction) and the vertical direction (second direction) orthogonal to the horizontal direction of the stereo camera 10. The coordinate representing the position in the horizontal direction is called the u coordinate. The coordinate representing the position in the vertical direction is called the v coordinate. The u coordinate and the v coordinate are called image coordinates. In each figure represented by the image of the present disclosure, the u coordinate is the coordinate from left to right. The v coordinate is the coordinate from top to bottom. The origin of the uv coordinate space can be the upper left end of the first parallax image. The u coordinate and the v coordinate can be expressed in pixels.
[0086] The parallax image generating unit 25 can replace the difference in display parallax with the difference in brightness or color of the pixels. Figure 6 In , parallax is represented by different shading for ease of illustration. Figure 6 In , the darker the shadow, the smaller the parallax; the lighter the shadow, the larger the parallax. Figure 6 In the image, equally shaded areas indicate they are within the predetermined parallax range. In actual first-parallax images, there are areas in the UV coordinate space where parallax is easily obtained and areas where it is difficult to obtain. For example, parallax is difficult to obtain in spatially uniform areas of a subject, such as car windows, or in areas overexposed due to sunlight reflection. In first-parallax images, if objects or structures are present, they appear with a different brightness or color than the more distant background.
[0087] The parallax image generating unit 25 does not need to display the first parallax image as an image after calculating the parallax. The parallax image generating unit 25 may generate and store information of the first parallax image as a first parallax map within the image processing unit 22 and perform necessary processing.
[0088] After step S101, the image processing unit 22 performs a first process of estimating the position of the road surface 41 based on the first parallax image (step S102). The first process is performed by the road surface detection unit 26. Figure 7 、 8 14 and 15, the road surface position estimation process performed by the road surface detection unit 26 is described. First, the road surface detection unit 26 extracts a road surface candidate parallax d from the first parallax image. c (Step S201). Road surface candidate disparity d c is the road disparity d collected from the first disparity image r The parallax with a high probability of corresponding. Road parallax d r It refers to the parallax of the road surface 41 area. Road surface parallax d r is the parallax excluding objects on the road surface 41. Road surface parallax d r Indicates the distance to the corresponding position on the road. Road parallax d r are collected as having close values at positions with the same v coordinate.
[0089] Road surface candidate disparity d c The details of the extraction process are in Figure 8 As shown in the flow chart. Figure 8As shown, the road surface detection unit 26 calculates an initial value d0 of the road surface candidate disparity based on the installation position of the stereo camera 10, that is, an initial value of the disparity used to calculate the road surface candidate disparity (step S301). The initial value d0 of the road surface candidate disparity is the initial value of the road surface candidate disparity at the road surface candidate disparity extraction position closest to the stereo camera 10. The road surface candidate disparity extraction position closest to the stereo camera 10 can be set, for example, within a range of 1 to 10 meters from the stereo camera 10.
[0090] like Figure 9 As shown, in the stereo camera 10, the road surface height Y is the height of the stereo camera 10 from the photographed road surface 41 in the vertical direction. In addition, the road surface height Y0 is the height of the installation position of the stereo camera 10 from the road surface 41. Due to the undulations of the road, the road surface height Y may vary depending on the distance from the stereo camera 10. Therefore, the road surface height Y at a position far from the stereo camera 10 is inconsistent with the road surface height Y0 at the installation position of the stereo camera 10. In one embodiment of the multiple embodiments, the optical axes OX of the first camera 11 and the second camera 12 of the stereo camera 10 are arranged parallel to each other and face forward. Figure 9 In this case, Z represents the horizontal distance to a specific road surface position. Let the baseline length of the stereo camera 10 be B, and the longitudinal image size be TOTALv. In this case, the road surface parallax d of the road surface 41 captured at any longitudinal coordinate (v coordinate) is s The relationship with the road surface height Y is given by the following mathematical formula and is independent of the horizontal distance Z.
[0091] d s =B / Y×(v-TOTALv / 2) (1)
[0092] The road parallax d calculated by mathematical formula (1) is s It is also called “geometrically estimated road surface parallax”. Hereinafter, geometrically estimated road surface parallax may be referred to with reference numeral d s To express.
[0093] The road surface candidate disparity initial value d0 is obtained by assuming that the road surface 41 is flat and parallel to the stereo camera 10 and the road surface candidate disparity d is closest to the stereo camera 10. cThe coordinates d0 of the road candidate parallax are calculated based on the optical axis OX between the extraction positions of the first camera 11 and the second camera 12. In this case, the v coordinate at the extraction position of the road candidate parallax closest to the position of the stereo camera 10 on the first parallax image is set to a specific coordinate v0. The coordinate v0 is the initial value of the v coordinate when the road candidate parallax is extracted. The coordinate v0 is between TOTALv / 2 and TOTALv. The coordinate v0 is located at the bottom (the larger side of the v coordinate) within the range of image coordinates capable of calculating the parallax. The coordinate v0 can be TOTALv corresponding to the bottom row in the first parallax image. The road candidate parallax initial value d0 can be determined by substituting v0 and Y0 into v and Y in mathematical formula (1).
[0094] The road surface detection unit 26 calculates the disparity collection threshold for the first row with a vertical v-coordinate of v0 based on the road surface candidate disparity initial value d0 (step S302). A row is an array of horizontally arranged pixels with the same v-coordinate on the first disparity image. The disparity collection threshold includes an upper threshold, which serves as the upper limit for collecting disparities, and a lower threshold, which serves as the lower limit for collecting disparities. The disparity collection threshold is set above and below the road surface candidate disparity initial value d0 based on a predetermined rule to include the road surface candidate disparity initial value d0. Specifically, the road surface disparity is defined as the upper and lower thresholds of the disparity collection threshold when the road surface height Y changes by a predetermined road surface height change ΔY from the state where the road surface candidate disparity initial value d0 is calculated. That is, the lower threshold of the disparity collection threshold is obtained by subtracting the disparity corresponding to the road surface height change ΔY from the road surface candidate disparity initial value d0. The upper threshold of the disparity collection threshold is obtained by adding the disparity corresponding to the road surface height change ΔY to the road surface candidate disparity initial value d0. The specific lower and upper thresholds of the disparity collection threshold can be obtained by changing the value of Y in the mathematical formula (1).
[0095] Next, the road surface detection unit 26 repeats the processing between step S303 and step S307. First, the road surface detection unit 26 processes the line at the bottom of the first parallax image where the v coordinate is v0 (step S303).
[0096] The road surface detection unit 26 collects disparities using the disparity collection threshold (step S304). The road surface detection unit 26 collects disparity pixels having a disparity between the lower threshold and the upper threshold of the disparity collection threshold as candidate road surface disparities d for each disparity pixel arranged in the horizontal direction when the v coordinate is v0 in the first disparity image. cThat is, the road surface detection unit 26 determines that the parallax pixels having the parallax within the predetermined limit range based on the road surface candidate parallax initial value d0 calculated using the mathematical formula (1) are candidates for the parallax pixels representing the correct parallax of the road surface 41. The road surface detection unit 26 determines that the parallax pixels corresponding to the parallax pixels determined as candidates for the parallax pixels are candidates for the road surface parallax d c , where the candidate pixel determined to be a parallax pixel represents the correct parallax of the road surface 41. In this way, the road surface detection unit 26 can reduce the possibility of erroneously determining parallax outside of the road surface 41, such as objects or structures on the road surface 41, as parallax of the road surface 41. This improves the accuracy of detecting the road surface 41.
[0097] When all the disparity pixels when the v coordinate is v0 are determined in step S304, the road surface detection unit 26 calculates the collected road surface candidate disparity d c Take the average to calculate the average road candidate disparity d av , that is, the road surface candidate disparity d c The road surface detection unit 26 can calculate the average value of each road surface candidate disparity d c and its uv coordinates and the average road candidate disparity d when the v coordinate is v0 av Stored in memory 23.
[0098] After step S305, the road surface detection unit 26 calculates the average road surface candidate disparity d when the v coordinate is v0 calculated in step S305. av , the parallax collection threshold is calculated for each parallax pixel of the row above, that is, the row with v coordinate v=v0-1 (step S306). The road surface detection unit 26 changes the road surface height Y so that the average road surface candidate parallax d calculated in step S305 when the v coordinate is v0 is av The road surface detection unit 26 calculates the geometrically estimated road surface parallax d when the v coordinate is v0-1 by substituting v0-1 into the road surface height Y instead of v0 in the road surface detection unit 26. s Similar to step S302, the road surface detection unit 26 can estimate the road surface parallax d from the geometry. s The parallax obtained by subtracting the parallax corresponding to the predetermined road surface height change ΔY is set as the lower limit threshold of the parallax collection threshold. The road surface detection unit 26 can use the geometrically estimated road surface parallax d s The parallax obtained by adding the parallax corresponding to the predetermined road surface height change amount ΔY is set as the upper limit threshold of the parallax collection threshold.
[0099] After step S306, the road surface detection unit 26 determines the geometrically estimated road surface parallax d calculated by equation (1). sIs it greater than a predetermined value? The predetermined value is, for example, one pixel. The road surface detection unit 26 geometrically estimates the road surface parallax d s If it is greater than 1, the process returns to step S303 (step S307). In step S303, the road surface detection unit 26 calculates the road surface candidate parallax d c The extracted object moves to a row above one pixel. That is, the road surface candidate disparity d c When the extraction target is in the row with v coordinate v0, the road surface detection unit 26 changes the v coordinate of the row of the road surface detection target to v0-1. Figure 10 As shown, the road candidate disparity d c When the calculation target of is the nth row, the road surface detection unit 26 changes the row of the road surface detection target to the n+1th row. Figure 10 For ease of illustration, the vertical width of each row is exaggerated. In reality, each row is one pixel high. In this case, the v coordinate of row (n+1) is 1 less than the v coordinate of row (n).
[0100] Each process of steps S304 to S306 regarding the n+1th row is the same as the process performed on the row when the v coordinate is v0. In step S304, the road surface detection unit 26 collects the road surface candidate disparity d using the disparity collection threshold value calculated for the nth row in step S306. c In step S305, the road surface detection unit 26 collects the road surface candidate disparity d c Take the average to calculate the average road candidate disparity d av In step S306, the road surface detection unit 26 uses the average road surface candidate parallax d av The road surface height Y in the mathematical formula (1) is changed. The road surface detection unit 26 calculates the geometrically estimated road surface parallax d using the mathematical formula (1) with the changed road surface height Y. s Furthermore, in order to extract the road surface candidate disparity d in the n+2th row, c The road surface detection unit 26 estimates the road surface parallax d in consideration of the geometry. s The parallax collection threshold is calculated based on the road height change ΔY in .
[0101] The road surface detection unit 26 calculates the road surface candidate disparity d c The object to be extracted is the candidate disparity d from the road surface closest to the stereo camera 10. c The rows corresponding to the extraction position of are moved upward (in the negative direction of the v coordinate) and the candidate road disparity d corresponding to the v coordinate is extracted. c The road surface detection unit 26 can extract the road surface candidate disparity d c and the corresponding u coordinate, v coordinate, and the average road surface candidate disparity d corresponding to the v coordinate av Stored together in the memory 23.
[0102] When the geometrically estimated road surface parallax d calculated by equation (1) in step S307 is s If the road surface candidate disparity d is less than the predetermined value, the road surface detection unit 26 ends the process of c Extraction processing and return Figure 7 The predetermined value may be, for example, one pixel.
[0103] As mentioned above, in Figure 8 In the flowchart, the road candidate disparity d is extracted c The initial value of the v coordinate is set to v0 corresponding to the position on the near side as viewed from the stereo camera 10, and the road surface candidate disparity d on the far side is sequentially extracted. c Generally speaking, the stereo camera 10 has a higher parallax detection accuracy at the short distance side than at the long distance side. Therefore, by sequentially extracting the road surface candidate parallax d from the short distance side to the long distance side, c , which can improve the detected road candidate disparity d c precision.
[0104] In the above extraction of the candidate road disparity d c of Figure 8 In the flowchart, the road candidate disparity d is calculated for each coordinate in the longitudinal direction. c In other words, when extracting the road candidate disparity d c In the flowchart, the road candidate disparity d is calculated for each row of pixels in the vertical direction. c The calculation unit of the road candidate disparity is not limited to this. The road candidate disparity d may also be calculated in units of multiple coordinates in the longitudinal direction. c .
[0105] The road surface candidate disparity d in steps S301 to S307 c After the extraction process, the road surface detection unit 26 enters Figure 7 The road surface detection unit 26 estimates the road surface parallax d in order from the short distance side to the long distance side. r When the road parallax d r The Kalman filter is applied sequentially. Therefore, the road surface detection unit 26 initializes the Kalman filter (step S202). As the initial value of the Kalman filter, the road surface parallax d calculated in step S305 can be used. r The average road surface candidate disparity d corresponding to the lowest row (row with v coordinate value v0) among the estimated rows av The value of .
[0106] The road surface detection unit 26 changes the target row from the short distance side to the long distance side of the road surface 41 and sequentially executes the following processes of steps S203 to S210 (step S203 ).
[0107] First, the road surface detection unit 26 detects the target line in the first parallax image based on the road surface candidate parallax d within a certain width range in the real space. c , generating the road parallax d r The range of a certain width in the actual space is a range that takes into account the width of the lane of the road. The certain width can be, for example, a value such as 2.5m or 3.5m. The range for obtaining the parallax is initially set to, for example, Figure 11 The range surrounded by the solid frame line 50 in the image is pre-stored in the memory 23 of the object detection device 20. By limiting the range of obtaining parallax to this range, the possibility of the road surface detection unit 26 mistakenly identifying objects other than the road surface 41 or structures such as soundproof walls as the road surface 41 and extracting them is reduced. As a result, the accuracy of road surface detection can be improved. As will be described later, Figure 11 The parallax acquisition range indicated by the solid line in the figure can be changed sequentially from the initially set frame line 50 according to the situation on the road ahead.
[0108] The road surface detection unit 26 uses the road surface parallax d r The predicted value of the road surface parallax d is set for the target line r The acquisition range of road parallax d r The acquisition range is based on the Kalman filter to predict the road parallax d of the next line r The reliability is calculated based on the range of the reliability. The reliability is calculated based on the variance σ of the Gaussian distribution. 2 (σ is the road parallax d r The road surface detection unit 26 can calculate the road surface parallax d based on the predicted value ±2σ, etc. r The road surface detection unit 26 obtains the road surface candidate disparity d generated in step S204 from the road surface candidate disparity d c The histogram of the road disparity d is extracted in the Kalman filter-based setting r The road parallax d with the maximum frequency within the acquisition range r The road surface detection unit 26 extracts the road surface parallax d r The road parallax d as the target line r Observation value (step S205).
[0109] Next, the road surface detection unit 26 confirms the road surface parallax d determined in step S205. r is the correct road parallax d rThe road surface detection unit 26 generates a dv correlation map in which all road surface parallaxes d detected in each row up to the row currently being processed are converted into a dv correlation map. r Associated with the road parallax d r and v coordinates as the dv coordinate space of the coordinate axis. When the correct road surface 41 is detected, the dv correlation diagram is as follows Figure 12 As shown by the dotted line in the figure, as the value of v coordinate decreases, the road parallax d r It also decreases linearly.
[0110] On the other hand, when the parallax of the object is mistakenly recognized as the road surface 41, as shown in FIG. Figure 13 As shown in the dv correlation diagram, the parallax d at the portion representing the parallax of the object is almost constant regardless of the change in the longitudinal coordinate (v coordinate). Generally speaking, since the object includes a portion perpendicular to the road surface 41, it is displayed on the first parallax image as containing a large amount of parallax at equal distances. Figure 13 In the first portion R1, parallax d decreases as the v coordinate value changes. The first portion R1 represents the portion where the road surface 41 is correctly detected. In the second portion R2, parallax d remains constant even when the v coordinate changes. The second portion R2 is considered to be where an object is incorrectly detected. The road surface detection unit 26 can determine that an object has been incorrectly detected when a predetermined number of lines with approximately equal parallax d values persist.
[0111] If it is determined in step S206 that the parallax is not the correct road parallax d r (Step S206: No), the road surface detection unit 26 starts searching again for the road surface parallax d starting from the row where the object is determined to be erroneously recognized. r (Step S207). In step S207, the road surface detection unit 26 re-searches the road surface parallax histogram in the area of the row where the parallax d does not change even when the v coordinate value changes. If there is a high-frequency parallax in the area where the parallax is smaller than the parallax d determined in step S205, the road surface detection unit 26 can determine that the parallax is the correct road surface parallax d. r Observed value of .
[0112] If it is determined in step S206 that the road parallax d r is correct (step S206: yes), or the road parallax d is completed in step S207. r If the road surface detection unit 26 does not re-search, it proceeds to step S208. In step S208, the road surface detection unit 26 determines the lateral range of the road surface 41 on the first parallax image, which is the target of generating the next row of histograms that are offset by one pixel in the vertical direction. Figure 11As shown, when there is another vehicle 42 on the road surface 41, the road surface detection unit 26 cannot obtain the correct road surface parallax d at the portion of the road surface 41 that overlaps with the other vehicle 42. r When the road parallax d can be obtained in the road surface 41 r When the range of becomes narrow, it is difficult for the road surface detection unit 26 to obtain accurate road surface parallax d r Therefore, if Figure 11 As shown by the dotted line in FIG, the road surface detection unit 26 sequentially changes the number of pixels for obtaining the candidate road surface disparity d in the horizontal direction. c Specifically, when the road surface detection unit 26 determines that an object is included in step S206, it detects which side of the object in the lateral direction represents the correct road surface parallax d r The candidate disparity of the road surface d c In the next row, the range of parallax is shifted to the horizontal direction with more parallax d indicating the correct road surface. r The candidate disparity of the road surface d c One side ( Figure 11 to the right in the middle).
[0113] Then, the road surface detection unit 26 uses the road surface parallax d of the current line determined in step S205 or S207 to calculate the road surface parallax d of the current line. r To update the Kalman filter (step S209). That is, the Kalman filter is based on the road parallax d of the current row. r The observation value is used to calculate the road parallax d r When the estimated value of the current line is calculated, the road surface detection unit 26 calculates the road surface parallax d of the current line. r The estimated value of is added as part of the past data and used to calculate the road parallax d of the next line. r The height of the road surface 41 is assumed not to change suddenly relative to the horizontal distance Z from the stereo camera 10. Therefore, in the estimation using the Kalman filter in this embodiment, the road surface parallax d of the next row is estimated. r The road parallax d present in the current row r In this way, the road surface detection unit 26 limits the range of the histogram parallax used to generate the next row to the road surface parallax d of the current row. r The possibility of erroneously recognizing objects other than the road surface 41 is reduced. In addition, the amount of calculation performed by the road surface detection unit 26 can be reduced, thereby increasing the processing speed.
[0114] When the road surface parallax d estimated by the Kalman filter in step S209 r If the road parallax d estimated by the Kalman filter is greater than the predetermined value, the road surface detection unit 26 returns to step S203 and repeats the processing of steps S203 to S209. rIf the value is less than or equal to a predetermined value (step S210), the road surface detection unit 26 proceeds to the next process (step S211). The predetermined value may be, for example, one pixel.
[0115] In step S211, the road surface detection unit 26 maps the longitudinal image coordinate v and the estimated road surface parallax d on the dv correlation map. r The relationship between is approximated by two straight lines. r The value of the v coordinate is related to the distance z from the stereo camera 10 and the road surface height Y. Therefore, the v coordinate and the road surface parallax d r The relationship between the two straight lines can be considered to be the relationship between the distance to the stereo camera 10 and the height of the road surface 41 approximated by two straight lines. Figure 14 The flowchart shows the processing of step S211 in detail.
[0116] First, through the Figure 7 The process up to step S210 obtains the road parallax d r and the correlation between the v coordinate. For example, the v coordinate and the road parallax d r The correlation between Figure 15 The dashed line graph 51 is shown in the dv coordinate space. In real space, if the road surface 41 is flat and its inclination does not change, the graph 51 is a straight line. However, in real road surface 41, the inclination of road surface 41 may change due to ups and downs. If the inclination of road surface 41 changes, the graph 51 in the dv coordinate space cannot be represented by a straight line. If the inclination change of road surface 41 is approximated by three or more straight lines or curves, the processing load of object detection device 20 increases. Therefore, in this application, the graph 51 is approximated by two straight lines.
[0117] like Figure 15 As shown, the road surface detection unit 26 uses the first straight line 52 to calculate the estimated road surface parallax d on the lower side (short distance side) in the dv coordinate space by the least square method. r Approximation processing is performed (step S401). Approximation using the first straight line 52 enables the object detection device 20 to detect the object within the distance range up to the road surface parallax d corresponding to the predetermined distance. r The predetermined distance may be half the distance range within which object detection device 20 is to detect objects. For example, if object detection device 20 is designed to detect objects up to 100 meters ahead, first straight line 52 may be determined by the least squares method to be closest to figure 51 within a range from the shortest distance measurable by stereo camera 10 to 50 meters.
[0118] Next, the road surface detection unit 26 determines whether the inclination of the road surface 41 represented by the first straight line 52 approximated in step S401 is a possible inclination of the road surface 41 (step S402). The inclination angle of the first straight line 52 becomes a plane when converted to real space. The inclination of the first straight line 52 corresponds to the inclination angle of the road surface 41 in the YZ plane, determined based on conditions such as the road surface height Y0 at the installation location of the stereo camera 10 and the baseline length B. If the inclination of the road surface 41 in real space corresponding to the first straight line 52 is within a predetermined angle range relative to the horizontal plane in real space, the road surface detection unit 26 can determine that the inclination of the road surface 41 is a possible inclination. If the inclination of the road surface 41 in real space corresponding to the first straight line 52 is outside the predetermined angle range relative to the horizontal plane in real space, the road surface detection unit 26 can determine that the inclination of the road surface 41 is not a possible inclination. The predetermined angle can be appropriately set based on consideration of the driving environment of the mobile object 30.
[0119] If it is determined in step S402 that the inclination of the first straight line 52 is an inclination that cannot exist as the road surface 41 (step S402: No), the road surface detection unit 26 determines the first straight line 52 based on a theoretical road surface assuming that the road surface 41 is flat (step S403). The theoretical road surface can be calculated based on the installation conditions such as the road surface height Y0 at the installation position of the stereo camera 10, the installation angle, and the baseline length B. The road surface detection unit 26 can calculate the road surface parallax d calculated from the image. r For example, the road surface detection unit 26 may mistakenly use the parallax of an object or structure other than the road surface 41 as the road surface parallax d r In the case of extraction, it is determined that the road surface 41 has an inclination that cannot exist, thereby eliminating the error. In this way, it is possible to reduce the error of misjudging the parallax of objects or structures other than the road surface 41 as the road surface parallax d r possibility.
[0120] If, in step S402, the road surface detection unit 26 determines that the inclination of the first straight line 52 is a possible inclination of the road surface 41 (step S402: Yes), or after step S403, the road surface detection unit 26 proceeds to step S404. In step S404, the road surface detection unit 26 determines an approximation starting point 53 for approximating the second straight line 55. The road surface detection unit 26 can calculate the approximation error with the graph 51 from the smallest v coordinate of the first straight line 52 (the far side) toward the larger v coordinate (the near side), and select as the approximation starting point 53 the coordinate on the first straight line 52 where the approximation error continuously decreases to a predetermined value. Alternatively, the road surface detection unit 26 can calculate the approximation error with the graph 51 from the largest v coordinate of the first straight line 52 (the near side) toward the smaller v coordinate (the far side), and select as the approximation starting point 53 the coordinate on the first straight line 52 where the approximation error exceeds a predetermined value. The v coordinate of the approximation starting point 53 is not fixed to a specific value. Approximation starting point 53 can be set on first straight line 52 at a position corresponding to a v-coordinate closer to the stereo camera 10 than half the distance range within which object detection device 20 performs object detection. For example, when first straight line 52 approximates road surface 41 within a range from the closest measurable distance to 50 meters ahead, approximation starting point 53 can be set at a position corresponding to a v-coordinate 40 meters ahead of 50 meters.
[0121] After step S404, the road surface detection unit 26 repeatedly executes steps S405 to S407. Figure 16 As shown, the road surface detection unit 26 uses the angle difference from the first straight line 52 as an angle selected from a predetermined angle range to sequentially select candidates for the second straight line 55 starting from the approximation starting point 53 (step S405). The predetermined angle range is set to the angle within which the road slope can change within the distance range of the measurement target. The predetermined angle range can be, for example, ±3 degrees. For example, the road surface detection unit 26 may increment the angle of the candidate straight line 54 by 0.001 degrees, starting from -3 degrees from the angle of the first straight line 52 to +3 degrees from the angle of the first straight line 52.
[0122] For each selected candidate straight line 54, the road surface detection unit 26 calculates the error between the selected candidate straight line 54 and the upper (far-distance) portion of the approximate starting point 53 of the graph 51 in the dv coordinate space (step S406). The error can be calculated as the mean square error of the parallax d relative to the v coordinate. The road surface detection unit 26 may store the error calculated for each candidate straight line 54 in the memory 23.
[0123] After the road surface detection unit 26 completes the error calculation for all candidate straight lines 54 within the angle range (step S407), it searches for the minimum error among the errors stored in the memory 23. Figure 17 As shown, the road surface detection unit 26 selects the candidate straight line 54 having the smallest error as the second straight line 55 (step S408).
[0124] When the second straight line 55 is determined in step S408, the road surface detection unit 26 determines whether the error between the second straight line 55 and the graph 51 is within a predetermined value (step S409). The predetermined value may be appropriately set to obtain desired road surface estimation accuracy.
[0125] In step S409, if the error is within the predetermined value (step S409: Yes), the road parallax d r Approximation processing is performed using the first straight line 52 and the second straight line 55 .
[0126] In step S409, if the error exceeds the predetermined value (step S409: No), the road surface detection unit 26 extends the first straight line 52 upward (far distance side) and rewrites the approximation result (step S410). r Use two straight lines for approximation.
[0127] The road parallax d relative to the v coordinate is approximated by using two straight lines r , the road surface position is approximated by two straight lines. This reduces the subsequent computational load and speeds up object detection compared to approximating the road surface position using a curve or three or more straight lines. Furthermore, the error with the actual road surface is smaller than when the road surface is approximated using a single straight line. Furthermore, by not fixing the v coordinate of the approximation starting point 53 of the second straight line 55 at a predetermined coordinate, the accuracy of the approximation with the actual road surface can be improved compared to when the coordinates of the approximation starting point 53 are fixed in advance.
[0128] When the error is within the predetermined value in step S409 (step S409: Yes) or after step S410, the road parallax d r The linear approximation process is completed and returns to Figure 7 Step S212.
[0129] In step S212, the road parallax d removed from the first parallax image is determined. r The road parallax d removed from the first parallax image is r The threshold value corresponds to the first height described later. The first height can be calculated so that the road parallax d is removed in the subsequent processing of step S103. r .
[0130] Then, the processing of the image processing unit 22 returns to Figure 5The unnecessary parallax removal unit 27 of the image processing unit 22 obtains from the road surface detection unit 26 the image of the v coordinate in the dv coordinate space and the road surface parallax d approximated by two straight lines. r The approximate formula of the relationship between v coordinate in dv coordinate space and road parallax d r The relationship between the front distance Z of the stereo camera 10 and the road surface height Y in the real space can be obtained by using an approximate expression for the relationship. The unnecessary parallax removal unit 27 performs the second process (step S103) based on the approximate expression. The second process is to remove the parallax pixels corresponding to the subject whose height from the road surface 41 is less than the first height and the parallax pixels corresponding to the subject whose height from the road surface 41 is more than the second height in the real space from the first parallax image. In this way, the unnecessary parallax removal unit 27 removes the parallax pixels corresponding to the subject whose height from the road surface 41 is less than the second height in the real space. Figure 6 The first parallax image shown is generated Figure 18 The second parallax image is shown. Figure 18 The actual second parallax image based on the image acquired from the stereo camera 10 is as shown in FIG. Figure 19 As shown. Figure 19 In the image, the size of the parallax is expressed by the shades of black and white.
[0131] The first height can be set to, for example, a value greater than 15 cm and less than 50 cm. If the first height is set to less than 15 cm, the parallax removal unit 27 may easily detect objects other than objects on the road due to factors such as unevenness of the road surface 41 and / or errors in the approximation formula. This may result in detection errors or reduced detection speed. Furthermore, if the first height is greater than 50 cm, the parallax removal unit 27 may not be able to detect children and / or large obstacles on the road surface 41.
[0132] The second height is set based on the upper limit of vehicle heights permitted on the road. This permitted vehicle height is regulated by traffic laws. For example, under Japan's Road Traffic Law, truck heights are generally limited to 3.8 meters or less. For example, the second height could be 4 meters. If the second height is higher than 4 meters, unnecessary objects, such as overhead structures like traffic lights and information display boards, may be detected.
[0133] By removing unnecessary parallax, the unnecessary parallax removal unit 27 can pre-remove parallax for subjects other than the road surface 41 or objects on the road before the subsequent object detection process (third and fourth processes). This improves the accuracy of object detection. Furthermore, by removing unnecessary parallax, the amount of computation associated with parallax unrelated to objects on the road is reduced, thereby speeding up processing. Therefore, the object detection device 20 of the present disclosure can improve the performance of the object detection process.
[0134] The unnecessary parallax removal unit 27 transmits the second parallax image to the clustering unit 28. Based on the second parallax image and the approximate position of the road surface 41 calculated by the road surface detection unit 26, the clustering unit 28 performs a third process (step S104) to determine the object parallax associated with the object for each horizontal u-coordinate or range of u-coordinates. Specifically, the clustering unit 28 generates a histogram representing the number of pixels for each parallax for each horizontal u-coordinate range in the second parallax image. A u-coordinate range is a range that includes one or more horizontal pixels.
[0135] like Figure 20 As shown, the clustering unit 28 extracts a longitudinal region having one or more horizontal pixel widths Δu from the two-parallax image. The longitudinal region is used to determine whether there is an object parallax associated with the object in the region and to obtain distance information corresponding to the object parallax. Therefore, if the horizontal width Δu of the longitudinal region is made thinner, the resolution of detecting horizontal objects becomes higher. Figure 20 For ease of explanation, the width of Δu is shown as relatively wide, but Δu can be anywhere from one pixel to several pixels. Vertical regions are sequentially acquired from one horizontal end of the second parallax image to the other, and the following processing is performed to detect object parallax across the entire horizontal area of the second parallax image.
[0136] Figure 21 It is a diagram showing an example of a histogram. Figure 21 The horizontal axis is the disparity d expressed in pixels. The disparity d is larger on the left side of the horizontal axis and gradually decreases towards the right. The minimum value of the disparity d can be, for example, one pixel or less. The larger the disparity, the finer the resolution of the distance, and the smaller the disparity, the coarser the resolution of the distance. Therefore, Figure 21 The horizontal axis of the histogram gathers more disparity d on the side with larger disparity. Figure 21 The vertical axis of the histogram represents the number of occurrences of disparity pixels with the horizontal axis representing disparity d.
[0137] exist Figure 21 In FIG, a threshold curve is also shown, which represents a threshold value for determining whether each disparity d is an object disparity. Disparity d can be a representative value of an interval of disparity having a width. When the number of occurrences of pixels of each disparity d exceeds the threshold curve, it means that a predetermined number or more of disparity pixels representing the same distance specified by the threshold curve are included in a longitudinal area of width Δu. Figure 21In the figure, the area with slashes (columnar portion) exceeds the threshold curve. The threshold curve can be set to the number of occurrences (number of pixels) corresponding to a predetermined height in the y direction in the actual space for each parallax d. For example, the predetermined height can be 50 cm. When the parallax is small at a long distance, the object displayed on the image is smaller than the object at a close distance. Therefore, as the parallax becomes smaller, the value of the number of occurrences on the vertical axis of the threshold curve becomes smaller. When the number of occurrences exceeds the predetermined threshold corresponding to the parallax d, the clustering unit 28 determines the parallax d corresponding to the pixel as the object parallax d. e .
[0138] Then, the clustering unit 28 classifies the object based on the parallax d e The distribution of the road surface 41 and the position estimated in step S102 are used to calculate the parallax d with the object. e Related height information (step S105). The process of step S105 may be included in the third process.
[0139] Specifically, it is assumed that the image acquired by the first camera 11 or the second camera 12 includes the following Figure 22 The portion of the second parallax image corresponding to the portion surrounded by the frame line 61 of the other vehicle 42 is shown in FIG. Figure 23 Shown enlarged.
[0140] The clustering unit 28 is for objects with parallax d e The u coordinate is based on the object parallax d e The distance information and the road surface position estimated in step S102 are used to calculate the parallax d with the object. e When there is an object on the road surface 41, there is an object parallax d e The parallax pixels of d are arranged above the estimated road surface position on the second parallax image. The clustering unit 28 scans the parallax pixels of the second parallax image upward from the road surface position of the u coordinate to detect the object with parallax d. e The clustering unit 28 is based on the distribution of the disparity pixels of the same object in the vertical direction (v coordinate direction). e The height information on the second disparity image is determined by the number or distribution of the disparity pixels arranged. e Even when the parallax pixels are partially interrupted in the vertical direction, the clustering unit 28 can determine the height information based on a predetermined determination criterion.
[0141] The clustering unit 28 may associate the object disparity d with respect to each range of the horizontal coordinate (u coordinate) including one or more coordinates. e and height information, and stores it in the memory 23. Figure 24As shown in one example, the clustering unit 28 can store multiple object disparities d e The clustering unit 28 represents the point cluster distribution in a two-dimensional space (ud coordinate space) with the u coordinate and the parallax d as the horizontal and vertical axes. e The information is transmitted to the grouping unit 29.
[0142] The grouping unit 29 divides the object parallax d in the ud coordinate space into e The information is transformed into a point group in the xz coordinate space (real space) of the world coordinate system, and the point group (group) is extracted to perform the process of detecting the object, that is, the fourth process (step S106). The grouping unit 29 collects multiple adjacent points according to predetermined conditions and extracts them into a point group. Figures 25 to 32 An example of the processing performed by the grouping unit 29 will be described.
[0143] exist Figure 25 In the figure, a moving body 30 equipped with the object detection system 1 and other vehicles 42, 43, 44, a side wall 45, and a street light 46 are included, which are traveling on a road surface 41. Figure 25 In FIG, the moving object 30 is a vehicle.
[0144] Figure 26 4 is a flowchart showing the operation of the grouping unit 29. In step S501, the grouping unit 29 of the object detection device 20 mounted on the mobile body 30 Figure 24 The disparity d of multiple objects in the ud coordinate space shown e Transformed into Figure 27 The point group in the actual space (xz coordinate space) shown in . Figure 27 In the xz coordinate space, each point indicates whether an object exists. That is, the presence of a point means the object exists, and the absence of a point means the object does not exist. Figure 27 Corresponding to Figure 25 , other vehicles 42, 43, 44, side walls 45, and streetlights 46 are shown as point clusters. The grouping unit 29 groups each point with the object disparity d obtained by the clustering unit 28. e The x direction of the world coordinate system is a direction perpendicular to the direction of travel of the mobile body 30. The z direction of the world coordinate system is a direction horizontal to the direction of travel of the mobile body 30.
[0145] In step S502, the grouping unit 29 extracts point clusters that are continuously arranged in the z direction from the point cluster in the xz coordinate space. For example, if the number of points in the z direction is continuous for a predetermined number or more (a predetermined length or more in real space), the grouping unit 29 extracts this point cluster. The grouping unit 29 then deletes the extracted point cluster from the point cluster in the xz coordinate space. Figure 28An example of a point group 71 and a point group 72 continuously arranged in the z direction extracted by the grouping unit 29 is shown. Figure 29 An example of a point group in the xz coordinate space after the grouping unit 29 deletes the point group 71 and the point group 72 that are continuously arranged in the z direction from the xz coordinate space is shown.
[0146] Point groups arranged continuously in the z direction indicate the presence of objects estimated to be parallel to the traveling direction of the moving object 30. Objects parallel to the traveling direction of the moving object 30 include roadside structures such as guardrails and side walls, and the sides of other vehicles. Figure 28 The point group 71 in Figure 25 The side of the body of the other vehicle 44. Figure 28 Point group 72 in the representation Figure 25 The grouping unit 29 can reduce the detection rate of objects parallel to the traveling direction by performing the process of step S502.
[0147] In step S503, the grouping unit 29 extracts and combines point groups that are continuous in the x-direction from the point groups in the xz coordinate space. For example, if the number of points in the x-direction is continuous for a predetermined number or more (a predetermined length in real space or more), the grouping unit 29 extracts point groups that are continuous in the x-direction. The grouping unit 29 then combines the point groups that are continuous in the x-direction with those in the z-direction and determines a quadrilateral region (combined region) that includes the set of point groups. Figure 30 FIG. 2 shows an example of a combined region determined by the grouping unit 29. It should be noted that the grouping unit 29 may divide the combined region when the width of the combined region in the x direction exceeds a threshold.
[0148] The combined area indicates the presence of an object estimated to be perpendicular to the traveling direction of the moving object 30. The object perpendicular to the traveling direction of the moving object 30 is a structure on the road such as a streetlight or the back of another vehicle. Figure 30 The binding regions 73-1 to 73-4 represent Figure 25 A portion of the side wall 45 in. Figure 30 The binding region 74 represents Figure 25 The back of the vehicle body of the other vehicle 42. Figure 30 The binding region 75 in Figure 25 However, the combined area 75 shows only the portion of the back of the vehicle body of the other vehicle 43 that is not blocked by the other vehicle 42 when viewed from the mobile body 30. Figure 30 The bonding area 76 shows Figure 25 The back of the other vehicle 44. Figure 30 The bonding area 77 in FIG. Figure 25The streetlight 46 in FIG. 4 is thus detected by the grouping unit 29 through the processing of steps S501 to S503.
[0149] It should be noted that the extraction of the point group continuously arranged in the z direction and the extraction of the point group continuously arranged in the x direction can be performed by any method. Figure 31A and Figure 31B , which illustrates an example of extracting a point group that is continuously arranged in the z direction. A cell in the figure is a quantized unit in the xz coordinate space. Here, the window size in the x direction is set to three cells. The grouping unit 29 starts counting when there is a point in the window size, for example, Figure 31A As shown in , when there are more than i consecutive points in the unit cell, these points are determined to be a point group that is continuously arranged in the z direction. Figure 31B As shown in FIG, even if there is a window without points after counting starts, as long as the window without points is continuous in the z direction and within j cells, the grouping unit 29 can consider the points to be continuous. In this case, for example, when it is considered that the points are continuous in the z direction and are within k cells or more, the grouping unit 29 determines that these points are a point group continuously arranged in the z direction.
[0150] In step S504, the grouping unit 29 groups the object parallax d associated with each point in the xz coordinate space. e The distance from the stereo camera 10 mounted on the mobile body 30 to an object perpendicular to the direction of travel of the mobile body 30 (the object detected by the processing of steps S501 to S503) is determined. For example, the grouping unit 29 calculates the average value of the distances from the stereo camera 10 to the points of the longest continuous point cluster in the x-direction among the set of point clusters included in the combined area as the distance to the object perpendicular to the direction of travel of the mobile body 30.
[0151] Furthermore, in step S504, the grouping unit 29 determines the height of an object perpendicular to the direction of travel of the mobile body 30 based on the height information associated with each point in the xz coordinate space. In the point group in the xz coordinate space, a point with a high height indicated by the height information may represent an aerial object, etc., or may also be noise. Here, the grouping unit 29 may not use the point with a high height indicated by the height information in the point group in the xz coordinate space to determine the height of the object. For example, the grouping unit 29 may rearrange the heights of each point in the set of point groups included in the combined area in order of height, and set the predetermined sequential heights as the heights of the object perpendicular to the direction of travel of the mobile body 30. Thus, the grouping unit 29 can determine the heights of objects other than objects such as aerial objects that are taller than the mobile body 30. Furthermore, the influence of noise can be reduced, and the height of the object can be determined with high accuracy.
[0152] Furthermore, grouping unit 29 can determine the lateral width of the object based on the distribution of the point cluster in the x-direction. In other words, grouping unit 29 can determine the lateral width of the object based on the width of the combined area in the x-direction. Thus, grouping unit 29 can identify the position, lateral width, and height of the identified object in the x-z coordinate space.
[0153] The objects represented by the combined area determined in step S503 may include a moving object, a side wall higher than the moving object, and other stationary objects. Figure 32 As shown, when another vehicle 47 is traveling on a highway, there is a possibility that both the other vehicle 47 and the highway's soundproofing wall 48 are displayed as a single combined area. When a single combined area represents both the other vehicle 47 and the highway's soundproofing wall 48, the heights corresponding to the two ends of the combined area in the x-direction in the xz coordinate space are compared, with one being the height of the other vehicle 47 and the other being the height of the soundproofing wall 48. In step S504, the grouping unit 29 calculates the average height of the point cluster associated with a predetermined percentage (e.g., 30%) of the combined area at each end of the x-direction. If the difference between the two average heights exceeds a threshold (e.g., several meters), the lower value is determined as the height of the object. In other words, if the height difference between the two ends of the detected object exceeds the threshold, the grouping unit 29 determines the lower height as the object's height. This allows the grouping unit 29 to determine the height of the moving object even when a moving object and a stationary object taller than the moving object are detected as a single object.
[0154] The image processing unit 22 can output the position, lateral width, and height information of the object identified by the grouping unit 29 to other devices in the mobile body 30 through the output unit 24 (step S107). For example, the image processing unit 22 can output this information to a display device in the mobile body 30. Figure 33 As shown, the display device in the mobile body 30 displays a frame line surrounding the image of the other vehicle 42 on the image of the first camera 11 or the second camera 12 based on the information acquired from the object detection device 20 . Figure 33 The frame lines in represent the position of the detected objects and the range they occupy in the image.
[0155] As described above, the object detection device 20 of the present disclosure can achieve fast processing speed and high-precision object detection. In other words, the object detection device 20 of the present disclosure can improve object detection performance. Furthermore, the object detection device 20 does not limit the objects to be detected to specific types. The object detection device 20 can detect any object on the road surface. The image processing unit 22 of the object detection device 20 can perform the first, second, third, and fourth processes without using information from the images captured by the stereo camera 10 other than the first parallax image. Therefore, the object detection device 20 does not need to perform processing to separately identify objects from the captured images, in addition to processing the first and second parallax images. Consequently, the object detection device 20 of the present disclosure can reduce the processing load on the image processing unit 22 involved in object recognition. This does not preclude the object detection device 20 of the present disclosure from being combined with image processing performed directly on images obtained from the first camera 11 or the second camera 12. The object detection device 20 can also be combined with image processing techniques such as template matching.
[0156] In the above description of the processing performed by the image processing unit 22, to facilitate understanding of the present disclosure, the processing includes determinations and operations using various images. This processing using images does not necessarily include actual image rendering. Processing having substantially the same content as that using these images is performed through information processing within the image processing unit 22.
[0157] Although the embodiments according to the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art can easily make various deformations or modifications based on the present disclosure. Therefore, it should be noted that these deformations or modifications are included in the scope of the present disclosure. For example, the functions included in each component or each step, etc. can be reconfigured in a logically non-contradictory manner, and multiple components or steps, etc. can be combined into one or divided. Although the embodiments according to the present disclosure are described with the device as the center, the embodiments according to the present disclosure can also be implemented as a method including steps performed by each component of the device. The embodiments according to the present disclosure can also be implemented as a method, a program, or a storage medium with a program recorded thereon, which is executed by a processor included in the device. It should be understood that the scope of the present disclosure also includes these.
[0158] In the present disclosure, the descriptions such as "first" and "second" are identifiers for distinguishing the configurations. The configurations distinguished by the descriptions such as "first" and "second" in the present disclosure can have their numbers swapped. For example, "first," which is an identifier for the first lens, can be replaced with "second," which is an identifier for the second lens. The exchange of identifiers is performed simultaneously. The configurations can be distinguished even after the identifiers are swapped. Identifiers can be deleted. Configurations from which identifiers have been deleted are distinguished by reference numerals. The descriptions based solely on identifiers such as "first" and "second" in the present disclosure cannot be used to explain the order of the configurations or to justify the existence of identifiers with smaller numbers.
[0159] In the disclosure, the x-direction, y-direction, and z-direction are provided for ease of description and are interchangeable. The components of the present disclosure are described using an orthogonal coordinate system with the x-direction, y-direction, and z-direction as the directions of the respective axes. The positional relationship of the components of the present disclosure is not limited to being in an orthogonal relationship. The u-coordinate and v-coordinate representing the coordinates of the image are provided for ease of description and are interchangeable. The origin and direction of the u-coordinate and the v-coordinate are not limited to those described in the disclosure.
[0160] In the above embodiment, the first camera 11 and the second camera 12 of the stereo camera 10 are arranged side by side in the x-direction. The arrangement of the first camera 11 and the second camera 12 is not limited to this. The first camera 11 and the second camera 12 may be arranged side by side in a direction perpendicular to the road surface (y-direction) or in a direction inclined relative to the road surface 41. The number of cameras constituting the stereo camera 10 is not limited to two. The stereo camera 10 may include three or more cameras. For example, by using two cameras arranged horizontally on the road surface and two cameras arranged vertically, i.e., a total of four cameras, more accurate distance information can be obtained.
[0161] In the above embodiment, the stereo camera 10 and the object detection device 20 are mounted on the mobile object 30. However, the stereo camera 10 and the object detection device 20 are not limited to being mounted on the mobile object 30. For example, the stereo camera 10 and the object detection device 20 may be mounted on roadside equipment installed at an intersection, etc., and configured to capture images including the road surface. For example, the roadside equipment may detect a first vehicle approaching from one of the intersecting roads at the intersection and provide a notification of the first vehicle's approach to a second vehicle approaching on the other road.
[0162] Description of reference numerals:
[0163] 1: Object Detection System
[0164] 10: Stereo Camera
[0165] 11: First Camera
[0166] 12: Second Camera
[0167] 20: Object detection device
[0168] 21: Get Department
[0169] 22: Image processing unit (processor)
[0170] 23: Memory
[0171] 24: Output
[0172] 25: Parallax image generation unit
[0173] 26: Road Surface Detection Department
[0174] 27: Unnecessary parallax removal unit
[0175] 28: Clustering Department
[0176] 29: Grouping Department
[0177] 30: Mobile
[0178] 41: Road surface
[0179] 42, 43, 44, 47: Other vehicles (objects)
[0180] 45: Sidewall
[0181] 46: Streetlight
[0182] 48: Soundproof wall
[0183] 50, 61: frame line
[0184] 51: Graphics
[0185] 52: First straight line
[0186] 53: Approximate starting point
[0187] 54: Candidate line
[0188] 55: Second straight line
[0189] 71 and 72 point groups
[0190] 73-1, 73-2, 73-3, 73-4, 74, 75, 76, 77: Binding area
[0191] R1: Part 1
[0192] R2: Part 2
Claims
1. An object detection device, wherein: The device includes a processor configured to perform the following processing: Road surface detection processing for estimating a position of a road surface in real space based on a first disparity map generated based on output from a stereo camera that captures an image containing the road surface, wherein the first disparity map is a map associating disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to a horizontal direction of the image; an object disparity determination process of determining the disparity as an object disparity when the number of occurrences of the disparity in each coordinate range of the first direction of the first disparity map exceeds a predetermined threshold value corresponding to the disparity; and The object detection process transforms the object parallax information into a point group in the xz coordinate space of the world coordinate system, where each point indicates whether the object exists. After deleting the continuous point group in the z direction from the point group in the xz coordinate space, the object is detected by extracting a set of point groups.
2. The object detection device according to claim 1, wherein: The processor is configured to detect the object by extracting a set of point groups continuous in the x-direction from the point group in the xz coordinate space in the object detection process.
3. The object detection device according to claim 1 or 2, wherein: The processor is configured to, after detecting the object in the object detection process, determine a distance from the stereo camera to the object based on the object parallax.
4. The object detection device according to claim 1 or 2, wherein: The processor is composed of: In the object disparity determination process, height information is calculated based on the determined distribution of the object disparity in the second direction on the first disparity map and the estimated position of the road surface. In the object detection process, after the object is detected, the height of the object is determined based on the height information.
5. The object detection device according to claim 4, wherein: The processor is configured to, in the object detection process, determine the lower height as the height of the object when a difference between heights of both ends of the detected object exceeds a threshold.
6. The object detection device according to claim 1 or 2, wherein: The processor is configured to, after detecting the object in the object detection process, determine the width of the object based on the distribution of the set of the point group in the x-direction.
7. The object detection device according to claim 1 or 2, wherein: The processor is further configured to perform the following processing: unnecessary parallax removal processing, based on the position of the road surface estimated by the road surface detection processing, generating a second parallax map obtained by removing the parallax corresponding to the height within a predetermined range from the road surface in real space from the first parallax map; The object disparity determination process is configured to use the second disparity map instead of the first disparity map.
8. An object detection system, wherein: include: A stereo camera captures multiple images with parallax differences between them. An object detection device comprising at least one processor, The processor is configured to generate a first disparity map based on an output of the stereo camera that captures an image including a road surface, the first disparity map being a map associating disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to a horizontal direction of an image captured by the stereo camera. The processor is further configured to perform the following processing: road surface detection processing, estimating the position of the road surface in real space based on the first disparity map; an object disparity determination process of determining the disparity as an object disparity when the number of occurrences of the disparity in each coordinate range of the first direction of the first disparity map exceeds a predetermined threshold value corresponding to the disparity; and Object detection processing transforms the object disparity information in each coordinate range of the first direction into a point group in the xz coordinate space of the world coordinate system, each point indicating whether an object exists. After deleting the continuous point group in the z direction from the point group in the xz coordinate space, the object is detected by extracting a set of point groups.
9. A mobile device comprising an object detection system, wherein: The object detection system includes a stereo camera for capturing a plurality of images having parallax between them and an object detection device having at least one processor. The processor is configured to generate a first disparity map based on an output of the stereo camera that captures an image including a road surface, the first disparity map being a map associating disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to a horizontal direction of an image captured by the stereo camera. The processor is configured to also perform the following processing: road surface detection processing, estimating the position of the road surface in actual space based on the first disparity map; object disparity determination processing, determining the disparity as object disparity when the number of occurrences of disparity in each coordinate range of the first direction of the first disparity map exceeds a predetermined threshold value corresponding to the disparity; and object detection processing, transforming the object disparity information into a point group on the xz coordinate space of the world coordinate system, each point indicating whether an object exists, deleting continuous point groups in the z direction from the point group on the xz coordinate space, and then detecting the object by extracting a set of point groups.
10. A method for object detection, wherein: obtaining or generating a first disparity map, the first disparity map being generated based on output from a stereo camera that captures an image including a road surface, wherein the first disparity map is a map associating disparities obtained from the output of the stereo camera with two-dimensional coordinates consisting of a first direction and a second direction intersecting the first direction, wherein the first direction corresponds to a horizontal direction of the image, executing a road surface detection process of estimating the position of a road surface in real space based on the first disparity map, performing object disparity determination processing of determining the disparity as an object disparity when the number of occurrences of the disparity in each coordinate range of the first direction of the first disparity map exceeds a predetermined threshold corresponding to the disparity, The object detection process is performed by transforming the object disparity information of each coordinate range of the first direction into a point group on the xz coordinate space of the world coordinate system, each point indicating whether the object exists, deleting the continuous point group in the z direction from the point group on the xz coordinate space, and then detecting the object by extracting a set of point groups.
Citation Information
Patent Citations
Vehicle exterior monitoring device
JP1993265547A
Anchor device
JP2019173341A
Object detection device, mobile equipment control system and object detection program
JP2016206801A
Information processing device, imaging device, apparatus control system, movable body, information processing method, and program
JP2018092605A