Object detection device, object detection method, and program

By generating low-resolution images and aggregating feature differences in partial areas, the system addresses processing load and accuracy issues in object detection, ensuring robust obstacle identification for autonomous vehicles.

JP7819072B2Active Publication Date: 2026-02-24HONDA MOTOR CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022154768
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2026-02-24
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

Conventional object detection systems face issues of excessive processing load or insufficient accuracy, particularly in real-time applications like autonomous driving.

Method used

The system generates a low-resolution image from a captured image, defines multiple partial area sets with varying pixel counts, aggregates feature differences between these areas, and extracts points of interest to reduce processing load while maintaining accuracy.

Benefits of technology

This approach allows for efficient object detection with reduced processing load and high detection accuracy, effectively identifying potential obstacles on the road.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007819072000001
    Figure 0007819072000001
  • Figure 0007819072000002
    Figure 0007819072000002
  • Figure 0007819072000003
    Figure 0007819072000003
Patent Text Reader

Abstract

To perform object detection suitably while reducing the processing load.SOLUTION: An object detection device includes: an acquisition unit for acquiring an image of a surface on which a moving object can pass at an inclination with respect to the surface; a low-resolution image generation unit that generates a low-resolution image with a reduced image quality of the captured image; a definition unit that defines a plurality of partial region sets, each of which includes a partial region, where each of the plurality of partial region sets is defined to include a plurality of partial regions in a target region for each partial region set, and the target region is a vertically limited portion of the low-resolution image cut out such that at least a portion thereof does not overlap with any other partial region set with respect to the vertical direction; and an extraction unit that derives an aggregate value by aggregating the differences in feature values between the partial regions, included in each of the plurality of partial region sets, and the surrounding partial regions for each of the partial regions, and extracts a focus point based on the aggregate value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an object detection device, an object detection method, and a program. [Background technology]

[0002] A conventional invention has been disclosed for a driving obstacle detection system that divides the area of ​​an object in a surveillance area, such as on a road, obtained by photographing into blocks, extracts local features for each block, and determines whether or not an obstacle is present based on the extracted local features (Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2019-124986 Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques sometimes result in excessive processing load or insufficient accuracy.

[0005] The present invention has been made in consideration of these circumstances, and one of its objectives is to provide an object detection device, an object detection method, and a program that can perform object detection effectively while reducing the processing load. [Means for solving the problem]

[0006] The object detection device, the object detection method, and the program according to the present invention employ the following configuration. (1): An object detection device according to one embodiment of the present invention includes an acquisition unit that acquires an image of a surface on which a moving object can pass, the image being captured at an angle relative to the surface; a low-resolution image generation unit that generates a low-resolution image by reducing the image quality of the captured image; a definition unit that defines a plurality of partial area sets, each of which includes a partial area, wherein each of the plurality of partial area sets is defined to include a plurality of partial areas in a target area for each partial area set, and the target area is obtained by cutting out a portion of the low-resolution image limited in the vertical direction so that at least a portion of the target area does not overlap with other partial area sets in the vertical direction; and an extraction unit that derives an aggregate value that aggregates the differences in features between the partial areas included in each of the plurality of partial area sets and surrounding partial areas, and extracts a point of interest based on the aggregate value.

[0007] (2): In the aspect (1) above, the definition unit defines the plurality of partial area sets so that the more partial area defined in the partial area sets is closer to the front of the low-resolution image, the greater the number of pixels in the partial area.

[0008] (3): In the aspect (1) above, the extraction unit derives the aggregated value by aggregating the differences in features between the partial regions included in each of the plurality of partial region sets and other adjacent partial regions located above, below, to the left, right, and diagonally.

[0009] (4): In the aspect (3) above, the extraction unit further adds to the aggregated value, for each of the partial regions included in the plurality of partial region sets, the difference in feature values ​​between the partial regions adjacent to each other vertically, the difference in feature values ​​between the partial regions adjacent to each other horizontally, and the difference in feature values ​​between the partial regions adjacent to each other diagonally.

[0010] (5): In the above aspect (1), the present invention further includes a high-resolution processing unit that performs high-resolution processing on the target area in the captured image to determine whether an object on the road is an object with which the moving body should avoid contact.

[0011] (6): In the aspect (1) above, the object detection device is mounted on a moving body, and the definition unit changes the aspect ratio of the partial region based on the environment in which the moving body is located.

[0012] (7): In the aspect (6) above, when the speed of the moving body is greater than a reference speed, the definition unit changes the aspect ratio of the partial region to be vertically longer than when the speed of the moving body is equal to or less than the reference speed.

[0013] (8): In the aspect (6) above, when the rotation angle of the moving body is greater than a reference angle, the definition unit changes the aspect ratio of the partial region to be horizontally longer than when the rotation angle of the moving body is equal to or smaller than the reference angle.

[0014] (9): In the aspect (6) above, when the moving body is on a road surface with an uphill gradient of a predetermined gradient or more, the definition unit changes the aspect ratio of the partial region to be vertically longer than when the moving body is not on a road surface with an uphill gradient of a predetermined gradient or more.

[0015] (10): In the aspect (6) above, when the moving body is on a road surface with a downward gradient of a predetermined gradient or more, the definition unit changes the aspect ratio of the partial region to be horizontally longer than when the moving body is not on a road surface with a downward gradient of a predetermined gradient or more.

[0016] (11): In the above aspect (1), the definition unit defines the partial region as a horizontally long rectangle.

[0017] (12) In the above aspect (1), the extracting unit extracts the portion of interest by regarding the counted value less than the lower limit value as zero.

[0018] (13): Another aspect of the object detection method of the present invention is an object detection method executed by a computer, comprising: acquiring an image of a surface on which a moving object can pass, captured at an angle relative to the surface; generating a low-resolution image by reducing the image quality of the captured image; defining a plurality of partial area sets, each of which includes a partial area; deriving an aggregate value by aggregating differences in features between the partial area included in each of the plurality of partial area sets and surrounding partial areas; and extracting a point of interest based on the aggregate value, wherein each of the plurality of partial area sets is defined to include a plurality of partial areas in a target area for each partial area set, and the target area is a vertically limited portion of the low-resolution image cut out so that at least a portion of the target area does not overlap with other partial area sets in the vertical direction.

[0019] (14): Another aspect of the present invention is a program executed by a computer, which causes the computer to acquire an image of a surface on which a moving body can pass, captured at an angle relative to the surface, generate a low-resolution image by reducing the image quality of the captured image, define a plurality of partial area sets, each of which includes a partial area, derive an aggregated value by aggregating the differences in features between the partial areas included in each of the plurality of partial area sets and surrounding partial areas, and extract a point of interest based on the aggregated value, wherein each of the plurality of partial area sets is defined to include a plurality of partial areas in a target area for each partial area set, and the target area is a vertically limited portion of the low-resolution image cut out so that at least a portion of the target area does not overlap with other partial area sets in the vertical direction. [Effects of the Invention]

[0020] According to the above aspects (1) to (14), object detection can be suitably performed while reducing the processing load. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a diagram illustrating an example of the configuration of an object detection device 100 and peripheral devices. [Figure 2] 2 is a diagram schematically illustrating the function of each part of the object detection device 100. FIG. [Figure 3] 10 is a diagram for explaining the processing of a mask region determination unit 130, a grid definition unit 140, and an extraction unit 150. FIG. [Figure 4] 10 is a diagram for explaining the processing of a feature amount difference calculation unit 152, a counting unit 154, and an addition unit 156. FIG. [Figure 5] FIG. 10 is a diagram illustrating an example of a definition of a surrounding grid. [Figure 6] FIG. 10 is a diagram illustrating an example of rules for selecting a comparison target grid and a comparison source grid. [Figure 7] FIG. 10 is a diagram illustrating another example of rules for selecting a comparison target grid and a comparison source grid. [Figure 8] 10 is a diagram for explaining the processing of an adder 156 and a combiner 158. FIG. [Figure 9] FIG. 9 is a diagram for explaining the processing of the attention portion extraction section 160. In FIG. DETAILED DESCRIPTION OF THE INVENTION

[0022] Hereinafter, with reference to the drawings, embodiments of an object detection device, an object detection method, and a program of the present invention will be described. The object detection device is mounted on, for example, a mobile body. The mobile body may be, for example, a four-wheeled vehicle, a two-wheeled vehicle, micromobility, a robot that moves on its own, an aerial vehicle such as a drone, or a portable device such as a smartphone that is mounted on a mobile body that moves on its own or is carried by a person. In the following description, the mobile body is assumed to be a four-wheeled vehicle, and the mobile body will be referred to as a "vehicle." The object detection device is not limited to one mounted on a mobile body, but may also be one that performs the processing described below based on images captured by a fixed-point observation camera or a smartphone camera.

[0023] [composition] 1 is a diagram showing an example of the configuration of an object detection device 100 and peripheral devices. The object detection device 100 communicates with a camera 10, a driving control device 200, a notification device 210, and the like.

[0024] The camera 10 is attached to the rear surface of the vehicle's windshield or the like, captures an image of at least the road in the direction of travel of the vehicle, and outputs the captured image to the object detection device 100. A sensor fusion device or the like may be interposed between the camera 10 and the object detection device 100, but this will not be described here. The camera 10 is an example of a device that captures an image of a surface on which a mobile object can pass, at an angle relative to the surface. The mobile object has been described above. The "surface on which a mobile object can pass" may include outdoor surfaces such as roads (street surfaces) and public open spaces, as well as corridors and room floors when the mobile object moves indoors. "At an angle relative to the surface" means that the image is not captured by looking directly down from the sky using an aircraft. In other words, the image is captured at an angle of at least a predetermined angle. Specifically, this means capturing an image from a height of less than 5 m, for example, so that the ground plane is included in the captured image. In other words, "at an angle with respect to the surface" means capturing an image from a height of less than 5 m, for example, at a depression angle of less than about 20 degrees. Camera 10 may be mounted on a moving object that moves in contact with the "surface," or on a drone or the like that flies at low altitude.

[0025] The driving control device 200 is, for example, an automatic driving control device that causes the vehicle to drive autonomously, or a driving support device that performs inter-vehicle distance control, automatic braking control, and lane change control, etc. The notification device 210 is a speaker, vibrator, light-emitting device, display device, etc. that output information to the vehicle occupants.

[0026] The object detection device 100 includes, for example, an acquisition unit 110, a low-resolution image generation unit 120, a grid definition unit 140, an extraction unit 150, and a high-resolution processing unit 170. The extraction unit 150 includes a feature difference calculation unit 152, an aggregation unit 154, an addition unit 156, a synthesis unit 158, and a focus area extraction unit 160. These components are implemented by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be implemented by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or may be implemented by a combination of software and hardware. The program may be stored in advance in a storage device (a storage device with a non-transitory storage medium) such as an HDD (Hard Disk Drive) or flash memory, or may be stored in a removable storage medium (a non-transitory storage medium) such as a DVD or CD-ROM, and installed by inserting the storage medium into a drive device.

[0027] Fig. 2 is a diagram schematically illustrating the function of each unit of object detection device 100. Each unit of object detection device 100 will be described below with reference to Fig. 2. Acquisition unit 110 acquires a captured image from camera 10. Acquisition unit 110 stores (data of) the acquired captured image in a working memory such as RAM (Random Access Memory).

[0028] The low-resolution image generating unit 120 performs thinning processing on the captured image to generate a low-resolution image with lower image quality than the captured image. The low-resolution image is, for example, an image with fewer pixels than the captured image.

[0029] The mask region determination unit 130 determines a mask region that is not to be processed by the grid definition unit 140 and the following components. This will be described in detail later.

[0030] The grid definition unit 140 defines multiple partial area sets in the low-resolution image. "Defining" refers to determining boundaries for the low-resolution image. Each of the multiple partial area sets is defined by cutting out multiple partial areas (hereinafter referred to as grids) from the low-resolution image. The grids are set, for example, in a rectangular shape with no gaps. The grids are, for example, square, but may also be horizontally long rectangles. As will be described later, the grid definition unit 140 may change the size or aspect ratio of the grids based on the environment in which the moving object is located. The grid definition unit 140 defines the multiple partial area sets so that the number of pixels within a grid increases (i.e., the grid becomes larger) as the grids are defined closer to the front of the low-resolution image (i.e., lower in the image). Hereinafter, the multiple partial area sets may be referred to as the first partial area set PA1, the second partial area set PA2, ..., the kth partial area set PAk. The detailed function of the grid definition unit 140 will be described later.

[0031] The extraction unit 150 derives an aggregate value by aggregating differences in feature amounts between the grids included in each of the plurality of partial area sets and the surrounding grids, and adds up the aggregate values ​​between the plurality of partial area sets to extract a point of interest (a point of discontinuity with the surroundings in the figure). The detailed functions of each unit of the extraction unit 150 will be described later.

[0032] The high-resolution processing unit 170 cuts out a portion of the captured image corresponding to the location of interest (synchronous cutout in the figure), and performs high-resolution processing on this to determine whether or not the object on the road is an object with which the vehicle should avoid contact. For example, the high-resolution processing unit 170 uses a trained model that recognizes road markings (an example of an object that is not an object with which the vehicle should avoid contact) and fallen objects (an example of an object with which the vehicle should avoid contact) from the image to determine whether the image captured at the location of interest is a road marking, a fallen object, or is unknown (an untrained object). At this time, the high-resolution processing unit 170 may further narrow down the location of interest in the captured image to a portion that is recognized as corresponding to a road marking or a fallen object and perform processing on it.

[0033] FIG. 3 is a diagram illustrating the processing of the mask region determination unit 130, the grid definition unit 140, and the extraction unit 150. For example, the mask region determination unit 130 extracts edge points in the left-right direction in a low-resolution image and detects the positions of road dividing lines, road shoulders, etc. (white lines, lane boundaries) in the image by connecting the edge points arranged in a straight line. Then, the mask region determination unit 130 detects the area sandwiched between the left and right road dividing lines, etc. and including the center point in the left-right direction on the foreground side of the image as the vehicle's lane. Next, the mask region determination unit 130 determines the area other than the vehicle's lane (the area above the vanishing point where the road dividing lines, etc. intersect on the far side and the area closer to the left and right ends of the road dividing lines) as the mask region. The grid definition unit 140 and the extraction unit 150 perform processing excluding the mask region.

[0034] The grid definition unit 140 defines each of the multiple partial area sets so that the target area for each partial area set includes multiple partial areas. The target area is a vertically limited portion of the low-resolution image from which the mask area has been removed, extracted so that at least a portion of the target area does not overlap vertically with other partial area sets. In the following description, it is assumed that the partial area sets are extracted so that they do not overlap vertically with other partial area sets. As described above, the grid definition unit 140 defines partial area sets in order, starting with the first partial area set PA1, which has the largest number of grid pixels, followed by the second partial area set PA2, which has the next largest number of grid pixels, and so on, down to the k-th partial area set PAk, which has the fewest number of grid pixels.

[0035] The processing of the feature amount difference calculation unit 152, the aggregation unit 154, and the addition unit 156 will be described below. The processing of these functional units, which will be described with reference to FIGS. 4 to 7, is performed by first selecting one partial region set and then selecting one grid of interest from the selected partial region set. After all grids in the selected partial region set have been selected as grids of interest and processing is complete, the next partial region set is selected and processing is performed in the same manner. After processing is complete for all partial region sets, the synthesis unit 158 ​​synthesizes (combines) the processing results of each partial region set to generate extraction target data PT, which is a single image, and passes this data to the focus portion extraction unit 160. Note that if a partial region set is defined so that it partially overlaps with other partial region sets, the synthesis unit 158 ​​may add or average the processing results for the overlapping portions.

[0036] FIG. 4 is a diagram illustrating the processing of the feature difference calculation unit 152, the aggregation unit 154, and the addition unit 156. The feature difference calculation unit 152 calculates the difference in the feature for each pixel between the comparison target grid and the comparison source grid. The feature is, for example, a luminance value for each R, G, and B component, and a set of R, G, and B corresponds to one pixel. The comparison target grid and the comparison source grid are selected from the target grid and the surrounding grids. FIG. 5 is a diagram illustrating an example of the definition of the surrounding grids. As shown in the figure, grids 2 to 9 adjacent to the target grid in the up, down, left, right, and diagonal directions are defined as the surrounding grids. The method of selecting the surrounding grids (surrounding partial regions) is not limited to this; the grids above, below, left, and right may be selected as the surrounding grids, or the surrounding grids may be selected according to other rules.

[0037] The comparison target grid and comparison source grid are selected in order from the combinations shown in FIG. 6, for example. FIG. 6 is a diagram showing an example of rules for selecting the comparison target grid and comparison source grid. The comparison target grid is the grid of interest, and the comparison source grids are selected in order from grids 2 to 9. The relationship between the comparison target grid and the comparison source grid may be reversed. Then, the aggregation unit 154 calculates the sum of the differences in the feature amounts for each pixel and divides it by the number of pixels n in the grid to calculate a first aggregate value V1. If the grid of interest corresponds to a mask area, the first aggregate value V1 is replaced with zero and output. In other words, the feature difference calculation unit 152, the aggregation unit 154, and the addition unit 156 process grid 1 as the comparison target grid and grid 3 as the comparison source grid, process grid 1 as the comparison target grid and grid 8 as the comparison source grid, process grid 1 as the comparison target grid and grid 5 as the comparison source grid, process grid 1 as the comparison target grid and grid 6 as the comparison source grid, process grid 1 as the comparison target grid and grid 2 as the comparison source grid, process grid 1 as the comparison target grid and grid 4 as the comparison source grid, process grid 1 as the comparison target grid and grid 7 as the comparison source grid, and process grid 1 as the comparison target grid and grid 9 as the comparison source grid, in parallel or sequentially.

[0038] The comparison target grid and the comparison source grid may be selected in order from the combinations shown in Fig. 7. Fig. 7 is a diagram showing another example of the rules for selecting the comparison target grid and the comparison source grid. The combination of the comparison target grid and the comparison source grid is not limited to the combination of the grid of interest and a surrounding grid, but may also include combinations of surrounding grids (particularly, combinations of the upper grid and the lower grid, the left grid and the right grid, the upper left grid and the lower right grid, and the upper right grid and the lower left grid).

[0039] More specifically, a method for calculating the difference in feature amounts will be described. As a method for calculating the difference in feature amounts, the following patterns 1 to 4 can be considered, for example. In the following description, the identification number of a pixel in each of the comparison target grid and comparison source grid is represented by i (i=1 to k; k is the number of pixels in each of the comparison target grid and comparison source grid).

[0040] (Pattern 1) The feature amount difference calculation unit 152 calculates, for example, the difference ΔRi in the brightness of the R component, the difference ΔGi in the brightness of the G component, and the difference ΔBi in the brightness of the B component between pixels at the same position in both the comparison target grid and the comparison source grid (where i=1 to k as described above). 2 +ΔGi 2 +ΔBi 2 is calculated for each pixel, and the maximum value or average value of each pixel feature amount Ppi is calculated as the difference in feature amount between the comparison target grid and the comparison source grid.

[0041] (Pattern 2) The feature difference calculation unit 152 calculates, for example, the statistical value (mean value, median, mode, etc.) Raa of the R component brightness, the statistical value (same) Gaa of the G component brightness, and the statistical value (same) Baa of the B component brightness of each pixel in the comparison target grid, and calculates the statistical value (same) Rab of the R component brightness, the statistical value (same) Gab of the G component brightness, and the statistical value (same) Bab of the B component brightness of each pixel in the comparison source grid, and finds the differences ΔRa (=Raa-Rab), ΔGa (=Gaa-Gab), and ΔBa (=Baa-Bab) between them. Then, ΔRa, which is the sum of the squares of the brightness differences, is calculated. 2 +ΔGa 2 +ΔBa 2 , or the maximum value of the square of the brightness difference Max(ΔRa 2 ,ΔGa 2 ,ΔBa 2 ) is calculated as the difference in the feature amounts between the comparison target grid and the comparison source grid.

[0042] (Pattern 3) The feature amount difference calculation unit 152 calculates, for example, for each pixel i in the comparison target grid, a first index value W1ai (=(RB) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the B component by the sum of the luminance of the R, G, and B components, and a second index value W2ai (=(RG) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the G component by the sum of the luminance of the R, G, and B components. Furthermore, the feature amount difference calculation unit 152 calculates, for example, for each pixel i in the comparison source grid, a first index value W1bi (=(RB) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the B component by the sum of the luminance of the R, G, and B components, and a second index value W2bi (=(RG) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the G component by the sum of the luminance of the R, G, and B components. Next, the feature amount difference calculation unit 152 calculates each pixel feature amount Ppi=(W1ai-W1bi) 2 +(W2ai-W2bi) 2Then, the feature difference calculation unit 152 calculates the maximum or average value of each pixel feature Ppi as the difference in feature between the comparison target grid and the comparison source grid. Note that by combining the first index value and the second index value, the balance of the RGB components in each pixel can be expressed. Using the same concept as above, for example, the luminance of each R, G, and B component may be defined as the magnitude of a vector shifted by 120 degrees, and the vector sum may be used in the same way as the combination of the first index value and the second index value.

[0043] (Pattern 4) The feature difference calculation unit 152, for example, calculates the statistical value (same) Raa of the R component brightness, the statistical value (same) Gaa of the G component brightness, and the statistical value (same) Baa of the B component brightness for each pixel in the comparison target grid, and calculates the statistical value (same) Rab of the R component brightness, the statistical value (same) Gab of the G component brightness, and the statistical value (same) Bab of the B component brightness for each pixel in the comparison source grid. Next, the feature difference calculation unit 152 calculates a third index value W3a (=(Raa-Baa) / (Raa+Gaa+Baa)) for the comparison grid by dividing the difference between the statistical value Raa of the R component brightness and the statistical value Baa of the B component brightness by the sum of the statistical values ​​of the R, G, and B components brightness, and a fourth index value W4a (=(Raa-Gaa) / (Raa+Gaa+Baa)) by dividing the difference between the statistical value Raa of the R component brightness and the statistical value Gaa of the G component brightness by the sum of the statistical values ​​of the R, G, and B components brightness. Similarly, for the comparison source grid, the feature amount difference calculation unit 152 calculates a third index value W3b (=(Rab-Bab) / (Rab+Gab+Bab)) by dividing the difference between the R component luminance statistical value Rab and the B component luminance statistical value Bab by the sum of the R, G, and B component luminance statistical values, and a fourth index value W4b (=(Rab-Gab) / (Rab+Gab+Bab)) by dividing the difference between the R component luminance statistical value Rab and the G component luminance statistical value Gab by the sum of the R, G, and B component luminance statistical values. Then, the feature amount difference calculation unit 152 calculates the difference ΔW3 between the third index value W3a of the comparison target grid and the third index value W3b of the comparison source grid, and the difference ΔW4 between the fourth index value W4a of the comparison target grid and the fourth index value W4b of the comparison source grid, and calculates the sum of their squares, ΔW3 2 +ΔW42 , or the maximum value of the squares Max(ΔW3 2 ,ΔW4 2 ) is calculated as the difference in the feature amounts between the comparison target grid and the comparison source grid.

[0044] If the image to be processed is a black-and-white image, the feature difference calculation unit 152 may simply calculate the difference in luminance value as the difference in feature between the comparison target grid and the comparison source grid. Also, if the image to be processed is an RGB image, the feature difference calculation unit 152 may convert the RGB image into a black-and-white image and calculate the difference in trajectory value as the difference in feature between the comparison target grid and the comparison source grid.

[0045] 4, the adder 156 calculates a second aggregate value V2 by adding the first aggregate values ​​V1 calculated for the grid of interest. The second aggregate value V2 is an example of the "aggregate value" in the claims. When the process of calculating the second aggregate value V2 while changing the grid of interest is completed, data is generated in which the second aggregate value V2 is set for all grids for each partial region set.

[0046] Once data in which the second aggregate value V2 is set for all grids is generated, the synthesis unit 158 ​​combines the data to generate a single image, the extraction target data PT. FIG. 8 is a diagram illustrating the processing performed by the addition unit 156 and synthesis unit 158. In the diagram, the smallest rectangle represents one pixel of the low-resolution image. For ease of explanation, the first partial region set PA1 and the second partial region set PA2 are shown as representative partial region sets, and their horizontal sizes are assumed to be significantly smaller than their actual sizes. It is also assumed that the second aggregate value V2 has been normalized at some stage to a value between zero and one. In the illustrated example, the first partial region set PA1 is a set of first grids consisting of 16 pixels, and the second partial region set PA2 is a set of second grids consisting of 9 pixels. Although examples of grid sizes of 16 pixels, 9 pixels, and 4 pixels have been shown, when the above-mentioned pattern 2 or pattern 4 is adopted as the method for calculating the difference in feature amounts, the computation load for statistical values ​​can be reduced by using a grid size in which the number of pixels on one side is a power of 2, such as 4 pixels, 16 pixels, or 64 pixels.

[0047] 9 is a diagram for explaining the processing of the portion-of-interest extraction unit 160. The portion-of-interest extraction unit 160 sets a search area WA according to the grid size for the processing target data PT, for example, for each area with the same grid size (i.e., for each area divided depending on which partial area set the data originates from), and extracts the search area WA in which the sum of the second aggregated values ​​V2 within the search area WA is equal to or greater than a reference value, as the portion of interest. In this case, the search area WA is set to a fixed size, for example, two grids horizontally and one grid vertically.

[0048] Alternatively, the attention location extraction unit 160 may set the search area WA to a variable size, in which case the attention location extraction unit 160 may extract, as the attention location, the search area WA in which the difference in the second aggregate value V2 between the search area WA and the grids surrounding the search area WA is locally largest. The search area WA with the locally largest difference may appear in multiple places.

[0049] In either case, the portion of interest extraction section 160 may replace the second total value V2 that is less than the lower limit value with zero (regarding it as zero) before carrying out the above process.

[0050] As described above, the high-resolution processing unit 170 performs high-resolution processing on the area obtained by applying only the position of the point of interest to the captured image, and determines whether the object on the road is an object with which the vehicle should avoid contact.

[0051] The determination result of the high resolution processing unit 170 is output to the cruise control device 200 and / or the alarm device 210. The cruise control device 200 performs automatic braking control, automatic steering control, etc. to avoid contact between the vehicle and the object (actually an area on the image) determined to be a "falling object." The alarm device 210 outputs an alarm in various ways when the TTC (Time To Collision) between the object (same as above) determined to be a "falling object" and the vehicle falls below a threshold.

[0052] According to the embodiment described above, by providing an acquisition unit 110 that acquires an image of at least the road in the direction of travel of the vehicle, a low-resolution image generation unit 120 that generates a low-resolution image by reducing the image quality of the captured image, a grid definition unit 140 that defines one or more partial area sets, and an extraction unit 150 that derives an aggregate value by aggregating the differences in feature values ​​between the partial areas included in each of the one or more partial area sets and surrounding partial areas, and extracts points of interest based on the aggregate value, it is possible to maintain high detection accuracy while reducing the processing load.

[0053] If the processing performed by the feature amount difference calculation unit 152 and the aggregation unit 154 were performed on the captured image as is, the processing load would increase as the number of pixels increases, and there is a concern that the operation of the travel control device 200 and the alarm device 210 would not be able to keep up with the approach of a falling object. In this regard, the object detection device 100 of the embodiment generates a low-resolution image and then processes it, thereby making it possible to detect objects while reducing the processing load.

[0054] Furthermore, according to the embodiment, the grid definition unit 140 defines multiple partial area sets so that the number of pixels in the grid varies among the multiple partial area sets, and the extraction unit 150 extracts the target area by adding the aggregated values ​​for each pixel among the multiple partial area sets. This improves the robustness of detection performance against variations in the size of falling objects. Simply performing processing using low-resolution images can reduce image quality, potentially making the presence of falling objects unrecognizable. However, according to the embodiment, the above-described ingenuity makes it possible to expect that falling objects will appear as feature quantities in grids of any size. As described above, the object detection device 100 of the embodiment can maintain high detection accuracy while reducing processing load.

[0055] [Other examples of grid definitions] The grid definition unit 140 may change the aspect ratio of the grid based on the environment in which the vehicle is located. In this case, the aspect ratio of the search area WA is necessarily changed as well. In this case, the object detection device acquires various information required for the following processing from on-board sensors such as a vehicle speed sensor, steering angle sensor, yaw rate sensor, and gradient sensor.

[0056] For example, when the vehicle speed V is greater than the reference speed V1, the grid definition unit 140 changes the aspect ratio of the grid to be vertical compared to when the vehicle speed V is equal to or less than the reference speed V1. This is because as the speed V increases, the probability that the image captured by the camera 10 will be blurred vertically due to vibrations of the vehicle increases. By changing the aspect ratio of the grid to be vertical, even if a group of pixels whose feature values ​​differ significantly from their surroundings are stretched vertically due to image blur, the probability that the stretched portion will fit within the grid can be increased. "Changing the aspect ratio to be vertical" may mean enlarging the vertical size while maintaining the horizontal size, enlarging the vertical size while reducing the horizontal size, or maintaining the vertical size while reducing the horizontal size. "Changing the aspect ratio to be horizontal" means the opposite.

[0057] Furthermore, when the turning angle θ of the vehicle is greater than the reference angle θ1, the grid definition unit 140 may change the aspect ratio of the grid to be horizontally longer than when the turning angle θ of the vehicle is equal to or less than the reference angle θ1. Here, the turning angle θ is assumed to be information of an absolute value with the neutral position of the steering device set to zero. The turning angle θ may be an angular velocity or a steering angle. This is because as the turning angle θ of the vehicle increases, the probability that the image from the camera 10 will be blurred horizontally due to the turning behavior increases.

[0058] Furthermore, the grid definition unit 140 may change the aspect ratio of the grid to be more vertical when the vehicle is on an uphill road surface with a gradient of φ1 or greater than a predetermined gradient, compared to when the vehicle is not on an uphill road surface with a gradient of φ1 or greater. Alternatively, the grid definition unit 140 may change the aspect ratio of the grid to be more horizontal when the vehicle is on a downhill road surface with a gradient of φ2 or greater than a predetermined gradient, compared to when the vehicle is not on a downhill road surface with a gradient of φ2 or greater. The gradients φ1 and φ2 are both absolute values ​​(positive values ​​for both uphill and downhill directions) and may be the same or different values. This is because on an uphill slope, the portion of the image captured by camera 10 that captures the road surface extends relatively toward the top of the image (i.e., the portion of the image captured by camera 10 is stretched vertically compared to a flat road), whereas on a downhill slope, the portion of the image captured by camera 10 that captures the road surface extends relatively toward the bottom of the image (i.e., the portion of the image captured by camera 10 is compressed vertically compared to a flat road).

[0059] If the above conditions occur simultaneously, for example, if the vehicle speed V is greater than the reference speed V1 and the vehicle is on a road surface with a downward gradient of at least the predetermined gradient φ2, the grid definition unit 140 may determine the shape of the grid by offsetting the change in aspect ratio caused by the vehicle speed V being greater than the reference speed V1 and the change in aspect ratio caused by the vehicle being on a road surface with a downward gradient of at least the predetermined gradient φ2. The same applies when other conditions occur simultaneously.

[0060] Furthermore, since the optimal grid shape and size vary depending on the type of falling object, the grid definition unit 140 may set multiple partial area sets with different grid definitions depending on the expected size of the target falling object, and perform processing on them in parallel.

[0061] The above-described embodiment can be expressed as follows. a storage medium for storing computer-readable instructions; a processor connected to the storage medium; The processor executes the computer-readable instructions to: acquiring a captured image of a surface on which a moving object can pass, the captured image being inclined relative to the surface; generating a low-resolution image by reducing the image quality of the captured image; defining a plurality of subregion sets, each set including a subregion; deriving an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the plurality of partial region sets and surrounding partial regions, and extracting a portion of interest based on the aggregate value; each of the plurality of partial region sets is defined to include a plurality of partial regions in a target region for each partial region set; the target region is a vertically limited portion of the low-resolution image cut out so as not to overlap at least a portion of the target region with other partial region sets in the vertical direction; Object detection device.

[0062] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0063] 10 Camera 100 Object detection device 110 Acquisition Department 120 Low-resolution image generation unit 130 Mask area determination unit 140 Grid definition section 150 Extraction part 152 Feature Difference Calculation Unit 154 Counting Department 156 Addition section 158 Synthesis Section 160 Focus point extraction unit 170 High-resolution processing section

Claims

1. an acquisition unit that acquires an image of a surface on which a moving object can pass, the image being captured at an angle with respect to the surface; a low-resolution image generating unit that generates a low-resolution image by reducing the image quality of the captured image; a definition unit that defines a plurality of subregion sets, each of which includes a subregion; each of the plurality of partial region sets is defined to include a plurality of partial regions in a target region for each partial region set; a definition unit, wherein the target region is a vertically limited portion of the low-resolution image cut out so that at least a portion of the target region does not overlap with other partial region sets in the vertical direction; an extraction unit that derives an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the plurality of partial region sets and surrounding partial regions, and extracts a portion of interest based on the aggregate value; a high-resolution processing unit that performs high-resolution processing on the target area in the captured image and detects objects on the road; An object detection device comprising:

2. the definition unit defines the plurality of partial region sets such that the number of pixels in a partial region increases as the partial region defined in the partial region sets becomes closer to the front of the low-resolution image. The object detection device according to claim 1 .

3. the extraction unit derives the aggregated value by aggregating differences in feature amounts between the partial regions included in each of the plurality of partial region sets and other partial regions adjacent thereto above, below, left, right, and diagonally; The object detection device according to claim 1 .

4. the extraction unit further adds, for each of the partial regions included in the plurality of partial region sets, differences in feature amounts between the partial regions adjacent to each other vertically, differences in feature amounts between the partial regions adjacent to each other horizontally, and differences in feature amounts between the partial regions adjacent to each other diagonally to each other to the aggregated value.

4. The object detection device according to claim 3.

5. a high-resolution processing unit that performs high-resolution processing on the target location in the captured image and determines whether or not an object on the road is an object with which the moving body should avoid contact; The object detection device according to claim 1 .

6. the object detection device is mounted on a moving body, the definition unit changes the aspect ratio of the partial region based on an environment in which the moving object is located. The object detection device according to claim 1 .

7. the definition unit changes the aspect ratio of the partial region to be vertically longer when the speed of the moving object is higher than a reference speed compared to when the speed of the moving object is equal to or lower than the reference speed; 7. The object detection device according to claim 6.

8. the definition unit changes the aspect ratio of the partial region to be horizontally longer when the turning angle of the moving body is larger than a reference angle compared to when the turning angle of the moving body is equal to or smaller than the reference angle; 7. The object detection device according to claim 6.

9. the definition unit, when the moving object is on a road surface with an upslope of a predetermined gradient or more, changes the aspect ratio of the partial region to be vertically longer than when the moving object is not on a road surface with an upslope of a predetermined gradient or more.

7. The object detection device according to claim 6.

10. the definition unit, when the moving object is on a road surface with a downward gradient of a predetermined gradient or more, changes the aspect ratio of the partial region to be horizontally longer than when the moving object is not on a road surface with a downward gradient of a predetermined gradient or more.

7. The object detection device according to claim 6.

11. the definition unit defines the partial region as a horizontally long rectangular shape; The object detection device according to claim 1 .

12. the extraction unit extracts the portion of interest by regarding the aggregated value less than a lower limit value as zero. The object detection device according to claim 1 .

13. 1. A computer-implemented method for object detection, comprising: acquiring a captured image of a surface on which a moving object can pass, the captured image being inclined relative to the surface; generating a low-resolution image by reducing the image quality of the captured image; defining a plurality of subregion sets, each set including a subregion; deriving an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the plurality of partial region sets and surrounding partial regions, and extracting a portion of interest based on the aggregate value; performing high-resolution processing on the location of interest in the captured image to detect an object on the road; each of the plurality of partial region sets is defined to include a plurality of partial regions in a target region for each partial region set; the target region is a vertically limited portion of the low-resolution image cut out so as not to overlap at least a portion of the target region with other partial region sets in the vertical direction; Object detection methods.

14. A program executed by a computer, the program comprising: acquiring a captured image of a surface on which a moving object can pass, the captured image being inclined relative to the surface; generating a low-resolution image by reducing the image quality of the captured image; defining a plurality of subregion sets, each set including a subregion; deriving an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the plurality of partial region sets and surrounding partial regions, and extracting a portion of interest based on the aggregate value; each of the plurality of partial region sets is defined to include a plurality of partial regions in a target region for each partial region set; the target region is a vertically limited portion of the low-resolution image cut out so as not to overlap at least a portion of the target region with other partial region sets in the vertical direction; program.

Citation Information

Patent Citations

  • External monitoring device having fail / safe function

    JP2001028746A

  • Device and method for reading number plate

    JP2001273461A

  • Image processing system, control method and control program of the same

    JP2014153866A

  • Image processing device, image processing method, and image processing system

    JP2018107759A

  • Failure detection system

    JP2019124986A