OBJECT DETECTION DEVICE, OBJECT DETECTION METHOD, AND PROGRAM

By generating low-resolution images and calculating feature differences between sub-regions, and extracting points of interest, the problem of excessive load processing and insufficient accuracy of object detection systems in the prior art is solved, and efficient and accurate object detection is achieved.

JP7676193B2Active Publication Date: 2025-05-14HONDA MOTOR CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021060571
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-05-14
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

In the prior art, object detection systems are dealing with problems such as excessive load or insufficient accuracy.

Method used

By generating low-resolution images and defining multiple sub-region sets, each of which consists of sub-region areas in the low-resolution image, the feature differences between sub-regions are calculated and summed into aggregate values, and points of interest are extracted to reduce processing load.

Benefits of technology

It realizes the accuracy of object detection while reducing processing load, and improves the robustness of detection performance of objects of different sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676193000001
    Figure 0007676193000001
  • Figure 0007676193000002
    Figure 0007676193000002
  • Figure 0007676193000003
    Figure 0007676193000003
Patent Text Reader

Abstract

To provide an object detection device, an object detection method, and a program capable of maintaining high detection accuracy while reducing processing load.SOLUTION: The object detection device includes: an acquisition unit that acquires a captured image of a road; a low-resolution image generation unit that generates a low-resolution image obtained by degrading the image quality of the captured image; a definition unit that defines one or more partial area sets in which each of the one or more partial area sets is defined by cutting out a plurality of partial areas from the low-resolution image; and an extracting unit that, for the partial areas included in each of the one or more partial areas, derives a total value obtained by totaling differences in feature amounts between the partial area and peripheral partial areas, and extracts a point of interest based on the total value.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an object detection device, an object detection method, and a program. [Background technology]

[0002] Conventionally, an invention has been disclosed for a driving obstacle detection system that divides the area of ​​objects in a surveillance area, such as on a road, obtained by photographing into blocks, extracts local features for each block, and determines the presence or absence of an obstacle based on the extracted local features (Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2019-124986 A Summary of the Invention [Problem to be solved by the invention]

[0004] Conventional techniques can result in excessive processing loads or insufficient accuracy.

[0005] The present invention has been made in consideration of the above circumstances, and one of its objectives is to provide an object detection device, an object detection method, and a program that are capable of performing object detection while reducing the processing load. [Means for solving the problem]

[0006] The object detection device, the object detection method, and the program according to the present invention employ the following configuration. (1): An object detection device according to one embodiment of the present invention includes an acquisition unit that acquires an image of a road, a low-resolution image generation unit that generates a low-resolution image by reducing the image quality of the captured image, a definition unit that defines one or more partial area sets, each of which is defined by cutting out a plurality of partial areas from the low-resolution image, and an extraction unit that derives an aggregate value by aggregating differences in features between the partial areas included in each of the one or more partial area sets and surrounding partial areas, and extracts a point of interest based on the aggregate value.

[0007] (2): In the above aspect (1), the definition unit defines a plurality of partial area sets so that the number of pixels within the partial area differs between the plurality of partial area sets, and the extraction unit extracts the point of interest by adding up the aggregated value for each pixel between the plurality of partial area sets.

[0008] (3): In the above aspect (1) or (2), the extraction unit derives the aggregated value by aggregating the differences in features between the partial regions included in each of the one or more partial region sets and other partial regions adjacent thereto above, below, to the left, to the right, and diagonally.

[0009] (4): In the aspect of (3) above, the extraction unit further adds to the aggregated value, for each of the partial regions included in the one or more partial region sets, the difference in features between the partial regions adjacent vertically, the difference in features between the partial regions adjacent horizontally, and the difference in features between the partial regions adjacent diagonally.

[0010] (5): In any of the above aspects (1) to (4), the present invention further includes a high-resolution processing unit that performs high-resolution processing on the point of interest in the captured image to determine whether an object on the road is an object with which the moving body should avoid contact.

[0011] (6): An object detection method according to another aspect of the present invention is an object detection method executed using a computer, comprising: acquiring an image of a road; generating a low-resolution image by reducing the image quality of the captured image; defining one or more partial area sets; deriving an aggregated value by aggregating differences in features between partial areas included in each of the one or more partial area sets and surrounding partial areas; and extracting points of interest based on the aggregated value, wherein each of the one or more partial area sets is defined by cutting out a plurality of partial areas from the low-resolution image.

[0012] (7): A program according to another aspect of the present invention is a program executed by a computer, which causes the computer to acquire an image of a road surface, generate a low-resolution image by reducing the image quality of the captured image, define one or more partial area sets, derive an aggregated value by aggregating the differences in features between the partial areas included in each of the one or more partial area sets and surrounding partial areas, and extract a point of interest based on the aggregated value, wherein each of the one or more partial area sets is defined by cutting out a plurality of partial areas from the low-resolution image. Effect of the Invention

[0013] According to the above aspects (1) to (7), object detection can be performed while reducing the processing load. According to the above aspect (2), it is possible to further improve the robustness of the detection performance against variations in the size of the object captured in the focused area. [Brief description of the drawings]

[0014] [Figure 1] 1 is a diagram illustrating an example of a configuration of an object detection device 100 and peripheral devices. [Diagram 2] 2 is a diagram illustrating a schematic diagram of the function of each part of the object detection device 100. FIG. [Diagram 3]10 is a diagram for explaining the processing of a mask region determination unit 130, a grid definition unit 140, and an extraction unit 150. FIG. [Figure 4] 13 is a diagram for explaining the processing of a feature amount difference calculation unit 152, a counting unit 154, and a first addition unit 156. FIG. [Diagram 5] FIG. 13 is a diagram illustrating an example of a definition of a surrounding grid. [Figure 6] FIG. 13 is a diagram illustrating an example of a rule for selecting a comparison target grid and a comparison source grid. [Figure 7] FIG. 11 is a diagram illustrating another example of the rule for selecting a comparison target grid and a comparison source grid. [Figure 8] 13 is a diagram for explaining the processing of a first adder 156 and a second adder 158. FIG. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, with reference to the drawings, an embodiment of an object detection device, an object detection method, and a program of the present invention will be described. The object detection device is mounted on a moving body, for example. The moving body is, for example, a four-wheeled vehicle, a two-wheeled vehicle, a micromobility, a robot that moves by itself, or a portable device such as a smartphone that is placed on a moving body that moves by itself or is carried by a person. In the following description, the moving body is a four-wheeled vehicle, and the moving body is referred to as a "vehicle". The object detection device is not limited to being mounted on a moving body, and may be a device that performs the processing described below based on an image captured by a fixed-point observation camera or a smartphone camera.

[0016] 1 is a diagram showing an example of the configuration of an object detection device 100 and peripheral devices. The object detection device 100 communicates with a camera 10, a driving control device 200, a notification device 210, and the like.

[0017] The camera 10 is attached to the rear surface of the windshield of the vehicle, captures an image of at least the road in the traveling direction of the vehicle, and outputs the captured image to the object detection device 100. Note that a sensor fusion device or the like may be interposed between the camera 10 and the object detection device 100, but a description thereof will be omitted.

[0018] The driving control device 200 is, for example, an automatic driving control device that drives the vehicle autonomously, a driving support device that performs vehicle distance control, automatic braking control, lane change control, etc. The notification device 210 is a speaker, vibrator, light emitting device, display device, etc. for outputting information to the vehicle occupants.

[0019] The object detection device 100 includes, for example, an acquisition unit 110, a low-resolution image generation unit 120, a grid definition unit 140, an extraction unit 150, and a high-resolution processing unit 170. The extraction unit 150 includes a feature amount difference calculation unit 152, an aggregation unit 154, a first addition unit 156, a second addition unit 158, and a focus area extraction unit 160. These components are realized by, for example, a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware (including circuitry) such as an LSI (Large Scale Integration), an ASIC (Application Specific Integrated Circuit), an FPGA (Field-Programmable Gate Array), or a GPU (Graphics Processing Unit), or may be realized by cooperation between software and hardware. The program may be stored in advance in a storage device (a storage device with a non-transient storage medium) such as an HDD (Hard Disk Drive) or flash memory, or may be stored in a removable storage medium (non-transient storage medium) such as a DVD or CD-ROM and installed by inserting the storage medium into a drive device.

[0020] Fig. 2 is a diagram illustrating a schematic diagram of the function of each unit of the object detection device 100. Each unit of the object detection device 100 will be described below with reference to Fig. 2. The acquisition unit 110 acquires a captured image from the camera 10. The acquisition unit 110 stores (data of) the acquired captured image in a working memory such as a RAM (Random Access Memory).

[0021] The low-resolution image generating unit 120 performs a process such as thinning on the captured image to generate a low-resolution image with lower image quality than the captured image. The low-resolution image is, for example, an image with fewer pixels than the captured image.

[0022] The mask region determination unit 130 determines a mask region that is not to be processed by the grid definition unit 140 and the subsequent components. This will be described in detail later.

[0023] The grid definition unit 140 defines a plurality of partial region sets in the low-resolution image. Each of the plurality of partial region sets is defined by cutting out a plurality of partial regions (hereinafter, grids) from the low-resolution image. The partial regions are set, for example, in a rectangular shape with no gaps. The grid definition unit 140 defines the plurality of partial region sets such that the number of pixels in the grid differs among the plurality of partial region sets. Hereinafter, the plurality of partial region sets may be referred to as a first partial region set PA1, a second partial region set PA2, ..., a k-th partial region set PAk. Detailed functions of the grid definition unit 140 will be described later.

[0024] The extraction unit 150 derives an aggregate value by tallying up the differences in feature amounts between the grids included in each of the plurality of partial region sets and the surrounding grids, and adds up the aggregate values ​​between the plurality of partial region sets to extract a point of interest (a point of discontinuity with the surroundings in the figure). The detailed functions of each unit of the extraction unit 150 will be described later.

[0025] The high-resolution processing unit 170 cuts out a portion of the captured image corresponding to the location of interest (synchronous cutout in the figure), and performs high-resolution processing on the cutout to determine whether or not the object on the road is an object with which the vehicle should avoid contact. For example, the high-resolution processing unit 170 uses a trained model that recognizes road markings (an example of an object that is not an object with which the vehicle should avoid contact) and fallen objects (an example of an object with which the vehicle should avoid contact) from the image to determine whether the image captured at the location of interest is a road marking, a fallen object, or is unclear (an untrained object). At this time, the high-resolution processing unit 170 may further narrow down the processing to a portion of the location of interest in the captured image that is recognized as corresponding to a road marking or a fallen object.

[0026] FIG. 3 is a diagram for explaining the processing of the mask region determination unit 130, the grid definition unit 140, and the extraction unit 150. For example, the mask region determination unit 130 extracts edge points in the left and right direction in a low-resolution image, and detects the positions of road division lines, road shoulders, etc. (white lines, lane boundaries) in the image by connecting the edge points arranged in a straight line. Then, the mask region determination unit 130 detects an area sandwiched between the left and right road division lines, etc. and including a center point in the left and right direction on the near side of the image as the vehicle's lane. Next, the mask region determination unit 130 determines a portion other than the vehicle's lane (a portion above the vanishing point where the road division lines, etc. intersect on the far side, and a portion closer to the left and right ends than the road division lines) as the mask region. The grid definition unit 140 and the extraction unit 150 perform processing excluding the mask region.

[0027] As described above, the grid definition unit 140 defines partial region sets in the order of the first partial region set PA1 having the largest number of grid pixels, the second partial region set PA2 having the next largest number of grid pixels, and so on up to the k-th partial region set PAk having the smallest number of grid pixels. "Define" means to determine the grid boundary line for the low-resolution image.

[0028] The processing of the feature amount difference calculation unit 152, the counting unit 154, and the first addition unit 156 will be described below. The processing of these functional units described with reference to Figs. 4 to 7 is performed after first selecting one partial region set and selecting one grid of interest from the selected partial region set at a time. Then, when all grids in the selected partial region set are selected as grids of interest and processing is completed, the next partial region set is selected and processing is performed in the same manner. When processing is completed for all partial region sets, the second addition unit 158 ​​adds up the count values ​​(second count value V2 described later) between the partial region sets to generate extraction target data PT, which is a single image, and passes it to the focus portion extraction unit 160.

[0029] FIG. 4 is a diagram for explaining the processing of the feature amount difference calculation unit 152, the aggregation unit 154, and the first addition unit 156. The feature amount difference calculation unit 152 calculates the difference in the feature amount for each pixel between the comparison target grid and the comparison source grid. The feature amount is, for example, a luminance value for each component of R, G, and B, and a set of R, G, and B is one pixel. The comparison target grid and the comparison source grid are selected from the grid of interest and the surrounding grids. FIG. 5 is a diagram showing an example of the definition of the surrounding grid. As shown in the figure, grids 2 to 9 adjacent to the grid of interest in the up, down, left, right, and diagonal directions are defined as the surrounding grids. The method of selecting the surrounding grids (surrounding partial regions) is not limited to this, and the grids above, below, left, and right may be selected as the surrounding grids, or the surrounding grids may be selected according to another rule.

[0030] The comparison target grid and the comparison source grid are selected, for example, in order from the combinations shown in FIG. 6. FIG. 6 is a diagram showing an example of a rule for selecting the comparison target grid and the comparison source grid. The comparison target grid is a grid of interest, and the comparison source grids are selected in order from grids 2 to 9. The relationship between the comparison target grid and the comparison source grid may be reversed. Then, the counting unit 154 calculates a first count value V1 by finding the sum of the differences in the feature amounts for each pixel and dividing it by the number of pixels n in the grid. If the grid of interest corresponds to a mask area, the first count value V1 is replaced with zero and output. In other words, the feature difference calculation unit 152, the aggregation unit 154, and the first addition unit 156 process grid 1 as the comparison target grid and grid 3 as the comparison source grid, process grid 1 as the comparison target grid and grid 8 as the comparison source grid, process grid 1 as the comparison target grid and grid 5 as the comparison source grid, process grid 1 as the comparison target grid and grid 6 as the comparison source grid, process grid 1 as the comparison target grid and grid 2 as the comparison source grid, process grid 1 as the comparison target grid and grid 4 as the comparison source grid, process grid 1 as the comparison target grid and grid 7 as the comparison source grid, and process grid 1 as the comparison target grid and grid 9 as the comparison source grid, in parallel or sequentially.

[0031] The comparison target grid and the comparison source grid may be selected in order from the combinations shown in Fig. 7. Fig. 7 is a diagram showing another example of the rules for selecting the comparison target grid and the comparison source grid. The combination of the comparison target grid and the comparison source grid is not limited to the combination of the grid of interest and the surrounding grid, but may also include combinations of surrounding grids (particularly, combinations of the upper grid and the lower grid, the left grid and the right grid, the upper left grid and the lower right grid, and the upper right grid and the lower left grid).

[0032] More specifically, a method for calculating the difference in feature amounts will be described. As a method for calculating the difference in feature amounts, for example, the following Pattern 1 to Pattern 4 are possible. In the following description, the identification number of a pixel in each of the comparison target grid and comparison source grid is represented by i (i=1 to k; k is the number of pixels in each of the comparison target grid and comparison source grid).

[0033] (Pattern 1) The feature amount difference calculation unit 152 calculates, for example, the R component luminance difference ΔRi, the G component luminance difference ΔGi, and the B component luminance difference ΔBi between pixels at the same position in both the comparison target grid and the comparison source grid (where i=1 to k as described above). 2 +ΔGi 2 +ΔBi 2 is calculated for each pixel, and the maximum value or average value of each pixel feature amount Ppi is calculated as the difference in feature amount between the comparison target grid and the comparison source grid.

[0034] (Pattern 2) The feature amount difference calculation unit 152 calculates, for example, the R component luminance statistical value (mean value, median, mode, etc.) Raa, the G component luminance statistical value (same) Gaa, and the B component luminance statistical value (same) Baa for each pixel in the comparison target grid, calculates the R component luminance statistical value (same) Rab, the G component luminance statistical value (same) Gab, and the B component luminance statistical value (same) Bab for each pixel in the comparison source grid, and obtains the differences ΔRa (=Raa-Rab), ΔGa (=Gaa-Gab), and ΔBa (=Baa-Bab) between them.Then, ΔRa, which is the sum of squares of the luminance differences, is calculated. 2 +ΔGa 2 +ΔBa 2 , or the maximum value of the squared difference in brightness Max(ΔRa 2 ,ΔGa 2 ,ΔBa 2 ) is calculated as the difference in the feature amounts between the comparison target grid and the comparison source grid.

[0035] (Pattern 3) The feature amount difference calculation unit 152 calculates, for example, for each pixel i in the comparison target grid, a first index value W1ai (=(RB) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the B component by the sum of the luminance of each of the R, G, and B components, and a second index value W2ai (=(RG) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the G component by the sum of the luminance of each of the R, G, and B components. In addition, the feature amount difference calculation unit 152 calculates, for example, for each pixel i in the comparison source grid, a first index value W1bi (=(RB) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the B component by the sum of the luminance of each of the R, G, and B components, and a second index value W2bi (=(RG) / (R+G+B)) obtained by dividing the difference in luminance between the R component and the G component by the sum of the luminance of each of the R, G, and B components. Next, the feature amount difference calculation unit 152 calculates each pixel feature amount Ppi=(W1ai-W1bi) 2 +(W2ai-W2bi) 2 Then, the feature amount difference calculation unit 152 calculates the maximum or average value of each pixel feature amount Ppi as the difference between the feature amounts of the comparison target grid and the comparison source grid. The balance of the RGB components in each pixel can be expressed by combining the first index value and the second index value. Using the same concept as above, for example, the luminance of each of the R, G, and B components may be defined as the magnitude of a vector shifted by 120 degrees, and the vector sum may be used in the same way as the combination of the first index value and the second index value.

[0036] (Pattern 4) The feature difference calculation unit 152, for example, calculates a statistical value (same) Raa of the R component luminance, a statistical value (same) Gaa of the G component luminance, and a statistical value (same) Baa of the B component luminance for each pixel in the comparison target grid, and calculates a statistical value (same) Rab of the R component luminance, a statistical value (same) Gab of the G component luminance, and a statistical value (same) Bab of the B component luminance for each pixel in the comparison source grid. Next, the feature difference calculation unit 152 calculates a third index value W3a (=(Raa-Baa) / (Raa+Gaa+Baa)) obtained by dividing the difference between the R component luminance statistical value Raa and the B component luminance statistical value Baa by the sum of the luminance statistical values ​​of the R, G, and B components, and a fourth index value W4a (=(Raa-Gaa) / (Raa+Gaa+Baa)) obtained by dividing the difference between the R component luminance statistical value Raa and the G component luminance statistical value Gaa by the sum of the luminance statistical values ​​of the R, G, and B components. Similarly, for the comparison source grid, the feature amount difference calculation unit 152 calculates a third index value W3b (=(Rab-Bab) / (Rab+Gab+Bab)) by dividing the difference between the R component luminance statistical value Rab and the B component luminance statistical value Bab by the sum of the R, G, and B component luminance statistical values, and a fourth index value W4b (=(Rab-Gab) / (Rab+Gab+Bab)) by dividing the difference between the R component luminance statistical value Rab and the G component luminance statistical value Gab by the sum of the R, G, and B component luminance statistical values. Then, the feature amount difference calculation unit 152 calculates a difference ΔW3 between the third index value W3a of the comparison target grid and the third index value W3b of the comparison source grid, and a difference ΔW4 between the fourth index value W4a of the comparison target grid and the fourth index value W4b of the comparison source grid, and calculates the sum of their squares, ΔW3 2 +ΔW4 2 , or the maximum of the squares Max(ΔW3 2 ,ΔW4 2 ) is calculated as the difference in the feature amounts between the comparison target grid and the comparison source grid.

[0037] If the image to be processed is a black-and-white image, feature difference calculation unit 152 may simply calculate the difference in luminance values ​​as the difference in features between the comparison target grid and the comparison source grid. Also, if the image to be processed is an RGB image, it may convert the RGB image into a black-and-white image and calculate the difference in trajectory values ​​as the difference in features between the comparison target grid and the comparison source grid.

[0038] Returning to Fig. 4, the first adder 156 calculates a second aggregate value V2 by adding up the first aggregate value V1 calculated corresponding to the grid of interest. The second aggregate value V2 is an example of an "aggregate value" in the claims. When the process of calculating the second aggregate value V2 while changing the grid of interest is completed, data is generated in which the second aggregate value V2 is set for all grids for each partial region set.

[0039] When data in which the second total value V2 is set for all grids is generated, the second adder 158 adds them between the partial region sets to generate extraction target data PT. FIG. 8 is a diagram for explaining the processing of the first adder 156 and the second adder 158. In the figure, the smallest rectangle is one pixel of the low-resolution image. Here, in order to simplify the explanation, it is assumed that there are three partial region sets, namely, the first partial region set PA1, the second partial region set PA2, and the third partial region set PA3, and their sizes are also assumed to be much smaller than the actual sizes. In addition, it is assumed that the second total value V2 is normalized at some stage so as to be a value between zero and one. In the illustrated example, the first partial region set PA1 is a set of first grids consisting of 16 pixels, the second partial region set PA2 is a set of second grids consisting of 9 pixels, and the third partial region set PA3 is a set of third grids consisting of 4 pixels. Since the second total value V2 is set for a grid that bundles a plurality of pixels, the second adder 158 expands the second total value V2 set for the grid to all pixels in the grid, and then adds the pixel values ​​of the pixels between the partial region sets to generate the extraction target data PT. Note that there may be partial region sets in which pixels and grids correspond one-to-one, and in that case, it is not necessary to expand the second total value V2 for each pixel. Note that, although examples of grid sizes of 16 pixels, 9 pixels, and 4 pixels have been shown, when the above pattern 2 or pattern 4 is adopted as the calculation method for the difference in the feature amount, it is better to set the grid size to a power of 2 in the number of pixels on one side, such as 4 pixels, 16 pixels, or 64 pixels, in order to reduce the calculation load of the statistical value.

[0040] The portion of interest extraction section 160 extracts, from the processing target data PT, a circumscribing rectangle that includes a group of portions whose sum of the aggregate values ​​is equal to or greater than a threshold value, as a portion of interest.

[0041] As described above, the high resolution processing unit 170 performs high resolution processing on the area obtained by fitting only the position of the point of interest to the captured image, and determines whether or not the object on the road is an object with which the vehicle should avoid contact.

[0042] The discrimination result of the high resolution processing unit 170 is output to the driving control device 200 and / or the notification device 210. The driving control device 200 performs automatic brake control, automatic steering control, etc. to avoid contact between the vehicle and the object (actually an area on the image) discriminated as a "falling object". The notification device 210 outputs an alarm in various ways when the TTC (Time To Collision) between the object (same as above) discriminated as a "falling object" and the vehicle falls below a threshold.

[0043] According to the embodiment described above, by providing an acquisition unit 110 that acquires an image of at least the road in the direction of travel of the vehicle, a low-resolution image generation unit 120 that generates a low-resolution image by reducing the image quality of the captured image, a grid definition unit 140 that defines one or more partial area sets, and an extraction unit 150 that derives an aggregate value by aggregating differences in features between surrounding partial areas for each partial area included in one or more partial area sets, and extracts points of interest based on the aggregate value, it is possible to maintain high detection accuracy while reducing the processing load.

[0044] If the processing performed by feature amount difference calculation unit 152 and aggregation unit 154 were to be performed on the captured image as is, the processing load would increase as the number of pixels increases, raising concerns that the operation of driving control device 200 and alarm device 210 would not be able to keep up with the approach of a falling object. In response to this issue, the object detection device 100 of the embodiment generates a low-resolution image and then performs processing, thereby making it possible to detect objects while reducing the processing load.

[0045] Furthermore, according to the embodiment, the grid definition unit 140 defines a plurality of partial area sets so that the number of pixels in the grid is different between the plurality of partial area sets, and the extraction unit 150 extracts the target portion by adding the aggregated value for each pixel between the plurality of partial area sets, so that the robustness of the detection performance against the variation in the size of the falling object can be improved. If the processing is simply performed with a low-resolution image, there is a concern that the image quality will be reduced to a level where the presence of the falling object cannot be recognized, but according to the embodiment, the above-mentioned ingenuity makes it possible to expect that the falling object will appear as a feature in a grid of any size. As described above, according to the object detection device 100 of the embodiment, it is possible to maintain high detection accuracy while reducing the processing load.

[0046] The above-described embodiment can be expressed as follows. A storage device storing a program; a hardware processor; The hardware processor executes the program stored in the storage device, Obtaining an image of a road surface; generating a low-resolution image by reducing the image quality of the captured image; Defining one or more subregion sets; deriving an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the one or more partial region sets and surrounding partial regions, and extracting a portion of interest based on the aggregate value; Each of the one or more sub-region sets is defined by extracting a plurality of sub-regions from the low-resolution image. The object detection device is configured as follows.

[0047] The above describes the form for carrying out the present invention using an embodiment, but the present invention is not limited to such an embodiment, and various modifications and substitutions can be made within the scope that does not deviate from the gist of the present invention. [Explanation of symbols]

[0048] 10. Camera 100 Object detection device 110 Acquisition Department 120 Low-resolution image generator 130 Mask area determination unit 140 Grid definition section 150 Extraction part 152 Feature Difference Calculation Unit 154 Counting Department 156 First Addition Unit 158 Second Addition Unit 160 Extraction of focus area 170 High Resolution Processing Unit

Claims

1. An acquisition unit that acquires an image of a road surface; a low-resolution image generating unit that generates a low-resolution image by reducing the image quality of the captured image; A definition unit that defines one or more partial region sets, each of the one or more partial region sets being defined by cutting out a plurality of partial regions from the low-resolution image; and an extraction unit that derives an aggregate value by aggregating differences in feature amounts between the partial regions included in each of the one or more partial region sets and surrounding partial regions, and extracts a portion of interest based on the aggregate value; a high-resolution processing unit that performs high-resolution processing on a portion of the captured image corresponding to the location of interest and determines whether an object on a road included in the portion of the captured image corresponding to the location of interest is an object with which a moving body should avoid contact; An object detection device comprising:

2. the definition unit defines a plurality of partial region sets such that the number of pixels in the partial region differs among the plurality of partial region sets; the extraction unit extracts the portion of interest by adding up the aggregated values ​​for each pixel among a plurality of the partial region sets. The object detection device according to claim 1.

3. the extraction unit derives the aggregated value by aggregating differences in feature amounts between the partial regions included in each of the one or more partial region sets and other partial regions adjacent thereto vertically, horizontally, and diagonally; 3. The object detection device according to claim 1 or 2.

4. the extraction unit further adds, for each of the partial regions included in the one or more partial region sets, a difference in feature amount between the partial regions adjacent vertically, a difference in feature amount between the partial regions adjacent horizontally, and a difference in feature amount between the partial regions adjacent diagonally to the set of partial regions to the aggregated value; The object detection device according to claim 3.

5. 1. A computer-implemented method for object detection, comprising: Obtaining an image of a road surface; generating a low-resolution image by reducing the image quality of the captured image; Defining one or more subregion sets; deriving an aggregate value by aggregating differences in feature amounts between each of the partial regions included in the one or more partial region sets and surrounding the partial regions, and extracting a portion of interest based on the aggregate value; each of the one or more partial region sets is defined by cutting out a plurality of the partial regions from the low-resolution image; The method further includes performing high-resolution processing on a portion of the captured image corresponding to the location of interest, and determining whether or not an object on a road included in the portion of the captured image corresponding to the location of interest is an object with which the moving body should avoid contact. Object detection methods.

6. A program executed by a computer, the program comprising: Acquiring an image of at least a road surface in a traveling direction of a moving object; generating a low-resolution image by reducing the image quality of the captured image; Defining one or more subregion sets; deriving an aggregate value by aggregating differences in feature amounts between partial regions included in each of the one or more partial region sets and surrounding partial regions, and extracting a portion of interest based on the aggregate value; each of the one or more partial region sets is defined by cutting out a plurality of the partial regions from the low-resolution image; and causing the computer to perform high-resolution processing on a portion of the captured image corresponding to the location of interest, and to determine whether an object on a road included in the portion of the captured image corresponding to the location of interest is an object with which a moving body should avoid contact. program.

Citation Information

Patent Citations

  • External monitoring device having fail / safe function

    JP2001028746A

  • Device and method for reading number plate

    JP2001273461A

  • Image processing system, control method and control program of the same

    JP2014153866A

  • Image processing device, image processing method, and image processing system

    JP2018107759A

  • Failure detection system

    JP2019124986A