Object detection device

The object detection device addresses inefficiencies in setting multiple overlapping image regions by employing a two-stage reduction process, significantly reducing computational load and enhancing detection efficiency.

WO2025197121A1PCT designated stage Publication Date: 2025-09-25SUBARU CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/011503
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing object detection systems face inefficiencies due to the setting of multiple overlapping image regions for a single subject, leading to excessive computational requirements in processes like Non-Maximum Suppression (NMS).

Method used

An object detection device employs a two-stage reduction process: a pre-reduction process that reduces image regions based on overlap and reliability, followed by an NMS process, to set one image region per subject, thereby minimizing computational load.

Benefits of technology

The device effectively reduces the number of image regions with a small amount of calculation, optimizing computational efficiency and enabling accurate subject detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024011503_25092025_PF_FP_ABST
    Figure JP2024011503_25092025_PF_FP_ABST
Patent Text Reader

Abstract

An object detection device according to an embodiment of the present disclosure comprises a processing circuit capable of performing a first process for setting a plurality of image regions on the basis of a captured image and calculating the reliability of each of the plurality of image regions, and a second process for reducing the image regions. The second process includes first and second reduction processes. The first reduction process includes: sequentially selecting, as a first image region, a plurality of image regions in order beginning with an image region for which a first position corresponding to a first region end section in a prescribed direction is closest to an image end section that is one prescribed-direction end in the captured image; selecting, as a second image region, an image region for which a second position corresponding to a second region end section is farthest from the image end section among one or a plurality of image regions selected as the first image region in the past; and determining an image region having a lower reliability among the first and second image regions as a first object to be deleted, in cases where a first degree of overlap of the first and second image regions is equal to or greater than a first threshold value.
Need to check novelty before this filing date? Find Prior Art

Description

Object detection device

[0001] The present disclosure relates to an object detection device that detects an object based on a captured image.

[0002] There is a technique for detecting a subject based on a captured image including the subject. In this detection process, multiple overlapping image regions may often be set for a single subject. For example, Patent Document 1 discloses a technique for performing filtering by calculating an Intersection Over Union (IoU) value.

[0003] JP 2020-205039 A

[0004] An object detection device according to an embodiment of the present disclosure includes a processing circuit. The processing circuit is capable of performing a first process of setting a plurality of image regions corresponding to a plurality of subjects and calculating a reliability of each of the plurality of image regions based on a captured image including images of the plurality of subjects, and a second process of reducing the image regions based on a degree of overlap among the plurality of image regions. The second process includes a first reduction process and a second reduction process. The first reduction process includes sequentially selecting as the first image area from the multiple image areas set by the first process, starting with the image area whose first position corresponding to the first area end in a predetermined direction is closest to the image end, which is one end of the predetermined direction in the captured image; selecting as the second image area, from one or more image areas previously selected as the first image area, the image area whose second position corresponding to the second area end in the predetermined direction is farthest from the image end; calculating a first degree of overlap between the first image area and the second image area; if the first degree of overlap is equal to or greater than a first threshold, designating the image area with a low reliability from the first image area and the second image area as a first deletion target; and reducing the image area based on the first deletion target. The second reduction process includes sequentially selecting, as a third image area, the multiple image areas reduced by the first reduction process in order of reliability; calculating a second degree of overlap between the third image area and one or more image areas that have not yet been selected as the third image area; and removing, from the one or more image areas that have not yet been selected, one or more image areas whose second degree of overlap is equal to or greater than a second threshold value, thereby reducing the image area.

[0005] The accompanying drawings are included to provide a further understanding of the disclosure, and are incorporated in and constitute a part of this specification. The drawings illustrate one embodiment and, together with the description, serve to explain the principles of the disclosure.

[0006] FIG. 1 is an explanatory diagram illustrating an example configuration of a vehicle equipped with an object detection device according to an embodiment of the present disclosure. FIG. 2 is a block diagram illustrating an example configuration of the object detection device illustrated in FIG. 1. FIG. 3 is an explanatory diagram illustrating an example of a captured image generated by the imaging device illustrated in FIG. 1. FIG. 4 is an explanatory diagram illustrating an example operation of the image region setting unit illustrated in FIG. 2. FIG. 5 is an explanatory diagram illustrating an example operation of the image region reduction unit illustrated in FIG. 2. FIG. 6 is a flowchart illustrating an example operation of the processing device illustrated in FIG. 2. FIG. 7A is an explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7B is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7C is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7D is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7E is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7F is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7G is another explanatory diagram illustrating an example operation of the pre-processing unit illustrated in FIG. 2. FIG. 7H is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7I is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7J is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7K is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7L is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7M is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7N is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 7O is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 8 is an explanatory diagram illustrating an example of an intersection region and a union region in IoU determination. FIG. 9 is another explanatory diagram illustrating an example of an operation of the pre-processing unit shown in FIG. 2 . FIG. 10 is a flowchart illustrating an example of an operation of the pre-processing unit. FIG. 11A is an explanatory diagram illustrating an example of an operation of the NMS processing unit shown in FIG. 2 . FIG. 11B is another explanatory diagram illustrating an example of an operation of the NMS processing unit shown in FIG. 2 . Fig. 11C is another explanatory diagram showing an example of an operation of the NMS processing unit shown in Fig. 2. Fig. 11D is another explanatory diagram showing an example of an operation of the NMS processing unit shown in Fig. 2. Fig. 11E is another explanatory diagram showing an example of an operation of the NMS processing unit shown in Fig. 2.FIG. 11F is another explanatory diagram illustrating an example of an operation of the NMS processing unit shown in FIG. 2 . FIG. 11G is another explanatory diagram illustrating an example of an operation of the NMS processing unit shown in FIG. 2 . FIG. 12 is a flowchart illustrating an example of an operation of the NMS processing unit. FIG. 13 is a flowchart illustrating an example of an operation of a processing device according to a modified example. FIG. 14 is a flowchart illustrating an example of an operation of a pre-processing unit according to a modified example. FIG. 15 is another flowchart illustrating an example of an operation of the pre-processing unit according to a modified example. FIG. 16 is an explanatory diagram illustrating an example of an image region. FIG. 17 is an explanatory diagram illustrating an example of an operation of a processing device according to a modified example. FIG. 18 is a flowchart illustrating an example of an operation of a processing device according to another modified example. FIG. 19 is a flowchart illustrating an example of an operation of the pre-processing unit according to another modified example. FIG. 20 is an explanatory diagram illustrating an example of an operation of the pre-processing unit according to another modified example. FIG. 21 is another flowchart illustrating an example of an operation of the pre-processing unit according to another modified example. FIG. 22 is an explanatory diagram illustrating another example of an image region. FIG. 23 is an explanatory diagram illustrating an example of an operation of a processing device according to another modified example.

[0007] When detecting a subject based on a captured image, multiple overlapping image regions may be set for one subject. Therefore, in an object detection device, the image regions are reduced so that one image region is set for one subject. Such an object detection device is expected to reduce the image regions with a small amount of calculation.

[0008] It is desirable to provide an object detection device that can reduce the image area with a small amount of calculation.

[0009] Some exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Note that the following description illustrates one specific example of the present disclosure and should not be construed as limiting the present disclosure. For example, each element, including numerical values, shapes, materials, parts, the position of each part, and the connection method of each part, is merely an example and should not be construed as limiting the present disclosure. Furthermore, in the following exemplary embodiments, components not described in independent claims based on the highest concept of the present disclosure are optional and may be provided as needed. The drawings are schematic and are not intended to be drawn to scale. Throughout this specification and the drawings, components having substantially the same function and configuration are designated by the same reference numerals, and redundant description will be omitted. Furthermore, components not directly related to one embodiment of the present disclosure are not shown in the drawings.

[0010] <Embodiment> [Configuration Example] Fig. 1 shows an example configuration of a vehicle 1 equipped with an object detection device according to an embodiment. The vehicle 1 is a vehicle such as an automobile, and is equipped with an object detection device 10. The object detection device 10 is configured to detect objects around the vehicle 1. The object detection device 10 includes an imaging device 11 and a processing device 20.

[0011] The imaging device 11 is configured to generate captured images by capturing images of the area ahead of the vehicle 1. The imaging device 11 may be a monocular camera or a stereo camera. The imaging device 11 includes a lens and an image sensor. In this example, the imaging device 11 is disposed inside the vehicle 1 near the top of the windshield of the vehicle 1. The imaging device 11 generates a series of captured images by performing an imaging operation at a predetermined frame rate (e.g., 60 fps). The imaging device 11 then supplies the generated captured images to the processing device 20.

[0012] The processing device 20 is configured to detect subjects such as vehicles and people ahead of the vehicle 1 based on the captured image supplied from the imaging device 11, and to perform driving assistance processing based on the detection results. The processing device 20 is configured to include, for example, one or more processors and one or more memories, and is configured to perform processing by executing a program.

[0013] 2 shows an example of the configuration of the processing device 20. The processing device 20 has an image area setting unit 21, an image area reduction unit 22, and a driving support processing unit 25.

[0014] The image region setting unit 21 is configured to set an image region R corresponding to a subject based on a captured image supplied from the imaging device 11 and calculate a score SC for the image region R. Specifically, the image region setting unit 21 sets the image region R by searching for a region that is likely to include an image of a predetermined detection target, such as a vehicle, bicycle, or person, based on the captured image using a machine learning technique such as a deep neural network (DNN). The score SC can take a value between 0 and 1, for example. The higher the likelihood that the image region R includes an image of the predetermined detection target, such as a vehicle, bicycle, or person, the higher the score SC. In other words, the score SC is a reliability parameter that indicates the degree to which the image region R can be trusted as a region that includes an image of such a predetermined detection target.

[0015] 3 shows an example of a captured image supplied to the image area setting unit 21. In this example, the captured image includes, for example, images of a plurality of vehicles traveling ahead of the vehicle 1.

[0016] Fig. 4 shows an example of the operation of the image area setting unit 21. In this example, the image area setting unit 21 sets an image area R based on the captured image shown in Fig. 3. The image area setting unit 21 sets a plurality of overlapping image areas R for each of a plurality of vehicles.

[0017] The image region reduction unit 22 (FIG. 2) is configured to reduce the image region R set by the image region setting unit 21 using the score SC of the image region R and the degree of overlap of the image region R.

[0018] Fig. 5 shows an example of the operation of the image region reduction unit 22. As shown in Fig. 4, the image region setting unit 21 can set multiple overlapping image regions R for one subject. Therefore, the image region reduction unit 22 identifies the multiple overlapping image regions R based on the degree of overlap of the image regions R, and reduces the image regions R so that one image region R is set for one subject.

[0019] In this way, in the processing device 20, the image region setting unit 21 allows multiple image regions R to be set for one subject, thereby setting an excessive number of image regions R. Then, the image region reduction unit 22 reduces the image regions R so that one image region R is set for one subject. In this way, the processing device 20 sets one image region R for one subject while setting image regions R for all subjects without omission.

[0020] The image area reduction unit 22 includes a pre-processing unit 23 and an NMS (Non Maximum Suppression) processing unit 24 .

[0021] The pre-processing unit 23 is configured to perform a pre-reduction process A for reducing the image region R in advance based on the image region R set by the image region setting unit 21 before the NMS processing unit 24 processes the image region R.

[0022] The NMS processing unit 24 is configured to perform an NMS reduction process B that reduces the image region R using an NMS algorithm based on the image region R that has been processed by the pre-processing unit 23 .

[0023] That is, as will be described later, the NMS algorithm requires a large amount of calculation in processing, so in the processing device 20, the pre-processing unit 23 first reduces the image region R, and the NMS processing unit 24 further reduces the image region R using the NMS algorithm based on the multiple image regions R after processing by the pre-processing unit 23. In this way, the object detection device 10 is able to set one image region R for one subject while reducing the amount of calculation, as will be described later.

[0024] The driving assistance processing unit 25 is configured to detect the type and position of the subject based on the image area R after processing by the image area reduction unit 22, and based on the detection result, to assist the driver in driving the vehicle 1. Specifically, the driving assistance processing unit 25 is configured to control the operation of the vehicle 1 so as to warn the driver or brake the vehicle 1 when there is a pedestrian in front of the vehicle 1, for example.

[0025] Here, the processing device 20 corresponds to a specific example of a "processing circuit" in an embodiment of the present disclosure. The processing performed by the image region setting unit 21 corresponds to a specific example of a "first processing" in an embodiment of the present disclosure. The processing performed by the image region reduction unit 22 corresponds to a specific example of a "second processing" in an embodiment of the present disclosure. The pre-reduction processing A corresponds to a specific example of a "first reduction processing" in an embodiment of the present disclosure. The NMS reduction processing B corresponds to a specific example of a "second reduction processing" in an embodiment of the present disclosure. The score SC corresponds to a specific example of a "reliability" in an embodiment of the present disclosure. The IoU value corresponds to a specific example of a "degree of overlap" in an embodiment of the present disclosure.

[0026] [Operation and Function] Next, the operation and function of the object detection device 10 of this embodiment will be described.

[0027] (Overall Operation Overview) First, the operation of the object detection device 10 will be described with reference to Figures 1 and 2. The imaging device 11 generates a captured image by capturing an image of the area ahead of the vehicle 1. The image area setting unit 21 of the processing device 20 sets an image area R corresponding to a subject based on the captured image supplied from the imaging device 11, and calculates a score SC for the image area R. The image area setting unit 21 can set multiple overlapping image areas R for one subject. The image area reduction unit 22 reduces the image area R so that one image area R is set for one subject. The driving assistance processing unit 25 detects the type and position of the subject based on the image area R processed by the image area reduction unit 22, and provides assistance to the driver in driving the vehicle 1 based on the detection result.

[0028] 6 shows an example of the operation of the processing device 20 in the object detection device 10. The processing device 20 performs the process shown in FIG. 6 every time a captured image is supplied from the imaging device 11.

[0029] First, the image region setting unit 21 of the processing device 20 sets an image region R based on a captured image and calculates a score SC for the image region R (step S1). Specifically, the image region setting unit 21 sets a plurality of image regions R corresponding to a plurality of subjects, as shown in Fig. 4, based on a captured image including images of a plurality of subjects, as shown in Fig. 3, for example. The image region setting unit 21 also calculates a score SC for each of the plurality of image regions R.

[0030] Next, the image region setting unit 21 performs a filtering process based on the score SC (step S2). Specifically, the image region setting unit 21 performs a filtering process to delete, for example, image regions R having a score SC lower than a predetermined value from among the generated multiple image regions R. That is, if the score SC of an image region R is sufficiently low, the image region R is unlikely to include an image of a predetermined detection target such as a vehicle, bicycle, or person, and therefore the image region setting unit 21 deletes the image region R and excludes the image region R from the processing target.

[0031] Next, the pre-processing unit 23 of the image area reduction unit 22 performs pre-reduction processing A (step S3). Specifically, the pre-processing unit 23 reduces the image area R in advance, based on the image area R set by the image area setting unit 21, before processing by the NMS processing unit 24. This pre-reduction processing A will be described in detail later.

[0032] Next, the NMS processing unit 24 of the image area reduction unit 22 performs NMS reduction processing B (step S4). Specifically, the NMS processing unit 24 reduces the image area R using an NMS algorithm based on the image area R processed by the pre-processing unit 23. This NMS reduction processing B will be described in detail later. As a result, for example, one image area R is set for one subject.

[0033] Then, the driving assistance processing unit 25 performs driving assistance processing based on the image region R after processing by the image region reduction unit 22 (step S5). Specifically, the driving assistance processing unit 25 detects the type and position of the subject based on the image region R after processing by the image region reduction unit 22, and provides assistance to the driver in driving the vehicle 1 based on the detection result.

[0034] This completes the process.

[0035] Next, the advance reduction process A and the NMS reduction process B will be described in detail.

[0036] 7A to 7O show a specific example of the advance reduction process A. In this example, the advance processing unit 23 of the image area reduction unit 22 performs a process of reducing the image area R using the unregistered buffer B1 and the temporary registered buffer B2.

[0037] First, as shown in FIG. 7A, the pre-processing unit 23 stores the plurality of image regions R set by the image region setting unit 21 in the unregistered buffer B1.

[0038] Next, as shown in Fig. 7B, the pre-processing unit 23 extracts the image region R whose left end is the leftmost from among the multiple image regions R in the unregistered buffer B1. That is, the multiple image regions R in the unregistered buffer B1 are arranged at various positions in the captured image, as shown in Fig. 4, for example. The pre-processing unit 23 extracts the image region R whose left end is the leftmost from among the multiple image regions R arranged at various positions in this way from the unregistered buffer B1.

[0039] In this example, as shown in FIG. 7B, the temporary registration buffer B2 is empty, so the pre-processing unit 23 stores the image region R extracted from the unregistered buffer B1 in the temporary registration buffer B2 as shown in FIG. 7C.

[0040] Next, as shown in FIG. 7D, the pre-processing unit 23 extracts from the unregistered buffer B1 the image region R whose left end is located at the leftmost position among the plurality of image regions R in the unregistered buffer B1.

[0041] 7E, the pre-processing unit 23 extracts one image region R from the temporary registration buffer B2. In this example, since the temporary registration buffer B2 contains only one image region R, the pre-processing unit 23 extracts this image region R from the temporary registration buffer B2. The pre-processing unit 23 then performs an Intersection over Union (IoU) determination based on the image region R extracted from the non-registration buffer B1 and the image region R extracted from the temporary registration buffer B2.

[0042] FIG. 8 shows an example of IoU determination, where (A) shows the intersection region RI and (B) shows the union region RU. This example illustrates IoU determination based on image region R11 and image region R12. In this example, a portion of image region R11 and a portion of image region R12 overlap each other. This overlapping region is the intersection region RI. The entire region of image region R11 and image region R22 is the union region RU. The pre-processing unit 23 calculates the IoU values ​​of image region R11 and image region R12 by dividing the area of ​​the intersection region RI by the area of ​​the union region RU. This IoU value can take a value between 0 and 1. The wider the intersection region RI and the narrower the union region RU, the higher the IoU value.

[0043] 7F , if the IoU values ​​of two image regions R are lower than a predetermined threshold value TH1, the pre-processing unit 23 stores these two image regions R in the temporary registration buffer B2. That is, because the IoU values ​​of the two image regions R are low, the pre-processing unit 23 determines that these two image regions R correspond to different subjects, and stores these two image regions R in the temporary registration buffer B2.

[0044] Next, as shown in FIG. 7G, the pre-processing unit 23 extracts from the unregistered buffer B1 the image region R whose left end is located at the leftmost position among the plurality of image regions R in the unregistered buffer B1.

[0045] Next, as shown in FIG. 7H , the pre-processing unit 23 retrieves one image region R from the temporary registration buffer B2. In this example, there are two image regions R in the temporary registration buffer B2. As will be described below, the pre-processing unit 23 retrieves the image region R whose right edge is the rightmost of the two image regions R in the temporary registration buffer B2 from the temporary registration buffer B2. Then, the pre-processing unit 23 performs IoU determination based on the image region R retrieved from the unregistered buffer B1 and the image region R retrieved from the temporary registration buffer B2.

[0046] 9 shows an example of the operation of extracting one image region R from the temporary registration buffer B2. Note that, although all image regions R are drawn at the same size in Fig. 9, the size of the image region R may vary depending on the subject. In Fig. 9, the horizontal axis indicates the position of the image region R in the left-right direction of the captured image.

[0047] As described above, the pre-processing unit 23 extracts the image region R whose left edge is the leftmost from the unregistered buffer B1. Therefore, as shown in Fig. 9, the left edge of the image region R extracted from the unregistered buffer B1 (image region R21) is located to the left of the left edges of all the image regions R in the unregistered buffer B1.

[0048] Furthermore, as described above, the pre-processing unit 23 extracts the image region R whose right edge is located at the rightmost position from the temporary registration buffer B2. Therefore, as shown in FIG. 9 , the right edge of the image region R (image region R22) extracted from the temporary registration buffer B2 is located to the right of the right edges of all the image regions R in the temporary registration buffer B2. As a result, this image region R22 is more likely to overlap in the horizontal direction with the image region R21 extracted from the unregistered buffer B1 than the other image regions R in the temporary registration buffer B2. In FIG. 9 , the image regions R21 and R22 overlap each other in the horizontal direction within a range of width W. The image region R22 is the image region R whose width W is likely to be the widest between the image region R21 and the image region R21. As shown in FIG. 4 , because the vehicle is on the road surface, the multiple image regions R in the captured image are aligned in the horizontal direction. Therefore, the image regions R21 and R22 are more likely to have the widest intersection region RI within the two-dimensional image plane of the captured image. In this way, the pre-processing unit 23 selects the image region R22 in which the intersection region RI between the image region R21 and the image region R22 is likely to be the widest, using only the left-right direction, as shown in FIG.

[0049] In this way, the pre-processing unit 23 extracts the image region R whose right edge is the rightmost of the two image regions R in the temporary registration buffer B2 from the temporary registration buffer B2. Then, the pre-processing unit 23 performs IoU determination based on the image region R extracted from the unregistered buffer B1 and the image region R extracted from the temporary registration buffer B2.

[0050] 7I, if the IoU values ​​of two image regions R are lower than a predetermined threshold TH1, the pre-processing unit 23 stores these two image regions R in the temporary registration buffer B2. That is, because the IoU values ​​of the two image regions R are low, the pre-processing unit 23 determines that these two image regions R correspond to, for example, different subjects, and stores these two image regions R in the temporary registration buffer B2.

[0051] 7J , when the IoU values ​​of two image regions R are equal to or greater than a predetermined threshold value TH1, the pre-processing unit 23 performs score determination based on these two image regions R. That is, since the IoU values ​​of the two image regions R are high, the pre-processing unit 23 determines that these two image regions R correspond to one subject, and performs score determination based on these two image regions R to delete one of the two image regions R.

[0052] 7K, if the score SC of the image region R retrieved from the temporary registration buffer B2 is higher than the score of the image region R retrieved from the unregistered buffer B1, as shown in FIG. 7L, the pre-processing unit 23 returns the image region R retrieved from the temporary registration buffer B2 to the temporary registration buffer B2, and deletes the image region R retrieved from the unregistered buffer B1 as a deletion target. That is, in this case, the image region R retrieved from the temporary registration buffer B2 is more likely to include an image of a predetermined detection target such as a vehicle, bicycle, or person than the image region R retrieved from the unregistered buffer B1, so the pre-processing unit 23 leaves the image region R retrieved from the temporary registration buffer B2 and deletes the image region R retrieved from the unregistered buffer B1.

[0053] 7M, if the score of image region R retrieved from unregistered buffer B1 is higher than the score SC of image region R retrieved from temporary registration buffer B2, as shown in FIG. 7N, the pre-processing unit 23 stores image region R retrieved from unregistered buffer B1 in temporary registration buffer B2, and deletes image region R retrieved from temporary registration buffer B2 as a deletion target. That is, in this case, image region R retrieved from unregistered buffer B1 is more likely to include an image of a predetermined detection target such as a vehicle, bicycle, or person than image region R retrieved from temporary registration buffer B2, so the pre-processing unit 23 leaves image region R retrieved from unregistered buffer B1 and deletes image region R retrieved from temporary registration buffer B2.

[0054] The pre-processing unit 23 repeats this operation thereafter, as shown in Fig. 7O, until the unregistered buffer B1 becomes empty.

[0055] This completes the advance reduction process A.

[0056] 10 shows an example of the advance reduction process A. At the start of this advance reduction process A, a plurality of image regions R set by the image region setting unit 21 are stored in the unregistered buffer B1.

[0057] First, the pre-processing unit 23 extracts the image region R whose left edge is located at the leftmost side from the non-registration buffer B1, and stores this image region R in the temporary registration buffer B2 (step S101).

[0058] Next, the pre-processing unit 23 extracts the image region R whose left end is located at the leftmost position from the unregistered buffer B1 (step S102).

[0059] Next, the pre-processing unit 23 checks whether the temporary registration buffer B2 contains one image region R (step S103). If there is one image region R ("Y" in step S103), the pre-processing unit 23 extracts the one image region R from the temporary registration buffer B2 (step S104). If there are two or more image regions R ("N" in step S103), the pre-processing unit 23 extracts the image region R whose right edge is the rightmost from the temporary registration buffer B2 (step S105).

[0060] Next, the pre-processing unit 23 calculates the IoU values ​​of the two image regions R retrieved from the unregistered buffer B1 and the temporary registration buffer B2 (step S106).

[0061] Next, the pre-processing unit 23 checks whether the IoU value calculated in step S106 is higher than a threshold value TH1 (step S107). If the IoU value is not higher than the threshold value TH1 ("N" in step S107), the pre-processing unit 23 stores these two image regions R in the temporary registration buffer B2 (step S108). If the IoU value is higher than the threshold value TH1 ("Y" in step S107), the pre-processing unit 23 stores the image region R with the higher score SC in the temporary registration buffer B2, and designates the image region R with the lower score SC as a deletion target and deletes this image region R.

[0062] Next, the pre-processing unit 23 checks whether the unregistered buffer B1 is empty (step S110). If the unregistered buffer B1 is not empty ("N" in step S110), the process returns to step S102, and the pre-processing unit 23 repeats steps S102 to S109 until the unregistered buffer B1 is empty. If the unregistered buffer B1 is empty ("Y" in step S110), the process ends.

[0063] In this way, the pre-processing unit 23 reduces the image region R.

[0064] Here, the IoU value in the pre-reduction process A corresponds to a specific example of a "first degree of overlap" in an embodiment of the present disclosure. The threshold value TH1 corresponds to a specific example of a "first threshold" in an embodiment of the present disclosure. The left edge of the image region R corresponds to a specific example of a "first region edge" in an embodiment of the present disclosure. The right edge of the image region R corresponds to a specific example of a "second region edge" in an embodiment of the present disclosure. The left edge of the captured image corresponds to a specific example of an "image edge" in an embodiment of the present disclosure.

[0065] 11A to 11G show a specific example of the NMS reduction process B. In this example, the NMS processing unit 24 of the image area reduction unit 22 performs a process of reducing the image area R using the unregistered buffer B3 and the registered buffer B4.

[0066] First, as shown in Fig. 11A, the NMS processing unit 24 stores the plurality of image regions R processed by the pre-processing unit 23 in the unregistered buffer B3. That is, the plurality of image regions R stored in the unregistered buffer B3 in Fig. 11A are the same as the plurality of image regions R stored in the temporary registration buffer B2 when the pre-reduction process A is completed.

[0067] 11B, the NMS processing unit 24 extracts the image region R having the highest score SC from the plurality of image regions R in the unregistered buffer B3 and stores this image region R in the registered buffer B4. That is, the NMS processing unit 24 stores the image region R that is most likely to include an image of a predetermined detection target, such as a vehicle, bicycle, or person, in the registered buffer B4 from the plurality of image regions R in the unregistered buffer B3.

[0068] Next, as shown in FIG. 11C , the NMS processing unit 24 performs an IoU determination based on the image region R most recently stored in the registered buffer B4 and each of the multiple image regions R in the unregistered buffer B3. As shown in FIGS. 11C and 11D , the NMS processing unit 24 deletes from the unregistered buffer B3 any image region R with an IoU value higher than a predetermined threshold value TH2 from the unregistered buffer B3. That is, the image region R with an IoU value higher than the predetermined threshold value TH2 among the multiple image regions R in the unregistered buffer B3 overlaps with the image region R in the registered buffer B4, and these two image regions R are likely to correspond to a single subject. Furthermore, of these two image regions R, the image region R in the unregistered buffer B3 has a lower score SC than the image region R in the registered buffer B4. Therefore, of these two image regions R, the NMS processing unit 24 deletes the image region R in the unregistered buffer B3.

[0069] Next, as shown in FIG. 11E, the NMS processing unit 24 extracts the image region R with the highest score SC from the plurality of image regions R in the unregistered buffer B3, and stores this image region R in the registered buffer B4.

[0070] Next, as shown in FIG. 11F, the NMS processing unit 24 performs IoU determination based on the image region R stored in the registered buffer B4 and each of the multiple image regions R in the unregistered buffer B3. In this example, the two IoU values ​​are lower than the threshold value TH2. Therefore, in this example, neither of the two image regions R in the unregistered buffer B3 is deleted.

[0071] The NMS processing unit 24 repeats this operation thereafter, until the unregistered buffer B3 becomes empty, as shown in FIG.

[0072] This completes the NMS reduction process B.

[0073] 12 shows an example of the NMS reduction process B. At the start of this NMS reduction process B, a plurality of image regions R processed by the pre-processing unit 23 are stored in the unregistered buffer B3.

[0074] First, the NMS processing unit 24 extracts the image region R with the highest score SC from the unregistered buffer B3, and stores this image region R in the registered buffer B4 (step S201).

[0075] Next, the NMS processing unit 24 calculates the IoU value between the image region R just entered in the registered buffer B4 and each of the one or more image regions R in the unregistered buffer B3 (step S202).

[0076] Then, the NMS processing unit 24 deletes the image region R having an IoU value higher than the threshold value TH2 from one or more image regions R in the unregistered buffer B3 (step S203).

[0077] Next, the NMS processing unit 24 checks whether the unregistered buffer B3 is empty (step S204). If the unregistered buffer B3 is not empty ("N" in step S204), the process returns to step S201, and the NMS processing unit 24 repeats steps S201 to S203 until the unregistered buffer B3 is empty. If the unregistered buffer B3 is empty ("Y" in step S204), this process ends.

[0078] In this way, the NMS processing unit 24 reduces the image region R.

[0079] Here, the IoU value in the NMS reduction process B corresponds to a specific example of a "second degree of overlap" in an embodiment of the present disclosure. The threshold value TH2 corresponds to a specific example of a "second threshold value" in an embodiment of the present disclosure.

[0080] The image area reduction unit 22 reduces the image area R by performing the above-described preliminary reduction process A and NMS reduction process B. This allows the object detection device 10 to set, for example, one image area R for one subject. The driving assistance processing unit 25 then detects the type and position of the subject based on the image area R processed by the image area reduction unit 22, and provides assistance to the driver in driving the vehicle 1 based on the detection result.

[0081] In the pre-reduction process A, the pre-processing unit 23 performs an IoU determination once each time an image region R is retrieved from the unregistered buffer B1, as shown in FIGS. 10 and 7H. Therefore, the computational complexity of the IoU determination is expressed as O(N), where N is the number of image regions R before the start of the pre-reduction process A. Here, the function O represents the computational complexity, and O(N) represents the computational complexity of N operations. Furthermore, during this IoU determination, the pre-processing unit 23 retrieves one image region R from the unregistered buffer B1 (step S102), retrieves one image region R from the temporary registration buffer B2 (steps S104 and S105), and stores the image region R in the temporary registration buffer B2 (step S108). The computational complexity of this process is expressed as O(log N) by using a priority queue. Here, log is a logarithmic function with base 2. Therefore, the computational complexity of the pre-reduction process A is expressed as O(N log N). Then, through this pre-reduction process A, the pre-processing unit 23 reduces the number of image regions R from N to N'.

[0082] In the NMS reduction process B, as shown in FIGS. 12 and 11C, the NMS processing unit 24 performs the IoU determination the same number of times as the number of image regions R in the unregistered buffer B3, which is one or more, every time it extracts an image region R from the unregistered buffer B3. Therefore, using the number "N'" of image regions R before the start of the NMS reduction process B, the amount of calculation of the NMS reduction process B is O(N' 2 )

[0083] In this way, in the object detection device 10, based on the N image regions R generated by the image region setting unit 21, the pre-processing unit 23 reduces the number of image regions R from N to N' in advance by pre-reduction processing A, and the NMS processing unit 24 further reduces the number of image regions R to N' by NMS reduction processing B, thereby reducing the amount of calculation.

[0084] That is, if the image area reduction unit 22 performs only the NMS reduction process B without performing the preliminary reduction process A, the amount of calculation for the NMS reduction process B is O(N 2 ) The amount of calculation is proportional to the square of the number of image regions R generated by the image region setting unit 21, so the greater the number of image regions R, the greater the amount of calculation.

[0085] On the other hand, in the object detection device 10, the pre-reduction process A is performed before the NMS reduction process B. As a result, in the object detection device 10, after reducing the image region R by the pre-reduction process A, the NMS reduction process B is performed using the smaller image region R. Therefore, the amount of calculation for the NMS reduction process B is O(N' 2 ), the amount of calculation required for the NMS reduction process B can be significantly reduced. Furthermore, the amount of calculation required for the pre-reduction process A is O(N log N), so even if the number of image regions R increases, the amount of calculation required for the pre-reduction process A does not increase significantly. As a result, the object detection device 10 can reduce the number of image regions R with a small amount of calculation.

[0086] In this way, the object detection device 10 is provided with a processing circuit (processing device 20) that can perform a first process of setting a plurality of image regions R corresponding to a plurality of subjects based on a captured image including images of the plurality of subjects and calculating the reliability (score SC) of each of the plurality of image regions R, and a second process of reducing the image regions R based on the degree of overlap (IoU value) among the plurality of image regions R. This second process includes a first reduction process (pre-reduction process A) and a second reduction process (NMS reduction process B). The first reduction process (pre-reduction process A) includes the steps of: sequentially selecting, as the first image region, the image region R set by the first process, starting with the image region R whose first position corresponding to a first region end (e.g., the left end) in a predetermined direction (e.g., the left direction) is closest to the image end (the left end of the captured image), which is one end of the predetermined direction in the captured image; selecting, as the second image region, the image region R whose second position corresponding to a second region end (e.g., the right end) in the predetermined direction is farthest from the image end (the left end of the captured image); calculating a first overlapping degree (IoU value) between the first image region and the second image region; and, if the first overlapping degree (IoU value) is equal to or greater than a first threshold (threshold TH1), designating, as a first deletion target, an image region with a low reliability (score SC) from the first image region and the second image region; and reducing the image region R based on the first deletion target. The second reduction process (NMS reduction process B) includes sequentially selecting, as a third image area, the multiple image areas reduced by the first reduction process (pre-reduction process A) in order from the image area R with the highest reliability (score SC); calculating a second degree of overlap (IoU value) between the third image area and one or more image areas R that have not yet been selected as the third image area; and deleting, from the one or more image areas R that have not yet been selected, one or more image areas R whose second degree of overlap (IoU value) is equal to or greater than a second threshold value (threshold value TH2), thereby reducing the image area R.As a result, in object detection device 10, after reducing image regions R by pre-reduction process A, NMS reduction process B is performed using fewer image regions R, thereby significantly reducing the amount of calculation in NMS reduction process B. Furthermore, the amount of calculation in this pre-reduction process A is O(N log N), so the amount of calculation does not increase significantly even if the number of image regions R increases. As a result, object detection device 10 can reduce image regions R with a small amount of calculation.

[0087] As described above, the present embodiment includes a processing circuit that can perform a first process of setting a plurality of image regions corresponding to a plurality of subjects based on a captured image including images of the plurality of subjects and calculating the reliability of each of the plurality of image regions, and a second process of reducing the image regions based on the degree of overlap among the plurality of image regions. This second process includes a first reduction process and a second reduction process. The first reduction process includes sequentially selecting as the first image area from the multiple image areas set by the first process, starting with the image area whose first position corresponding to the first area end in a predetermined direction is closest to the image end, which is one end of the predetermined direction in the captured image; selecting as the second image area, from one or more image areas previously selected as the first image area, the image area whose second position corresponding to the second area end in the predetermined direction is farthest from the image end; calculating a first degree of overlap between the first image area and the second image area; and, if the first degree of overlap is equal to or greater than a first threshold, designating the image area with a low reliability from the first image area and the second image area as a first deletion target; and reducing the image area based on the first deletion target. The second reduction process includes sequentially selecting, as a third image area, the image areas reduced by the first reduction process in descending order of reliability, calculating a second degree of overlap between each of the third image areas and one or more image areas that have not yet been selected as the third image area, and deleting, from among the one or more image areas that have not yet been selected, one or more image areas whose second degree of overlap is equal to or greater than a second threshold value, thereby reducing the image areas. This makes it possible to reduce the image areas with a small amount of calculation.

[0088] [Variation 1] In the above embodiment, as shown in Fig. 6, one prior reduction process A is performed before the NMS reduction process B, but this is not limited to this. Instead, for example, multiple prior reduction processes A may be performed. This variation will be described in detail below with several examples.

[0089] 13 shows an example of the operation of the processing device 20 according to this modification. The processes in steps S1, S2, S4, and S5 are the same as those in the above embodiment (FIG. 6).

[0090] After the process of step S2, the pre-processing unit 23 performs a pre-reduction process A1 (step S11).

[0091] 14 shows an example of the preliminary reduction process A1. The processes other than step S119 of this preliminary reduction process A1 are the same as those in the above embodiment ( FIG. 10 ). In step S119, the preliminary processing unit 23 stores the image region R with the higher score SC in the temporary registration buffer B2, and designates the image region R with the lower score SC as the image region to be deleted.

[0092] Next, as shown in FIG. 13, the pre-processing unit 23 performs a pre-reduction process A2 (step S12).

[0093] FIG. 15 shows an example of the advance reduction process A2. The processing of this advance reduction process A2 is the same as that of the advance reduction process A1 ( FIG. 15 ) except for steps S121, S122, and S125. This advance reduction process A2 is a process in which the processing order of the image regions R among the multiple image regions R in the advance reduction process A1 is changed. Specifically, in the advance reduction process A1, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the leftmost image region R. However, in the advance reduction process A2, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the rightmost image region R. In other words, the processing methods of the advance reduction process A1 and the advance reduction process A2 are reversed in the horizontal direction.

[0094] In the preliminary reduction process A2, the preliminary processing unit 23 first extracts the rightmost image region R from the unregistered buffer B1 and stores this image region R in the temporary registration buffer B2 (step S121). Next, the preliminary processing unit 23 extracts the rightmost image region R from the unregistered buffer B1 (step S122). Next, the preliminary processing unit 23 checks whether the temporary registration buffer B2 contains only one image region R (step S103). If there is only one image region R ("Y" in step S103), the preliminary processing unit 23 extracts that single image region R from the temporary registration buffer B2 (step S104). If there are two or more image regions R ("N" in step S103), the preliminary processing unit 23 extracts the leftmost image region R from the temporary registration buffer B2 (step S125). The subsequent processing is the same as in the preliminary reduction process A1 ( FIG. 15 ).

[0095] 13, the pre-processing unit 23 deletes the image region R that was selected as the deletion target in both the pre-reduction processes A1 and A2 (step S21). The subsequent processing is the same as in the above embodiment (FIG. 6).

[0096] As described above, in the pre-reduction process A1, the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R whose left edge is the leftmost, and in the pre-reduction process A2, the pre-reduction process 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R whose right edge is the rightmost. Thus, in the object detection device 10 according to the modified example, the processing order of image regions R in the pre-reduction process A1 and the processing order of image regions R in the pre-reduction process A2 are different from each other. The pre-processing unit 23 then deletes image regions R that are targeted for deletion in both the pre-reduction processes A1 and A2. As a result, the object detection device 10 according to this modified example can reduce the possibility of excessive deletion of image regions R, as will be described below.

[0097] Fig. 16 shows an example of three image regions R (image regions R31, R32, and R33) processed by object detection device 10 according to this modification. Fig. 17 shows the processing results of advance reduction processes A1 and A2 and NMS reduction process B.

[0098] 16, image regions R31 and R32 overlap each other, and the IoU value of these two image regions R31 and R32 is 0.6. Image regions R32 and R33 overlap each other, and the IoU value of these two image regions R32 and R33 is 0.8. Image regions R31 and R33 overlap each other, and the IoU value of these two image regions R31 and R33 is 0.2.

[0099] For example, when the three image regions R31 to R33 shown in FIG. 16 are supplied, if the image region reduction unit 22 performs only NMS reduction process B, image region R32 is deleted, leaving image regions R31 and R33, as shown in FIG. 17C. That is, first, the NMS processing unit 24 selects image region R33 and calculates the IoU values ​​of this image region R33 and the two image regions R31 and R32. Since the IoU value of image regions R31 and R33 is 0.2, which is below threshold value TH2 (0.5 in this example), the NMS processing unit 24 leaves image region R31. On the other hand, the IoU value of image regions R32 and R33 is 0.8, which is above threshold value TH2 (0.5), the NMS processing unit 24 deletes image region R32. As a result, image regions R31 and R33 remain.

[0100] For example, when the image area reduction unit 22 performs the pre-reduction process A1, as shown in FIG. 17A , image areas R31 and R32 are deleted, leaving image area R33. That is, first, the pre-processing unit 23 selects image areas R31 and R32 from left to right and performs IoU determination. Since the IoU values ​​of image areas R31 and R32 are 0.6, which is higher than threshold value TH1 (0.5 in this example), the pre-processing unit 23 selects image area R31, which has the lowest score SC, as the image area to be deleted. Next, the pre-processing unit 23 selects image areas R32 and R33 and performs IoU determination. Since the IoU values ​​of image areas R32 and R33 are 0.8, which is higher than threshold value TH1 (0.5), the pre-processing unit 23 selects image area R32, which has the lowest score SC, as the image area to be deleted. As a result, image area R33 remains. Therefore, even if the NMS processing unit 24 subsequently performs NMS reduction process B, the image area R31 will remain lost, resulting in a different result from when only NMS reduction process B is performed (Figure 17 (C)).

[0101] For example, when the image area reduction unit 22 performs the pre-reduction process A2, as shown in FIG. 17B , image area R32 is deleted, leaving image areas R31 and R33. That is, first, the pre-processing unit 23 selects image areas R32 and R33 from the right and performs IoU determination. Since the IoU values ​​of image areas R32 and R33 are 0.8, which is higher than threshold value TH1 (0.5 in this example), the pre-processing unit 23 selects image area R32, which has the lowest score SC, as the image area to be deleted. Next, the pre-processing unit 23 selects image areas R31 and R33 and performs IoU determination. Since the IoU values ​​of image areas R31 and R33 are 0.2, which is lower than threshold value TH1 (0.5), the pre-processing unit 23 leaves image areas R31 and R33. After this, when the NMS processing unit 24 performs NMS reduction process B, image areas R31 and R33 remain as they are. That is, the processing result is the same as the processing result when only the NMS reduction processing B is performed (FIG. 17C).

[0102] In this way, when the pre-processing unit 23 performs the pre-reduction process A1, the image region R31 that should not have been deleted ends up being deleted. In other words, an excessive amount of the image region R is deleted. On the other hand, when the pre-processing unit 23 performs the pre-reduction process A2, this image region R31 is not deleted. In this way, the pre-reduction processes A1 and A2 may produce different processing results depending on the order in which the image regions R are processed.

[0103] 13 , in the object detection device 10 according to this modification, the pre-processing unit 23 performs the pre-reduction processes A1 and A2 (steps S11 and S12) and deletes the image region R that was targeted for deletion in both the pre-reduction processes A1 and A2 (step S21). As a result, in the example of FIGS. 16 and 17 , only the image region R32 that was targeted for deletion in both the pre-reduction processes A1 and A2 is deleted, and the image region R31 that was targeted for deletion only in the pre-reduction process A1 is not deleted. As a result, in the object detection device 10 according to this modification, the effect of the processing order of the image regions R can be reduced, and the possibility of excessive deletion of image regions R can be reduced.

[0104] 18 shows an example of the operation of another processing device 20 according to this modification. The processes of steps S1, S2, S4, S5, S11, and S12 are the same as those in the modification (FIG. 13).

[0105] After the process of step S12, the pre-processing unit 23 performs a pre-reduction process A3 (step S13).

[0106] 19 shows an example of the advance reduction process A3. The processes of this advance reduction process A2 are the same as those of the advance reduction process A1 ( FIG. 15 ) except for steps S131, S132, and S135. In the advance reduction process A1, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R whose left edge is the leftmost. In the advance reduction process A3, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R whose left edge is the rightmost.

[0107] In the preliminary reduction process A3, the preliminary processing unit 23 first extracts the image region R whose left edge is the rightmost from the unregistered buffer B1 and stores this image region R in the temporary registration buffer B2 (step S131). Next, the preliminary processing unit 23 extracts the image region R whose left edge is the rightmost from the unregistered buffer B1 (step S132). Next, the preliminary processing unit 23 checks whether the temporary registration buffer B2 contains only one image region R (step S103). If there is only one image region R ("Y" in step S103), the preliminary processing unit 23 extracts that one image region R from the temporary registration buffer B2 (step S104). If there are two or more image regions R ("N" in step S103), the preliminary processing unit 23 extracts the image region R whose right edge is the leftmost from the temporary registration buffer B2 (step S135).

[0108] Fig. 20 shows an example of the process of step S135. In Fig. 20, the horizontal axis indicates the position of the image region R in the left-right direction of the captured image.

[0109] As shown in step S132, the pre-processing unit 23 extracts the image region R whose left edge is located at the rightmost position from the unregistered buffer B1. Therefore, as shown in Fig. 20, the left edge of the image region R extracted from the unregistered buffer B1 (image region R41) is located to the right of the left edges of all the image regions R in the unregistered buffer B1.

[0110] Furthermore, as shown in step S135, the pre-processing unit 23 extracts from the temporary registration buffer B2 the image region R whose right edge is located at the leftmost position. Therefore, as shown in FIG. 20 , the right edge of the image region R (image region R42) extracted from the temporary registration buffer B2 is located to the left of the right edges of all the image regions R in the temporary registration buffer B2. As a result, this image region R42 is more likely to overlap horizontally with the image region R41 extracted from the unregistered buffer B1 than the other image regions R in the temporary registration buffer B2. In FIG. 20 , the image regions R41 and R42 overlap horizontally. If the distance between the left edge of the image region R41 and the right edge of the image region R42 is distance L, the image region R42 is the image region R whose distance L is likely to be the shortest between the image region R41 and the image region R41. As shown in FIG. 4 , since the vehicle is on the road surface, the multiple image regions R in the captured image are aligned horizontally. Therefore, image region R41 and image region R42 are likely to have the narrowest union region RU within the two-dimensional image plane of the captured image. In this way, the pre-processing unit 23 selects image region R42, which is likely to have the narrowest union region RU between image region R41 and image region R42, using only the left-right direction, as shown in Figure 9.

[0111] The subsequent processing is the same as the advance reduction processing A1 (FIG. 15).

[0112] Next, the pre-processing unit 23 performs a pre-reduction process A4 (step S14).

[0113] 21 shows an example of the advance reduction process A4. The processes of this advance reduction process A2, other than steps S141, S142, and S145, are the same as those of the advance reduction process A1 ( FIG. 15 ). In the advance reduction process A3, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1, starting with the image region R whose left edge is the rightmost. However, in the advance reduction process A4, the advance processing unit 23 retrieves image regions R from the unregistered buffer B1, starting with the image region R whose right edge is the leftmost. In other words, the processing methods of the advance reduction process A3 and the advance reduction process A4 are reversed in the horizontal direction.

[0114] In the preliminary reduction process A4, the preliminary processing unit 23 first extracts the image region R whose right edge is the leftmost from the unregistered buffer B1 and stores this image region R in the temporary registration buffer B2 (step S141). Next, the preliminary processing unit 23 extracts the image region R whose right edge is the leftmost from the unregistered buffer B1 (step S142). Next, the preliminary processing unit 23 checks whether the temporary registration buffer B2 contains only one image region R (step S103). If there is only one image region R ("Y" in step S103), the preliminary processing unit 23 extracts that one image region R from the temporary registration buffer B2 (step S104). If there are two or more image regions R ("N" in step S103), the preliminary processing unit 23 extracts the image region R whose left edge is the rightmost from the temporary registration buffer B2 (step S145). The subsequent processing is the same as in the preliminary reduction process A1 (FIG. 15).

[0115] 18, the pre-processing unit 23 then deletes the image region R that was selected as the deletion target in all of the pre-reduction processes A1 to A4 (step S22). The subsequent processing is the same as in the above embodiment (FIG. 6).

[0116] In this way, in the advance reduction process A1, the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R with the leftmost left edge. In the advance reduction process A2, the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R with the rightmost right edge. In the advance reduction process A3, the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R with the rightmost left edge. In the advance reduction process A4, the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R with the leftmost right edge. Thus, in the object detection device 10 according to the modified example, the processing order of image regions R in the advance reduction process A1, the processing order of image regions R in the advance reduction process A2, the processing order of image regions R in the advance reduction process A3, and the processing order of image regions R in the advance reduction process A4 are all different. The pre-processing unit 23 then deletes image regions R that were targeted for deletion in all of the advance reduction processes A1 to A4. As a result, the object detection device 10 according to this modified example can reduce the possibility of excessive deletion of image regions R, as described below.

[0117] Fig. 22 shows an example of three image regions R (image regions R51, R52, and R53) processed by object detection device 10 according to this modification. Fig. 23 shows the processing results of pre-reduction processes A1 to A4 and NMS reduction process B.

[0118] 22 , image region R52 is included in image region R51, and image region R53 is included in image region R52. Image region R51 and image region R52 overlap each other, and the IoU value of these two image regions R51 and R52 is 0.8. Image region R52 and image region R53 overlap each other, and the IoU value of these two image regions R52 and R53 is 0.6. Image region R51 and image region R53 overlap each other, and the IoU value of these two image regions R51 and R53 is 0.3.

[0119] For example, when the three image areas R51 to R53 shown in Figure 22 are supplied, if the image area reduction unit 22 performs only NMS reduction process B, image area R52 is deleted and image areas R51 and R53 remain, as shown in Figure 23 (E).

[0120] For example, when the image area reduction unit 22 performs the preliminary reduction process A1, image areas R51 and R52 are deleted, leaving image area R53, as shown in Fig. 23(A). Similarly, when the image area reduction unit 22 performs the preliminary reduction process A2, image areas R51 and R52 are deleted, leaving image area R53, as shown in Fig. 23(B). Therefore, even if the NMS processing unit 24 subsequently performs the NMS reduction process B, image area R51 remains lost, resulting in a different result from when only the NMS reduction process B is performed (Fig. 23(E)).

[0121] For example, when the image area reduction unit 22 performs the preliminary reduction process A3, image area R52 is deleted, leaving image areas R51 and R53, as shown in Fig. 23(C). Similarly, when the image area reduction unit 22 performs the preliminary reduction process A4, image area R52 is deleted, leaving image areas R51 and R53, as shown in Fig. 23(D). If the NMS processing unit 24 then performs the NMS reduction process B, image areas R51 and R53 remain. In other words, the result of this process is the same as the result of performing only the NMS reduction process B (Fig. 23(E)).

[0122] In this way, when the pre-processing unit 23 performs the pre-reduction process A1 or A2, image region R51 is deleted, even though it should not have been deleted. In other words, an excessive amount of image region R is deleted. On the other hand, when the pre-processing unit 23 performs the pre-reduction process A3 or A4, image region R51 is not deleted. In this way, the pre-reduction processes A1 and A2 and the pre-reduction processes A3 and A4 may produce different processing results due to the order in which image region R is processed.

[0123] 18 , in the object detection device 10 according to this modification, the pre-processing unit 23 performs the pre-reduction processes A1 to A4 (steps S11 and S14) and deletes the image region R that was targeted for deletion in all of the pre-reduction processes A1 to A4 (step S22). As a result, in the example of FIGS. 22 and 23 , only the image region R52 that was targeted for deletion in all of the pre-reduction processes A1 to A4 is deleted, and the image region R51 that was targeted for deletion only in the pre-reduction processes A1 and A2 is not deleted. As a result, in the object detection device 10 according to this modification, the effect of the processing order of the image region R can be reduced, and the possibility of excessive deletion of the image region R can be reduced.

[0124] Furthermore, in the example of Figure 13, the pre-processing unit 23 is configured to perform pre-reduction processing A1 and pre-reduction processing A2, but this is not limited to this, and instead, it may be configured to perform any two or more of the pre-reduction processing A1 to A4.

[0125] [Variation 2] In the above embodiment, as shown in, for example, FIGS. 6 and 10 , the pre-processing unit 23 retrieves image regions R from the unregistered buffer B1 in order, starting with the image region R with the leftmost left edge, in the pre-reduction process A. However, this is not limited to this. Alternatively, the pre-processing unit 23 may retrieve image regions R from the unregistered buffer B1 in order, starting with the image region R with the rightmost right edge, in the pre-reduction process A, similar to the pre-reduction process A2 ( FIG. 15 ). Furthermore, the pre-processing unit 23 may retrieve image regions R from the unregistered buffer B1 in order, starting with the image region R with the rightmost left edge, in the pre-reduction process A, similar to the pre-reduction process A3 ( FIG. 19 ). Furthermore, the pre-processing unit 23 may retrieve image regions R from the unregistered buffer B1 in order, starting with the image region R with the leftmost right edge, in the pre-reduction process A, similar to the pre-reduction process A3 ( FIG. 21 ).

[0126] [Variation 3] In the above embodiment, for example, the image region R whose left edge is the leftmost is extracted from the unregistered buffer B1 as shown in steps S101 and S102 of Fig. 10, but this is not limited to this. Instead, for example, the left edge of the image region R may be set as "0%" and the right edge of the image region R as "100%," and the image region R whose quantile point is the leftmost of a predetermined percentage (for example, 5%) may be extracted from the unregistered buffer B1.

[0127] Although several embodiments of the present disclosure have been described above by way of example with reference to the accompanying drawings, the present disclosure is by no means limited to the above-described embodiments. Those skilled in the art will understand that various modifications and variations can be made without departing from the scope defined by the appended claims. The present disclosure is intended to encompass such modifications and variations to the extent that they fall within the scope of the appended claims and their equivalents.

[0128] For example, in the above embodiment, the object detection device 10 is provided in the vehicle 1, but the present invention is not limited to this, and the object detection device 10 can be provided in various devices. In other words, the object detection device 10 can be used for various purposes such as detecting an object based on a captured image.

[0129] The effects described in this specification are merely examples, and the effects of the present disclosure are not limited to the effects described in this specification. Therefore, other effects may be obtained with respect to the present disclosure.

[0130] Furthermore, the present disclosure may take the following aspects.

[0131] (1) A processing circuit is provided that is capable of performing a first process of setting a plurality of image regions corresponding to a plurality of subjects based on a captured image including images of the plurality of subjects and calculating a reliability of each of the plurality of image regions, and a second process of reducing the image regions based on a degree of overlap in the plurality of image regions, wherein the second process includes a first reduction process and a second reduction process, and the first reduction process comprises: sequentially selecting, as a first image region, the plurality of image regions set by the first process in order from an image region whose first position corresponding to a first region end in a predetermined direction is closest to an image end that is one end of the predetermined direction in the captured image; selecting, as a second image region, an image region whose second position corresponding to a second region end in the predetermined direction is farthest from the image end; and calculating a first degree of overlap between the first image region and the second image region. an object detection device including: when the first degree of overlap is equal to or greater than a first threshold, designating an image area with a low reliability from the first image area and the second image area as a first deletion target; and reducing image areas based on the first deletion target; wherein the second reduction process includes: sequentially selecting, as a third image area, the plurality of image areas reduced by the first reduction process in order from the image area with the highest reliability; and calculating a second degree of overlap between the third image area and one or more image areas not yet selected as the third image area, and reducing the image areas by deleting, from the one or more image areas not yet selected, one or more image areas with the second degree of overlap that is equal to or greater than a second threshold. (2) The object detection device according to (1), wherein the plurality of subjects include a vehicle, and the predetermined direction is a left-right direction in the captured image.(3) The object detection device described in (1) or (2), wherein the first reduction process includes: sequentially selecting, as a fourth image area, from the plurality of image areas set by the first process, an image area whose second position corresponding to an end of the second area in the predetermined direction is farthest from the end of the image in the captured image; selecting, as a fifth image area, an image area whose first position corresponding to an end of the first area is closest to the end of the image, from one or more image areas previously selected as the fourth image area; calculating a second overlapping degree between the fourth image area and the fifth image area; and, when the second overlapping degree is equal to or greater than the first threshold, designating an image area with a low reliability from the fourth image area and the fifth image area as a second deletion target; and reducing image areas based on the first deletion targets and the second deletion targets. (4) The object detection device described in any of (1) to (3), wherein the first reduction process includes: sequentially selecting, as a sixth image area, from the plurality of image areas set by the first process, an image area in order from an image area whose first position corresponding to an end of the first area in the predetermined direction is farthest from the end of the image in the captured image; selecting, as a seventh image area, an image area whose second position corresponding to an end of the second area is closest to the end of the image, from one or more image areas previously selected as the sixth image area; calculating a third overlapping degree between the sixth image area and the seventh image area; and, when the third overlapping degree is equal to or greater than the first threshold, designating an image area of ​​the sixth image area and the seventh image area with a low reliability as a third deletion target; and reducing image areas based on the first deletion targets and the third deletion targets.(5) The object detection device according to any one of (1) to (4), wherein the first reduction process includes: sequentially selecting, as an eighth image area, from the plurality of image areas set by the first process in order from an image area whose second position corresponding to an end of the second area in the predetermined direction is closest to the end of the image in the captured image; selecting, as a ninth image area, an image area whose first position corresponding to an end of the first area in the predetermined direction is farthest from the end of the image, from one or more image areas previously selected as the eighth image area; calculating a fourth overlapping degree between the eighth image area and the ninth image area; and, when the fourth overlapping degree is equal to or greater than the first threshold, designating, as a fourth deletion target, an image area with a low reliability from the eighth image area and the ninth image area; and reducing image areas based on the first deletion targets and the fourth deletion targets. (6) The object detection device according to any one of (1) to (5), wherein the first process is a process using a neural network.

[0132] The processing device 20 shown in FIG. 2 can be implemented by circuitry including at least one semiconductor integrated circuit, such as at least one processor (e.g., a central processing unit (CPU)), at least one application-specific integrated circuit (ASIC), and / or at least one field-programmable gate array (FPGA). The at least one processor can be configured to perform all or a portion of the various functions of the processing device 20 shown in FIG. 2 by reading instructions from at least one non-transitory, tangible computer-readable medium. Such media can take various forms, including, but not limited to, various magnetic media such as hard disks, various optical media such as CDs or DVDs, and various semiconductor memories (i.e., semiconductor circuits) such as volatile or non-volatile memories. Volatile memories can include DRAM and SRAM. Non-volatile memories can include ROM and NVRAM. An ASIC is an integrated circuit (IC) specialized to perform all or a portion of the various functions of the processing device 20 shown in FIG. 2. An FPGA is an integrated circuit designed to be configurable after manufacture to perform all or a portion of the various functions of the processing device 20 shown in FIG. 2.

Claims

1. A processing circuit is provided that is capable of performing a first process of setting a plurality of image regions corresponding to a plurality of subjects based on a captured image including images of the plurality of subjects and calculating the reliability of each of the plurality of image regions, and a second process of reducing the image regions based on the degree of overlap in the plurality of image regions, wherein the second process includes a first reduction process and a second reduction process, and the first reduction process comprises: sequentially selecting, as a first image region, the plurality of image regions set by the first process, starting with the image region whose first position corresponding to a first region end in a predetermined direction is closest to an image end that is one end of the predetermined direction in the captured image; selecting, as a second image region, an image region whose second position corresponding to a second region end in the predetermined direction is farthest from the image end; and calculating a first degree of overlap between the first image region and the second image region. an object detection device including: when the first degree of overlap is equal to or greater than a first threshold, designating an image area with a low reliability from the first image area and the second image area as a first deletion target; and reducing image areas based on the first deletion target; wherein the second reduction process includes: sequentially selecting, as a third image area, the plurality of image areas reduced by the first reduction process in order from the image area with the highest reliability; and calculating a second degree of overlap between the third image area and one or more image areas that have not yet been selected as the third image area, and reducing the image areas by deleting, from the one or more image areas that have not yet been selected, one or more image areas whose second degree of overlap is equal to or greater than a second threshold.

2. The object detection device according to claim 1, wherein the plurality of subjects include a vehicle, and the predetermined direction is the left-right direction in the captured image.

3. The object detection device of claim 1, wherein the first reduction process includes: sequentially selecting as a fourth image area from the plurality of image areas set by the first process, starting with the image area whose second position corresponding to the end of the second area in the predetermined direction is farthest from the end of the image in the captured image; selecting as a fifth image area, from the one or more image areas previously selected as the fourth image area, the image area whose first position corresponding to the end of the first area is closest to the end of the image; calculating a second degree of overlap between the fourth image area and the fifth image area; and, if the second degree of overlap is equal to or greater than the first threshold, designating the image area with the low reliability from the fourth image area and the fifth image area as a second target for reduction; and reducing image areas based on the first target for reduction and the second target for reduction.

4. The object detection device of claim 1, wherein the first reduction process includes: sequentially selecting as a sixth image area from the plurality of image areas set by the first process, starting with the image area whose first position corresponding to the end of the first area in the predetermined direction is farthest from the end of the image in the captured image; selecting as a seventh image area, from one or more image areas previously selected as the sixth image area, the image area whose second position corresponding to the end of the second area is closest to the end of the image; calculating a third degree of overlap between the sixth image area and the seventh image area; if the third degree of overlap is equal to or greater than the first threshold, setting the image area with the low reliability from the sixth image area and the seventh image area as a third target for reduction; and reducing image areas based on the first target for reduction and the third target for reduction.

5. The object detection device of claim 1, wherein the first reduction process includes: sequentially selecting as an eighth image area from the plurality of image areas set by the first process, starting with an image area in which the second position corresponding to the end of the second area in the predetermined direction is closest to the end of the image in the captured image; selecting as a ninth image area, from one or more image areas previously selected as the eighth image area, an image area in which the first position corresponding to the end of the first area is farthest from the end of the image; calculating a fourth degree of overlap between the eighth image area and the ninth image area; and, if the fourth degree of overlap is equal to or greater than the first threshold, designating an image area with a low reliability from the eighth image area and the ninth image area as a fourth target for reduction; and reducing image areas based on the first target for reduction and the fourth target for reduction.

6. The object detection device according to claim 1, wherein the first processing is processing using a neural network.

Citation Information

Patent Citations

  • Object detection method, object detection device, and image processing apparatus

    JP2020205039A

  • Object detection device, object detection method, and program

    WO2021005898A1