Detection device, control method for detection device, model generation method by model generation device for generating learned model, information processing program, and recording medium

The detection device enhances the speed and accuracy of person detection in fish-eye lens images by dividing the image into regions and using regional probabilities, eliminating the need for multiple dictionaries, thus improving detection efficiency.

JP7703888B2Active Publication Date: 2025-07-08OMRON CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021074210
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-04-26
Publication Date
2025-07-08
Estimated Expiration
2041-04-26

AI Technical Summary

Technical Problem

Conventional methods for detecting a person using a fish-eye lens require longer processing times due to the use of multiple dictionaries indicating human characteristics, leading to reduced detection speed and accuracy.

Method used

A detection device that divides the captured image into multiple regions, calculates the probability of a person's presence in each region, and determines whether the detection target is a person using these probabilities, eliminating the need for multiple dictionaries and improving detection speed and accuracy.

Benefits of technology

The device enables rapid and accurate detection of persons from images captured with a fish-eye lens by verifying the estimated results using regional probabilities, reducing the time and labor required for dictionary preparation and memory storage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007703888000001
    Figure 0007703888000001
  • Figure 0007703888000002
    Figure 0007703888000002
  • Figure 0007703888000003
    Figure 0007703888000003
Patent Text Reader

Abstract

To provide a detection device that detects a person at high speed and with high accuracy from an image captured using a fisheye lens, a detection device control method, a model generation method using a model generation device for generating a learned model, an information processing program, and a recording medium.SOLUTION: A detection device 10 includes: a division unit 120 configured to divide an image captured by a ceiling camera using a fisheye lens into a plurality of areas; an area estimation unit 152 configured to calculate a probability that each of the plurality of areas includes a position where a detection object estimated to be a person exists; and a determination unit 190 configured to determine whether the detection object is a person based on the probability of each of the plurality of areas.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a detection device that detects a person being imaged from an imaging image captured by a ceiling camera using a fish-eye lens.

Background Art

[0002] Conventionally, various studies have been known for detecting a person by analyzing an image captured using a fish-eye lens. For example, Patent Document 1 listed below discloses an image sensor that accurately detects a person from an image in which a wide range is imaged by using a plurality of dictionary information.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, since the above-described conventional technology detects a person from an image using a plurality of dictionaries indicating human characteristics, there is a problem that the time required to detect a person becomes longer compared to the case where there is one dictionary indicating human characteristics.

[0005] One aspect of the present invention aims to detect a person from an image captured using a fish-eye lens at high speed and with high accuracy.

Means for Solving the Problems

[0006] In order to solve the above problems, a detection device according to an aspect of the present invention is a detection device that detects a person captured in a captured image captured by a ceiling camera using a fisheye lens, the detection device including: a dividing unit that divides the captured image into a plurality of regions; a region estimation unit that calculates, for each of the plurality of regions, a probability including a position where a detection target estimated to be a person exists; and a determination unit that determines whether the detection target is a person using the probability of each of the plurality of regions.

[0007] According to the above configuration, the detection device divides the captured image into a plurality of regions, and calculates, for each of the plurality of regions, a probability including the "position where the detection target estimated to be a person exists". Then, the detection device verifies the estimation result that the detection target is a person using the probability of each of the plurality of regions.

[0008] Here, generally, when trying to detect a person from an image using a dictionary indicating human features, analysis of the image is required for each dictionary. Therefore, when trying to improve the detection accuracy of a person using a plurality of dictionaries indicating human features, analysis of the image is also required multiple times, and the time required to detect a person becomes long.

[0009] In contrast, the detection device determines whether the detection target is a person using the probability calculated for each of the plurality of regions obtained by dividing the captured image, thereby improving the detection accuracy of a person. That is, the detection device does not detect a person from the captured image using a plurality of dictionaries indicating human features, but improves the detection accuracy of a person by determining whether the detection target estimated to be a person is a person using the probability of each of the plurality of regions.

[0010] While the method of using a plurality of dictionaries indicating human features improves the accuracy of the estimation itself, the detection device improves the detection accuracy by verifying the estimated result (the estimation that the detection target is a person) (that is, removing an incorrect estimation result).

[0011] Therefore, the detection device does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image, and can shorten the time required to detect a person from the captured image compared to the case of detecting a person from the captured image using a plurality of dictionaries.

[0012] In addition, since the detection device determines (verifies) whether the detection target estimated to be a person is actually a person by using the probabilities of each of the plurality of regions, a person can be detected from the captured image with high accuracy.

[0013] Therefore, the detection device has the effect of being able to detect a person from an image (captured image) captured using a fish-eye lens quickly and with high accuracy.

[0014] Also, as described above, the detection device does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image.

[0015] Therefore, the detection device has the effect of reducing the labor required to prepare a dictionary (for example, a learned model) necessary for detecting a person from an image captured using a fish-eye lens, and also reducing the memory capacity required to store the dictionary.

[0016] The detection device according to one aspect of the present invention may further include a specifying unit that specifies, as a specific region, a region among the plurality of regions in which a bounding box surrounding the detection target exists, and the determination unit may determine whether the detection target is a person by using the probabilities of each of the plurality of regions and the specific region.

[0017] According to the above configuration, the detection device specifies the specific region in which the bounding box exists among the plurality of regions, and determines whether the detection target is a person by using the probabilities of each of the plurality of regions and the specific region.

[0018] For example, when the detection target estimated to be a person is actually a person, it is considered that the consistency between the probability of each of the plurality of regions and the region (specific region) where the bounding box surrounding the detection target exists is also high.

[0019] Therefore, the detection device determines whether the detection target is a person by using the probability of each of the plurality of regions and the specific region, that is, verifies the estimation that the detection target is a person.

[0020] Therefore, the detection device can effectively detect a person with high accuracy from an image (captured image) captured using a fisheye lens.

[0021] The specific part may specify, as the specific region, the position surrounded by the bounding box and estimated to be "a person" of the detection target (for example, the position corresponding to the feet of the detection target estimated to be "a person"), or the region including the position of the detection target. The specific part may calculate the position (foot position) of the detection target from the bounding box, for example, from the position, shape, and size of the bounding box. Further, the specific part may specify the center position (or centroid position) of the bounding box as 'the position (foot position) of the detection target surrounded by the bounding box and estimated to be "a person"'.

[0022] In the detection device according to an aspect of the present invention, the determination unit may determine that the detection target is a person when the specific region matches the region having the highest probability among the plurality of regions.

[0023] According to the above configuration, when the specific region matches the region with the highest probability among the plurality of regions, the detection device determines that the detection target is a person. That is, the detection device determines that the detection target is a person when the "region (the specific region) where the bounding box surrounding the detection target estimated to be a person exists" matches the "region with the highest probability including the position where the detection target estimated to be a person exists".

[0024] When the detection target estimated to be a person is actually a person, it is considered highly likely that the specific region where the bounding box surrounding the detection target exists matches the "region with the highest probability including the position where the detection target exists".

[0025] Therefore, when the specific region matches the region with the highest probability among the plurality of regions, the detection device determines that the detection target is a person.

[0026] Therefore, the detection device can effectively detect a person with high accuracy from an image (captured image) captured using a fish-eye lens.

[0027] When the specific part specifies, as the specific region, the position (foot position) of the detection target surrounded by the bounding box and estimated to be "a person", the determination unit may determine whether the foot position is included in the region with the highest probability. The determination unit may determine that the detection target is a person when the foot position is included in the region with the highest probability.

[0028] In the detection device according to an aspect of the present invention, the determination unit may determine that the detection target is a person when the specific region matches the region with the highest probability among the plurality of regions or a region adjacent to the region with the highest probability among the plurality of regions.

[0029] According to the above configuration, when the specific area coincides with the area having the highest probability among the plurality of areas or an area adjacent to the area having the highest probability among the plurality of areas, the detection device determines that the detection target is a person.

[0030] Here, in the captured image, a situation is assumed where a person is captured so as to straddle two of the plurality of areas, or the position where the person is captured is near the boundary between the two areas.

[0031] Under such a situation, when the detection target estimated to be a person is actually a person, the specific area where the bounding box surrounding the detection target exists is considered likely to coincide with any of the following areas. That is, it is considered likely that the specific area coincides with the area having the highest probability among the plurality of areas or an area adjacent to the area having the highest probability among the plurality of areas.

[0032] Therefore, when the specific area coincides with the area having the highest probability among the plurality of areas, the detection device determines that the detection target is a person. Also, when the specific area coincides with an area adjacent to the area having the highest probability among the plurality of areas, the detection device determines that the detection target is a person.

[0033] Therefore, even when a person is captured near the boundary between two of the plurality of areas in the captured image, the detection device can effectively detect a person with high accuracy from the captured image.

[0034] When the specific part specifies the foot position as the specific area, the determination part may determine whether the foot position is included in the area having the highest probability or an area adjacent to the area having the highest probability. When the foot position is included in the area having the highest probability or an area adjacent to the area having the highest probability, the determination part may determine that the detection target is a person.

[0035] In the detection device according to one aspect of the present invention, when the class confidence value calculated by multiplying the probability of "the region corresponding to the specific region among the plurality of regions" by the object confidence value, which is a value indicating the likelihood that some object is surrounded by the bounding box, is greater than a predetermined value, the determination unit may determine that the detection target is a person.

[0036] According to the above configuration, when the class confidence value calculated by multiplying the probability of the region corresponding to the specific region among the plurality of regions by the object confidence value is greater than a predetermined value, the detection device determines that the detection target is a person.

[0037] When the detection target estimated to be a person is actually a person, it is considered that the probability that the specific region where the bounding box surrounding the detection target exists includes the position where the detection target exists is sufficiently high. Also, when the detection target estimated to be a person is actually a person, the object confidence value, which is a value indicating the likelihood that some object is surrounded by the bounding box surrounding the detection target, is also considered to be a sufficiently high value.

[0038] Therefore, when the class confidence value calculated by multiplying the probability of "the region corresponding to the specific region among the plurality of regions" by the object confidence value is greater than a predetermined value, the detection device determines that the detection target is a person.

[0039] Therefore, the detection device has the effect of being able to detect a person with high accuracy from an image (captured image) captured using a fisheye lens.

[0040] When the specific part specifies the foot position as the specific region, when the class confidence value calculated by multiplying the probability of the region including the foot position by the object confidence value is greater than a predetermined value, the determination unit may determine that the detection target is a person.

[0041] In the detection device according to one aspect of the present invention, when the class confidence value calculated by multiplying the average value of the probability of "the region corresponding to the specific region" among the plurality of regions and the probability of "the region adjacent to the region corresponding to the specific region" among the plurality of regions by the object confidence value, which is a value indicating the likelihood that an object is surrounded by the bounding box, is greater than a predetermined value, it may be determined that the detection target is a person.

[0042] According to the above configuration, the detection device calculates the average value of the probability of "the region corresponding to the specific region" among the plurality of regions and the probability of "the region adjacent to the region corresponding to the specific region" among the plurality of regions. Then, when the class confidence value calculated by multiplying the average value by the object confidence value, which is a value indicating the likelihood that an object is surrounded by the bounding box, is greater than a predetermined value, the detection device determines that the detection target is a person.

[0043] Here, in the captured image, a situation is assumed where a person is captured so as to straddle two regions among the plurality of regions, or the position where the person is captured is near the boundary of the two regions.

[0044] Under such a situation, when the detection target estimated to be a person is actually a person, it is considered that both the probability of "the region corresponding to the specific region" and the probability of "the region adjacent to the region corresponding to the specific region" are sufficiently high values. Similarly, when the detection target estimated to be a person is actually a person, the object confidence value, which is a value indicating the likelihood that an object is surrounded by the bounding box surrounding the detection target, is also considered to be a sufficiently high value.

[0045] Therefore, the detection device calculates the average value of the probability of "the region corresponding to the specific region" and the probability of "the region adjacent to the region corresponding to the specific region". Then, when the class confidence value calculated by multiplying the average value by the object confidence level is greater than a predetermined value, the detection device determines that the detection target is a person.

[0046] Therefore, even when a person is imaged near the boundary of two of the plurality of regions in the captured image, the detection device can effectively detect a person with high accuracy from the captured image.

[0047] When the specific part specifies the foot position as the specific region, the determination unit may calculate the average value of the probability of "the region including the foot position" and the probability of "the region adjacent to the region including the foot position". Then, when the class confidence value calculated by multiplying the calculated average value by the object confidence level is greater than a predetermined value, the determination unit may determine that the detection target is a person.

[0048] In the detection device according to one aspect of the present invention, the region estimation unit may calculate the probability of each of the plurality of regions from the captured image using a region prediction model, which is a learned model that takes the captured image as an input and outputs the probability of the presence position of the detection target for each of the plurality of regions.

[0049] According to the above configuration, the detection device calculates the probability of each of the plurality of regions from the captured image using a region prediction model, which is a learned model that takes the captured image as an input and outputs the probability of the presence position of the detection target for each of the plurality of regions.

[0050] Therefore, the detection device can effectively calculate the probability of each of the plurality of regions from the captured image with high accuracy by using the region prediction model.

[0051] The detection device according to one aspect of the present invention further includes a learning unit that constructs a region prediction model, which takes the captured image as an input and outputs, for each of the plurality of regions, a probability including "the position where the detection target exists", by machine learning with respect to teacher data to which information indicating a region including "the position where the person captured in the captured image exists" is attached as a label.

[0052] According to the above configuration, the detection device constructs the region prediction model by machine learning with respect to teacher data to which information indicating a region including "the position where a person exists" is attached as a label for the captured image.

[0053] Therefore, the detection device can construct, by machine learning with respect to the teacher data, the region prediction model that enables the probability of each of the plurality of regions to be calculated with high accuracy from the captured image.

[0054] In order to solve the above problems, a control method according to one aspect of the present invention is a control method for a detection device that detects a person captured in a captured image captured by a ceiling camera using a fisheye lens, the control method including: a dividing step of dividing the captured image into a plurality of regions; a region estimating step of calculating, for each of the plurality of regions, a probability including the position where a detection target estimated to be a person exists; and a determining step of determining whether the detection target is a person using the probabilities of each of the plurality of regions.

[0055] According to the above configuration, the control method divides the captured image into a plurality of regions and calculates, for each of the plurality of regions, a probability including "the position where the detection target estimated to be a person exists". Then, the control method verifies the estimation result that the detection target is a person using the probabilities of each of the plurality of regions.

[0056] Here, generally, when attempting to detect a person from an image using a dictionary indicating human characteristics, analysis of the image is required for each dictionary. Therefore, when attempting to improve the detection accuracy of a person using a plurality of dictionaries indicating human characteristics, analysis of the image is also required multiple times, and the time required to detect a person becomes long.

[0057] In contrast, the control method determines whether the detection target is a person using the probability including "the position where the detection target estimated to be a person exists" calculated for each of the plurality of regions obtained by dividing the captured image, thereby improving the detection accuracy of a person. That is, the control method does not detect a person from the captured image using a plurality of dictionaries indicating human characteristics, but improves the detection accuracy of a person by determining whether the detection target estimated to be a person is a person using the probabilities of each of the plurality of regions.

[0058] While the method using a plurality of dictionaries indicating human characteristics improves the accuracy of the estimation itself, the control method improves the detection accuracy by verifying the estimated result (the estimation that the detection target is a person) (that is, removing an incorrect estimation result).

[0059] Therefore, the control method does not need to use a plurality of dictionaries indicating human characteristics to detect a person from the captured image, and can shorten the time required to detect a person from the captured image compared to the case of detecting a person from the captured image using a plurality of dictionaries.

[0060] In addition, since the control method determines (verifies) whether the detection target estimated to be a person is a person using the probabilities of each of the plurality of regions, a person can be detected from the captured image with high accuracy.

[0061] Therefore, the control method has the effect of being able to detect a person from an image (captured image) captured using a fisheye lens quickly and with high accuracy.

[0062] Also, as described above, the control method does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image.

[0063] Therefore, the control method has the effect of reducing the labor required to prepare a dictionary (for example, a learned model) necessary for detecting a person from an image captured using a fish-eye lens, and also reducing the memory capacity required to store the dictionary.

[0064] In order to solve the above problems, a model generation method according to an aspect of the present invention is a model generation method by a model generation device that generates a learned model. For a captured image divided into a plurality of regions, (A) information indicating the position, shape, and size of a bounding box that surrounds a person captured in the captured image, and (B) information for identifying an area including 'the position where the person captured in the captured image exists' are obtained as teacher data with labels, and a learning step of constructing a learned model that takes the captured image as an input and outputs (C) information indicating the position, shape, and size of the bounding box and (D) the probability of each of the plurality of regions including 'the position where the person captured in the captured image exists' by machine learning on the teacher data.

[0065] According to the above configuration, the model generation method constructs the learned model by machine learning on the teacher data. When the learned model is input with the captured image, it outputs the following two pieces of information. That is, (C) information indicating the position, shape, and size of the bounding box (rectangle information), and (D) the probability of each of the plurality of regions including 'the position where the person captured in the captured image exists'.

[0066] The rectangular information is information indicating the position, shape, and size of the bounding box surrounding the detection target estimated to be "a person", and is information including the estimation that the detection target surrounded by the bounding box is "a person".

[0067] Therefore, when the imaging image is input, the model generation method can construct a learned model that outputs the rectangular information including the estimation that the detection target is "a person" and the probability of each of the plurality of regions, which has the effect of achieving this.

Effect of the Invention

[0068] According to one aspect of the present invention, there is an effect that a person can be detected from an image captured using a fish-eye lens at high speed and with high accuracy.

Brief Description of the Drawings

[0069]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Mode for Carrying Out the Invention

[0070] 〔Embodiment 1〕 Hereinafter, an embodiment according to one aspect of the present invention (hereinafter also referred to as "this embodiment") will be described with reference to FIGS. 1 to 19. In the drawings, the same or corresponding parts are denoted by the same reference numerals and their description will not be repeated. In this embodiment, for example, the detection device 10 will be described as a typical example of a detection device that detects a person imaged from a captured image PI captured by a ceiling camera 20 using a fish-eye lens.

[0071] In the following description, "m", "n", "p", "q", "x", and "y" each represent an integer of "1" or more. Also, "m" is an integer less than or equal to "n", "p" and "q" are different integers from each other, and "x" and "y" are different integers from each other.

[0072] When it is necessary to distinguish a plurality of regions AR from each other, subscripts such as "(1)", "(2)", "(3)",..., "(n)" are attached to the reference numerals for distinction. For example, they are described as "region AR(1)", "region AR(2)", "region AR(3)",..., "region AR(n)" for distinction. When it is not necessary to particularly distinguish each of the plurality of regions AR, they are simply referred to as "region AR". The same applies to the detection rectangle BB, the detection target OB, the position PS, etc.

[0073] To facilitate the understanding of the detection device 10 according to one aspect of the present invention, first, the outline of the person detection system 1 including the detection device 10 will be described with reference to FIG. 2.

[0074] §1. Application Example (Overall Outline of Person Detection System) FIG. 2 is a diagram showing the overall outline of the person detection system 1 including the detection device 10. As shown in FIG. 2, the person detection system 1 includes a ceiling camera 20 that generates a captured image PI, and a detection device 10 that executes image analysis on the captured image PI generated by the ceiling camera 20 and detects a person imaged in the captured image PI. The ceiling camera 20 and the detection device 10 are communicably connected to each other via a communication cable, which is, for example, a USB (Universal Serial Bus) cable.

[0075] The ceiling camera 20 is an imaging device that can image a wide range of imaging target spaces using a fish-eye lens (ultra-wide-angle lens). The ceiling camera 20 is installed, for example, on the ceiling of the factory Fa, and generates an imaging image PI that images a plurality of work areas Ar in the factory Fa from above (diagonally above). The ceiling camera 20 outputs the generated imaging image PI to the detection device 10.

[0076] The detection device 10 acquires the imaging image PI generated by the ceiling camera 20 from the ceiling camera 20. The detection device 10 performs image analysis on the acquired imaging image PI and detects a person imaged in the imaging image PI.

[0077] (Regarding the need to detect people with high precision) In order to improve the work process performed in the factory Fa, it is required to accurately detect workers (people) existing in the factory Fa from the imaging image PI obtained by imaging the factory Fa. However, it is not easy to accurately detect a person (human body) imaged in the imaging image PI by image analysis of the imaging image PI captured by a fish-eye camera such as the ceiling camera 20.

[0078] (Occurrence of false detection) FIG. 3 is a diagram showing an example in which false detection occurs when trying to detect a person from the imaging image PI captured by the ceiling camera 20. In the image analysis of the imaging image PI using a learned model, for example, a detection rectangle BB (also referred to as a "bounding box") is arranged around the detection target OB estimated to be a "person (human body)" so as to surround the detection target OB.

[0079] In the imaging image PI illustrated in FIG. 3, two detection rectangles BB surrounding the detection target OB estimated to be a person are set. Specifically, a detection rectangle BB(1) and a detection rectangle BB(2) are set.

[0080] In the captured image PI illustrated in FIG. 3, among the detection rectangles BB surrounding the detection target OB estimated to be a person, the detection target OB surrounded by the detection rectangle BB(1) is actually a "person". However, the detection target OB surrounded by the detection rectangle BB(2) is not actually a "person".

[0081] As illustrated in FIG. 3, even when using a learned model, it is not easy to accurately detect (estimate) "the person captured in the captured image PI" from the captured image PI.

[0082] (Suppression of false detection) Therefore, when detecting a person (human body) by image analysis of the captured image PI, the detection device 10 uses a learned model 140 constructed by multi-class learning to reduce false detection of a person. That is, the detection device 10 divides the captured image PI into a plurality of regions AR, and estimates the region AR including the position PS (for example, the position corresponding to the feet of the "person", the foot position) where the detection target OB estimated to be a person exists as the class of the detection target OB. Then, the detection device 10 verifies the correctness of the estimation that "the detection target OB is a person" using the estimated class of the detection target OB.

[0083] That is, the detection device 10 learns, in addition to the rectangle information given as a label to the captured image PI in the teacher data DT, the "type (class) of the person captured in the captured image PI" given as a label to the captured image PI, and constructs the learned model 140.

[0084] Specifically, in the teacher data DT learned by the detection device 10, rectangle information specifying the position, shape, and size of the detection rectangle BB surrounding the person captured in the captured image PI is given as a label to the captured image PI.

[0085] Further, the detection device 10 divides the captured image PI into a plurality of regions AR. In the teacher data DT that the detection device 10 learns, information (class) for specifying the region AR to which "the position PS of the person captured in the captured image PI (for example, the position corresponding to the person's feet. Foot position)" belongs is given as a label to the captured image PI. The detection device 10 learns, as the class of the person, the identification information of the region AR to which "the position PS where the person captured in the captured image PI exists" belongs, in other words, the identification information of the region AR including "the position PS where the person captured in the captured image PI exists".

[0086] The detection device 10 performs learning (machine learning) on a data set DS that is a set of "teacher data DT in which rectangular information and a class are given as labels to the captured image PI", and constructs a learned model 140.

[0087] Although details will be described later, for example, the captured image PI is divided into a plurality of regions AR by a plurality of straight lines (parts of the plurality of straight lines) intersecting at a predetermined angle with each other at the approximate center of the captured image PI, and at least one of the distances from the approximate center.

[0088] The detection device 10 performs estimation using the learned model 140 constructed by learning the detection information and class of the person captured in the captured image PI, and thereby estimates, from the captured image PI, a detection rectangle BB surrounding the person and the region AR (class) to which the position PS where the person exists belongs.

[0089] The multi-class learning performed by the detection device 10 refers to machine learning in which, in addition to learning "where to place the detection rectangle BB" (rectangular information), the class of the detection target OB (object), that is, the region AR including "the position PS where the detection target OB exists" is learned. By using the learned model 140 constructed by multi-class learning, the detection device 10 estimates, in addition to estimating "where to place the detection rectangle BB", the class of the detection target OB, that is, the region AR to which "the position PS where the detection target OB exists" belongs.

[0090] The detection device 10 uses the estimated class of the detection target OB to verify whether the estimation that "the detection target OB is a person" is correct, and thereby detects (estimates) a person imaged in the captured image PI from the captured image PI with high accuracy.

[0091] In particular, the detection device 10 specifies the area AR where "the detection rectangle BB arranged for the captured image PI using the learned model 140" exists as the specific area IA. Then, the detection device 10 uses the estimated class of the detection target OB and the specified specific area IA to verify whether the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct.

[0092] For example, the detection device 10 compares the consistency between the class output from the learned model 140 and the position of the detection rectangle BB (more precisely, the area AR where the detection rectangle BB exists), and if the two are different, removes the detection rectangle BB as a false detection.

[0093] Also, for example, the detection device 10 determines the validity of the detection rectangle BB using the class output from the learned model 140, the area AR where the detection rectangle BB exists, and the object reliability OR of the detection rectangle BB.

[0094] Summarizing the outline of the detection device 10 described so far with reference to FIGS. 2 and 3, it is as follows. That is, the detection device 10 is a detection device that detects a person imaged in the captured image PI from the captured image PI captured by the ceiling camera 20 using a fish-eye lens. The detection device 10 includes a division unit 120, a region estimation unit 152 (estimation unit 150), and a determination unit 190.

[0095] The division unit 120 divides the captured image PI into a plurality of regions AR. The estimation unit 150 (for example, the region estimation unit 152) calculates the probability PR including the position PS where the detection target OB estimated to be "a person" exists for each of the plurality of regions AR. The determination unit 190 determines whether "the detection target OB is a person" using the probability PR of each of the plurality of regions AR.

[0096] According to the above configuration, the detection device 10 divides the captured image PI into a plurality of regions AR, and for each of the plurality of regions AR, calculates the probability PR that includes the position PS where the detection target OB estimated to be a person exists. Then, the detection device 10 verifies the estimation result that "the detection target OB is a person" using the probability PR of each of the plurality of regions AR. The probability PR is the probability that the region AR includes the position PS where the detection target OB estimated to be a person exists.

[0097] Here, generally, when trying to detect a person from an image using a dictionary indicating human features, analysis of the image is required for each dictionary. Therefore, when trying to improve the detection accuracy of a person using a plurality of dictionaries indicating human features, analysis of the image is also required multiple times, and the time required to detect a person becomes longer.

[0098] In contrast, the detection device 10 determines whether the detection target OB is a person using the probability PR that includes the position PS where the detection target OB estimated to be a person exists, which is calculated for each of the plurality of regions AR obtained by dividing the captured image PI, thereby improving the detection accuracy of a person. That is, the detection device 10 does not detect a person from the captured image PI using a plurality of dictionaries indicating human features, but determines whether the detection target OB estimated to be a person is actually a person by using the probability PR of each of the plurality of regions AR, thereby improving the detection accuracy of a person.

[0099] While the method of using a plurality of dictionaries indicating human features improves the accuracy of the estimation itself, the detection device 10 improves the detection accuracy by verifying the estimated result (the estimation that the detection target OB is a person) (that is, removing an incorrect estimation result).

[0100] Therefore, the detection device 10 does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image PI, and can shorten the time required to detect a person from the captured image PI compared to the case of detecting a person from the captured image PI using a plurality of dictionaries.

[0101] In addition, since the detection device 10 determines (verifies) whether the detection target OB estimated to be a person is actually a person by using the probability PR of each of the plurality of regions AR, a person can be detected with high accuracy from the captured image PI.

[0102] Therefore, the detection device 10 has the effect of being able to detect a person from the image (captured image PI) captured using the fisheye lens quickly and with high accuracy.

[0103] In addition, as described above, the detection device 10 does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image PI.

[0104] Therefore, the detection device 10 has the effect of reducing the labor required for preparing a dictionary (for example, a learned model) necessary for detecting a person from the image captured using the fisheye lens, and also reducing the memory capacity required to store the dictionary.

[0105] The detection device 10 further includes a specifying unit 170 that specifies, as a specific region IA, a region AR among the plurality of regions AR in which a detection rectangle BB surrounding the detection target OB exists. Then, the determination unit 190 determines "whether the detection target OB is a person" by using the probability PR of each of the plurality of regions AR and the specific region IA.

[0106] According to the above configuration, the detection device 10 specifies a specific region IA that is a region AR in which the detection rectangle BB exists among the plurality of regions AR, and determines "whether the detection target OB is a person" by using the probability PR of each of the plurality of regions AR and the specific region IA.

[0107] For example, when the detection target OB estimated to be a person is actually a person, it is considered that the consistency between the probability PR of each of the plurality of regions AR and the region AR (that is, the specific region IA) in which the detection rectangle BB surrounding the detection target OB exists is also high.

[0108] Therefore, the detection device 10 determines whether the detection target OB is a person by using the probability PR of each of the plurality of regions AR and the specific region IA, that is, verifies the estimation that "the detection target OB is a person".

[0109] Therefore, the detection device 10 has the effect of being able to detect a person with high accuracy from the image (captured image PI) captured using the fisheye lens.

[0110] The specifying unit 170 may specify, as the specific region IA, the position PS of the detection target OB (for example, the position corresponding to the feet of the detection target OB estimated to be "a person") surrounded by the detection rectangle BB and estimated to be "a person", or the region AR including the position PS of the detection target OB. The specifying unit 170 may calculate the position PS (foot position) of the detection target OB from the detection rectangle BB, for example, from the position, shape, and size of the detection rectangle BB. Further, the specifying unit 170 may specify the center position (or centroid position) of the detection rectangle BB as the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be "a person".

[0111] In the detection device 10, the region estimation unit 152 (estimation unit 150) calculates the probability PR of each of the plurality of regions AR from the captured image PI using the region prediction model 142 (learned model 140). The region prediction model 142 (learned model 140) is a learned model that takes the captured image PI as an input and outputs the probability PR of each of the plurality of regions AR including "the position PS where the detection target OB surrounded by the detection rectangle BB exists".

[0112] According to the above configuration, the detection device 10 calculates the probability PR of each of the plurality of regions AR from the captured image PI using the region prediction model 142 (learned model 140). Therefore, the detection device 10 has the effect of being able to calculate the probability PR of each of the plurality of regions AR from the captured image PI with high accuracy using the region prediction model 142 (learned model 140).

[0113] The detection device 10 may include a learning unit 220 that constructs a region prediction model 142 (trained model 140) by machine learning for the following teacher data DT. That is, the learning unit 220 performs machine learning on the captured image PI with respect to the teacher data DT to which "information indicating the region AR including the position PS where the person captured in the captured image PI exists (that is, the class)" is assigned as a label.

[0114] According to the above configuration, the detection device 10 constructs a region prediction model 142 (trained model 140) by machine learning for the teacher data DT to which "information indicating the region AR including the position PS where the person exists" is assigned as a label with respect to the captured image PI.

[0115] Therefore, the detection device 10 has the effect that it can construct a region prediction model 142 (trained model 140) by machine learning for the teacher data DT, which enables it to accurately calculate the probability PR of each of the plurality of regions AR from the captured image PI.

[0116] §2. Configuration Example So far, the outline of the person detection system 1 has been described with reference to FIGS. 2 and 3. Next, the details of the configuration of the detection device 10 will be described with reference to FIG. 1.

[0117] (Detailed Configuration of Detection Device) FIG. 1 is a block diagram showing the main configuration of the detection device 10. As shown in FIG. 1, the detection device 10 includes, as functional blocks, a storage unit 100, an image acquisition unit 110, a segmentation unit 120, an estimation unit 150, a specification unit 170, a determination unit 190, a teacher data generation unit 210, and a learning unit 220.

[0118] In addition to the above-described functional blocks, the detection device 10 may include a display unit that displays the captured image PI and the like as a result of the human detection process, a communication unit that outputs these as data to an external device, and the like. However, for the sake of simplicity of description, configurations not directly related to the present embodiment are omitted from the description and the block diagram. However, depending on the actual implementation situation, the detection device 10 may include the omitted configuration.

[0119] The above-described functional blocks included in the detection device 10 can be realized, for example, by a computing device reading a program stored in a storage device (storage unit 100) realized by a ROM (read only memory), NVRAM (non-Volatile random access memory), etc. into a RAM (random access memory) not shown and executing it. Examples of devices that can be used as the computing device include a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating point number Processing Unit), PPU (Physics Processing Unit), a microcontroller, or a combination thereof.

[0120] First, the details of each of the image acquisition unit 110, segmentation unit 120, estimation unit 150, specification unit 170, determination unit 190, teacher data generation unit 210, and learning unit 220 will be described below.

[0121] (Details of Functional Blocks Other than the Storage Unit) The image acquisition unit 110 acquires the captured image PI captured by the ceiling camera 20 from the ceiling camera 20. The image acquisition unit 110 outputs the captured image PI acquired from the ceiling camera 20 to the segmentation unit 120.

[0122] The splitting unit 120 acquires the captured image PI from the image acquisition unit 110, and splits the acquired captured image PI into a plurality of regions AR by a predetermined splitting method. The "predetermined splitting method" may be set by the user in advance, or the user may be able to update (change) the "predetermined splitting method" set in advance. For example, the "predetermined splitting method" set by the user in advance is stored in the storage unit 100. The splitting unit 120 stores the captured image PI split into a plurality of regions AR in the storage unit 100 as split image information 130.

[0123] The estimation unit 150 performs image analysis on the captured image PI (that is, the captured image PI stored in the split image information 130) split into a plurality of regions AR by the splitting unit 120, and generates (outputs) rectangle information and a class. The estimation unit 150 may output rectangle information and a class from the captured image PI that has not been split into a plurality of regions AR by the splitting unit 120 by using the learned model 140. That is, the captured image PI input by the estimation unit 150 to the learned model 140 may or may not be split into a plurality of regions AR by the splitting unit 120.

[0124] In the example shown in FIG. 1, the estimation unit 150 includes a person estimation unit 151 and a region estimation unit 152. The person estimation unit 151 outputs rectangle information from the captured image PI by using the learned model 140 (particularly, the person prediction model 141). The region estimation unit 152 outputs the class of the detection target OB (specifically, the probability PR of each of the plurality of regions AR calculated for each detection rectangle BB) surrounded by the detection rectangle BB from the captured image PI by using the learned model 140 (particularly, the region prediction model 142).

[0125] In FIG. 1, for the sake of easy understanding of the detection device 10, the person estimation unit 151 and the area estimation unit 152 are separately described. However, the estimation unit 150 including the person estimation unit 151 and the area estimation unit 152 may be realized by one neural network. In other words, the functions of the person estimation unit 151 and the area estimation unit 152 may be realized by one neural network. Further, the neural network realizing the function of the person estimation unit 151 and the neural network realizing the function of the area estimation unit 152 may be separated.

[0126] In the following description, an example of realizing the estimation unit 150 as one neural network will be described. The estimation unit 150 may be realized as, for example, a CNN (Convolutional Neural Network), a Faster R-CNN (Regional Convolutional Neural Network), or a DNN (Deep Neural Network).

[0127] The estimation unit 150 realized as a neural network outputs rectangle information and a class from the captured image PI using the learned model 140. Further, for each detection rectangle BB, the estimation unit 150 outputs an object reliability OR which is information indicating the likelihood that some object is surrounded by the detection rectangle BB.

[0128] The object reliability OR is information indicating whether the detection rectangle BB contains some object (detection target OB) or only the background, and the value thereof indicates the likelihood that the detection rectangle BB contains some object. That is, the object reliability OR is information indicating the magnitude of the possibility that the detection rectangle BB surrounds some detection target OB (the possibility that some object is surrounded by the detection rectangle BB).

[0129] First, the estimation unit 150 refers to the storage unit 100 to obtain the learned model 140. Next, the estimation unit 150 refers to the divided image information 130 of the storage unit 100 to obtain the captured image PI divided into a plurality of regions AR. Then, the estimation unit 150 uses the learned model 140 to output rectangle information, a class, and an object reliability OR from the captured image PI. That is, the estimation unit 150 arranges one or more detection rectangles BB in the captured image PI, and for each arranged detection rectangle BB, calculates the probability PR that "each of the plurality of regions AR includes the position PS of the detection target OB surrounded by the detection rectangle BB". Further, the estimation unit 150 calculates the object reliability OR for each detection rectangle BB.

[0130] The estimation unit 150 stores the estimation result estimated from the captured image PI using the learned model 140 in the estimation result information table 160. Specifically, the estimation unit 150 stores the rectangle information, the class, and the object reliability OR in the rectangle information table 161, the probability table 162, and the object reliability table 163 of the storage unit 100, respectively.

[0131] The specifying unit 170 specifies, for each of the one or more detection rectangles BB arranged in the captured image PI by the estimation unit 150, the region AR in which the detection rectangle BB exists, that is, the region AR including the detection rectangle BB, as a specific region IA.

[0132] Specifically, the specifying unit 170 refers to the rectangle information table 161 of the storage unit 100 to obtain the position, shape, and size of each of the one or more detection rectangles BB arranged in the captured image PI. Further, the specifying unit 170 refers to the divided image information 130 of the storage unit 100 to obtain the position, shape, and size of each of the plurality of regions AR into which the captured image PI is divided. The specifying unit 170 specifies, for each of the one or more detection rectangles BB arranged in the captured image PI, which region AR it is included in based on the "position, shape, and size of the detection rectangle BB" and the "position, shape, and size of each of the plurality of regions AR".

[0133] The specific part 170 may specify "the position PS of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'", for example, it may also specify "the position corresponding to the feet of the detection target OB estimated to be 'a person' (foot position)". The specific part 170 calculates the position PS (foot position) of the detection target OB from the detection rectangle BB, for example, it may calculate from the position, shape, and size of the detection rectangle BB. Also, the specific part 170 may specify the center position (or centroid position) of the detection rectangle BB as "the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'". The specific part 170 may specify the area AR where the position PS (foot position) of the detection target OB exists, that is, the area AR including the position PS, as the specific area IA.

[0134] The specific part 170 stores the specific area IA specified for each detection rectangle BB in the specific area table 180 of the storage part 100. The specific part 170 may specify "the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'" as the specific area IA for each detection rectangle BB, and store the position PS for each specified detection rectangle BB in the specific area table 180.

[0135] The determination part 190 determines whether "the detection target OB estimated to be a person by the estimation part 150 is actually a person" using the class of the detection target OB. That is, the determination part 190 determines the correctness of the rectangle information output from the captured image PI by the estimation part 150 using the learned model 140 using the class output from the captured image PI by the estimation part 150 using the learned model 140. In other words, the determination part 190 determines whether "the detection target OB surrounded by the detection rectangle BB is actually a person" using the class of the detection target OB.

[0136] When the determination unit 190 determines that "the detection target OB surrounded by the detection rectangle BB is actually a person", it adopts the detection rectangle BB (that is, the detection target OB) as a correct estimation. When the determination unit 190 determines that "the detection target OB surrounded by the detection rectangle BB is not actually a person", it removes the detection rectangle BB (that is, the detection target OB) as a false detection (false estimation).

[0137] In particular, the determination unit 190 determines whether "the detection target OB surrounded by the detection rectangle BB is actually a person" based on the class of the detection target OB surrounded by the detection rectangle BB and the specific area IA which is the area AR where the detection rectangle BB exists.

[0138] Specifically, the determination unit 190 refers to the estimation result information table 160, and in particular, refers to the probability table 162 to obtain the "PR of each of the plurality of areas AR" for each detection rectangle BB. Also, the determination unit 190 refers to the estimation result information table 160, and in particular, refers to the object reliability table 163 to obtain the object reliability OR for each detection rectangle BB. Furthermore, the determination unit 190 refers to the specific area table 180 to obtain the specific area IA for each detection rectangle BB. The determination unit 190 determines whether "the detection target OB surrounded by the detection rectangle BB is a person" by using "at least one of the 'PR of each of the plurality of areas AR' for each detection rectangle BB and the object reliability OR for each detection rectangle BB" and the specific area IA for each detection rectangle BB.

[0139] For example, in the first verification method, the determination unit 190 determines, for each detection rectangle BB, whether "the region AR with the highest probability PR that 'includes the position PS of the detection target OB surrounded by the detection rectangle BB' coincides with the specific region IA". When the determination unit 190 confirms that the region AR with the highest probability PR coincides with the specific region IA of the detection rectangle BB, it determines that "the detection target OB surrounded by the detection rectangle BB is a person". When the determination unit 190 confirms that the region AR with the highest probability PR does not coincide with the specific region IA of the detection rectangle BB, it determines that "the detection target OB surrounded by the detection rectangle BB is not a person (that is, the detection rectangle BB is a false detection)".

[0140] For example, in the second verification method, the determination unit 190 determines, for each detection rectangle BB, whether "the class confidence value CR of the detection rectangle BB, calculated by multiplying the probability PR of the region AR corresponding to the specific region IA by the object reliability OR of the detection rectangle BB, is greater than a predetermined value TH". When the determination unit 190 confirms that "the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH", it determines that "the detection target OB surrounded by the detection rectangle BB is a person". When the determination unit 190 confirms that "the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH", it determines that "the detection target OB surrounded by the detection rectangle BB is not a person (that is, the detection rectangle BB is a false detection)".

[0141] The teacher data generation unit 210 assigns, as labels, the "rectangle information and class" received from the user to the captured image PI (that is, the divided image information 130) divided into a plurality of regions AR by the division unit 120, and generates teacher data DT.

[0142] The learning unit 220 constructs a learned model 140 by supervised learning with respect to a data set DS that is a set of teacher data DT generated by the teacher data generation unit 210. The learning unit 220 stores the constructed learned model 140 in the storage unit 100.

[0143] In the example shown in FIG. 1, the learning unit 220 includes a person learning unit 221 that learns rectangle information and a region learning unit 222 that learns the class for each detection target OB surrounded by the detection rectangle BB. In FIG. 1, for the sake of easy understanding of the detection device 10, the person learning unit 221 and the region learning unit 222 are described separately, but the learning unit 220 including the person learning unit 221 and the region learning unit 222 may be realized by one neural network. In other words, the functions of the person learning unit 221 and the region learning unit 222 may be realized by one neural network. Also, the neural network that realizes the function of the person learning unit 221 and the neural network that realizes the function of the region learning unit 222 may be separated.

[0144] In the following description, an example in which the learning unit 220 is realized as one neural network will be described. The learning unit 220 may be realized, for example, as a CNN (Convolutional Neural Network), a Faster R-CNN (Regional Convolutional Neural Network), or a DNN (Deep Neural Network).

[0145] (Details of the memory unit) The storage unit 100 is a storage device that stores various data used by the detection device 10. Note that the storage unit 100 may non-temporarily store (1) a control program executed by the detection device 10, (2) an OS program, (3) an application program for executing various functions of the detection device 10, and (4) various data read when executing the application program. The data of (1) to (4) above is stored in a non-volatile storage device such as, for example, a ROM (read only memory), a flash memory, an EPROM (Erasable Programmable ROM), an EEPROM (registered trademark) (Electrically EPROM), or an HDD (Hard Disc Drive). The detection device 10 may include a temporary storage unit (not shown). The temporary storage unit is a so-called working memory that temporarily stores data used for calculations and calculation results during various processes executed by the detection device 10, and is composed of a volatile storage device such as a RAM (Random Access Memory). Which data is stored in which storage device is appropriately determined based on the purpose of use, convenience, cost, or physical constraints of the detection device 10. The storage unit 100 further stores divided image information 130, a learned model 140, an estimation result information table 160, and a specific area table 180.

[0146] The divided image information 130 is information indicating the captured image PI divided into a plurality of regions AR by the dividing unit 120, and is information stored by the dividing unit 120.

[0147] The learned model 140 is a learned model that takes the captured image PI as input and outputs rectangle information and, for each detection rectangle BB, the probability PR (i.e., class) that each of the plurality of regions AR "includes the position PS where the detection target OB surrounded by the detection rectangle BB exists". The learned model 140 is constructed by the learning unit 220 and stored in the storage unit 100 by the learning unit 220.

[0148] In the example shown in FIG. 1, the learned model 140 includes a person prediction model 141 and a region prediction model 142. The person prediction model 141 is a learned model that takes the captured image PI as input and outputs rectangular information (information indicating the position, shape, and size of the detection rectangle BB that encloses the detection target OB estimated as a "person"). The region prediction model 142 is a learned model that takes the captured image PI as input and outputs the class for each detection target OB surrounded by the detection rectangle BB (that is, the probability PR for each of the plurality of regions AR).

[0149] In FIG. 1, for the sake of easy understanding of the detection device 10, the person prediction model 141 and the region prediction model 142 are described separately, but a learned model 140 that integrally includes the person prediction model 141 and the region prediction model 142 may be stored in the storage unit 100.

[0150] The estimation result information table 160 stores the estimation results estimated (output) by the estimation unit 150 from the captured image PI using the learned model 140. The estimation result information table 160 includes a rectangular information table 161, a probability table 162, and an object reliability table 163.

[0151] The rectangular information table 161 stores the rectangular information output by the estimation unit 150 from the captured image PI using the learned model 140. In other words, the rectangular information table 161 stores information indicating the position, shape, and size of each of the one or more detection rectangles BB arranged by the estimation unit 150 in the captured image PI using the learned model 140.

[0152] The probability table 162 stores the classes output by the estimation unit 150 from the captured image PI using the learned model 140 (more precisely, the classes for each detection target OB surrounded by the detection rectangle BB). In other words, the probability table 162 stores, for each detection rectangle BB, the "probability PR for each of the plurality of regions AR that includes the position PS of the detection target OB surrounded by that detection rectangle BB".

[0153] For example, for the detection rectangle BB(x), probabilities PR(x - 0), PR(x - 1), PR(x - 2), ···, PR(x - n), which are the respective probabilities PR from region AR(0) to region AR(n), are stored in the probability table 162. Similarly, for example, for the detection rectangle BB(y), probabilities PR(y - 0), PR(y - 1), PR(y - 2), ···, PR(y - n), which are the respective probabilities PR from region AR(0) to region AR(n), are stored in the probability table 162.

[0154] In the object reliability table 163, for each of one or more detection rectangles BB arranged by the estimation unit 150 in the captured image PI using the learned model 140, object reliability OR, which is information indicating the likelihood that some object is surrounded by the detection rectangle BB, is stored. In other words, the object reliability table 163 stores the object reliability OR for each detection rectangle BB.

[0155] In the specific region table 180, information for identifying the specific region IA for each detection rectangle BB is stored. In other words, in the specific region table 180, for each detection rectangle BB, information for identifying the specific region IA, which is the region AR including the detection rectangle BB, is stored.

[0156] In the specific region table 180, as the specific region IA for each detection rectangle BB, "the position PS (e.g., the foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'" may be stored. The position PS (foot position) of the detection target OB is calculated from the detection rectangle BB, for example, from the position, shape, and size of the detection rectangle BB. Also, the specific region IA for each detection rectangle BB may be a region AR including "the position PS of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'". As described above, "the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be 'a person'" may be the center position or the centroid position of the detection rectangle BB.

[0157] §3. Operation Example Below, various processes executed by the detection device 10 will be described with reference to FIGS. 4 to 19. Specifically, the learning process executed by the detection device 10 will be described using FIGS. 4 to 6, and the human detection process executed by the detection device 10 will be described using FIGS. 7 to 19.

[0158] (Overview of the learning process) The detection device 10 executes, for example, a learning process (that is, a process for generating the learned model 140). In the learning process, the detection device 10 (particularly, the teacher data generation unit 210) first assigns rectangle information and a class as labels to the captured image PI divided into a plurality of regions AR, and generates teacher data DT.

[0159] The rectangle information is information indicating "the position, shape, and size of the detection rectangle BB surrounding the person captured in the captured image PI". For example, the rectangle information is "the X-axis coordinate and Y-axis coordinate of the upper right point of the detection rectangle BB" and "the X-axis coordinate and Y-axis coordinate of the lower left point of the detection rectangle BB".

[0160] The class is identification information of the region AR including "the position PS where the person captured in the captured image PI exists (for example, the position corresponding to the feet of that person, the foot position)". The class can also be rephrased as information for specifying the region AR including the "detection rectangle BB surrounding the person captured in the captured image PI".

[0161] Then, the detection device 10 (particularly, the learning unit 220) performs learning (supervised learning) on the data set DS which is a set of the teacher data DT, and constructs the learned model 140. The learning unit 220 stores the constructed learned model 140 in the storage unit 100.

[0162] As described above, in the detection device 10, the learning unit 220 constructs the learned model 140 by performing learning (for example, deep learning) on the data set DS which is a set of the teacher data generated by the teacher data generation unit 210. Thus, the detection device 10 also has a function as a model generation device for constructing the learned model 140.

[0163] (Example of learning process) FIG. 4 is a flowchart for explaining an example of a generation process (learning process) for generating a learned model 140 used by the detection device 10. The learned model 140 is generated, for example, by the detection device 10 performing learning (machine learning) on a data set DS that is a set of teacher data DT. However, it is not essential that the entity that generates the learned model 140 by learning on the data set DS is the detection device 10. Another device different from the detection device 10 may include a teacher data generation unit 210 and a learning unit 220, and generate the learned model 140 by learning on the data set DS.

[0164] Hereinafter, an example in which the detection device 10 generates (constructs) a learned model 140 by performing learning (for example, supervised learning) on the data set DS will be described.

[0165] As illustrated in FIG. 4, first, the teacher data generation unit 210 generates teacher data DT obtained by classifying the imaged person for each position PS of the imaged person (more precisely, the area AR including the position PS) (S110). That is, in the teacher data DT generated by the teacher data generation unit 210, the captured image PI is given the class of the person captured in the captured image PI (identification information of the area AR including the position PS (foot position) of that person).

[0166] Specifically, first, the teacher data generation unit 210 refers to the divided image information 130 of the storage unit 100 and acquires the "captured image PI divided into a plurality of areas AR". The teacher data generation unit 210 presents the acquired "captured image PI divided into a plurality of areas AR" to the user and prompts the user for an annotation for the "captured image PI divided into a plurality of areas AR". Then, the teacher data generation unit 210 acquires rectangular information and a class as labels (label information) to be given to the "captured image PI divided into a plurality of areas AR" from the annotation by the user.

[0167] The learning unit 220 acquires the teacher data DT generated by the teacher data generation unit 210 from the teacher data generation unit 210 (S120). The learning unit 220 performs multi-class learning (machine learning) on a data set DS that is a set of the teacher data DT acquired from the teacher data generation unit 210 (S130), and generates (constructs) a learned model 140. For example, the learning unit 220 generates the learned model 140 by supervised learning with respect to the data set DS.

[0168] The learning unit 220 stores the generated learned model 140 in the storage unit 100 (S140).

[0169] As described above with reference to FIG. 4, the model generation method executed by the detection device 10 is a "model generation method by a model generation device that generates a learned model", and includes the following two processes. That is, the model generation method includes an acquisition step (S120) of acquiring the teacher data DT and a learning step (S130) of constructing the learned model 140 by learning with respect to a data set DS that is a set of the teacher data DT.

[0170] In the teacher data DT acquired in the acquisition step, the following information is assigned as a label to the captured image PI divided into a plurality of regions AR. That is, (A) information indicating the position, shape, and size of a detection rectangle BB surrounding the person captured in the captured image PI (rectangle information), and (B) information (class) for identifying a region AR including the position PS where the person captured in the captured image PI exists, are assigned as labels.

[0171] The learned model 140 constructed in the learning step takes the captured image PI as an input, and outputs (C) the rectangle information and (D) the probability PR including the position PS where the person captured in the captured image PI exists for each of the plurality of regions AR, as a learned model.

[0172] According to the above configuration, the model generation method constructs a trained model 140 by machine learning on teacher data DT (more precisely, a data set DS which is a set of teacher data DT). When the trained model 140 is input with a captured image PI, it outputs the following two pieces of information. That is, (C) information indicating the position, shape, and size of the detection rectangle BB (rectangle information), and (D) for each of the plurality of regions AR, the probability PR (including the position PS where the person captured in the captured image PI exists) (probability).

[0173] The rectangle information is information indicating the position, shape, and size of the detection rectangle BB surrounding the detection object OB estimated to be "a person", and includes the estimation that the detection object OB surrounded by the detection rectangle BB is "a person".

[0174] Therefore, the model generation method can construct a trained model 140 that outputs, when input with a captured image PI, the rectangle information including the estimation that the detection object OB is "a person" and the probability PR (i.e., class) for each of the plurality of regions AR.

[0175] (Example of division of captured image) FIG. 5 is a diagram showing an example of a captured image PI divided into a plurality of regions AR by the detection device 10. For example, the detection device 10 (particularly, the division unit 120) may divide the captured image PI into a plurality of regions AR according to the distance from the center of the captured image PI. In other words, the division unit 120 may divide the captured image PI into a plurality of regions AR with the circumferences of one or more concentric circles (or ellipses) centered on "the center of the captured image PI" as the boundary lines.

[0176] Also, for example, the division unit 120 may divide the captured image PI into a plurality of sector-shaped regions AR with a central angle of a predetermined angle. In other words, the division unit 120 may divide the captured image PI into a plurality of regions AR with a part of a plurality of straight lines (a plurality of straight lines extending radially from the approximate center of the captured image PI and intersecting at a predetermined angle with each other) intersecting at a predetermined angle with each other at the approximate center of the captured image PI as the boundary lines.

[0177] Furthermore, for example, the dividing unit 120 may divide the captured image PI into a plurality of regions AR having as a boundary at least one of the circumferences of one or more concentric circles (ellipses) centered on the "substantially central point of the captured image PI" and a part of a plurality of straight lines intersecting at a predetermined angle with each other at the substantially central point of the captured image PI.

[0178] In the example shown in FIG. 5(A), the dividing unit 120 designates the inside of one circle (ellipse) centered on the "center of the captured image PI" within the captured image PI as the region AR(0). Further, the dividing unit 120 divides the region outside the above-described circle (ellipse) centered on the "center of the captured image PI" within the captured image PI into six regions AR from the region AR(1) to the region AR(6) by a part of a plurality of straight lines intersecting at 60 degrees with each other at the center of the captured image PI. Therefore, the captured image PI illustrated in FIG. 5(A) is divided by the dividing unit 120 into seven regions AR, namely, the regions AR(0), AR(1), AR(2), ···, AR(6).

[0179] In FIG. 5(A), the region AR(0) may be further divided by a part of a plurality of straight lines intersecting at a predetermined angle with each other at the center of the captured image PI by the dividing unit 120. For example, the dividing unit 120 may divide the region AR(0) in FIG. 5(A) into further regions AR(0-0), AR(0-1), AR(0-2), ···, AR(0-n).

[0180] In the example shown in FIG. 5(B), the dividing unit 120 divides the captured image PI into six regions AR from the region AR(0) to the region AR(5) by a plurality of straight lines (parts of straight lines) intersecting at 60 degrees with each other at the center of the captured image PI. Therefore, the captured image PI illustrated in FIG. 5(B) is divided by the dividing unit 120 into six regions AR, namely, the regions AR(0), AR(1), AR(2), ···, AR(5).

[0181] In the example shown in (C) of FIG. 5, the dividing unit 120 divides the captured image PI into 16 regions AR from region AR(0) to region AR(15) by a part of a plurality of straight lines intersecting at 22.5 degrees with each other at the center of the captured image PI. Therefore, the captured image PI illustrated in (C) of FIG. 5 is divided by the dividing unit 120 into 16 regions AR, namely, region AR(0), region AR(1), region AR(2), ···, region AR(15).

[0182] Dividing the captured image PI by a part of a plurality of straight lines intersecting at 22.5 degrees or 60 degrees with each other at the center of the captured image PI is not essential for the dividing unit 120. The dividing unit 120 may divide the captured image PI by a part of a plurality of straight lines (a part of the straight lines) intersecting at 1 degree with each other at a substantially center of the captured image PI, or may divide the captured image PI by a part of a plurality of straight lines (a part of the straight lines) intersecting at 45 degrees with each other at a substantially center of the captured image PI.

[0183] In the example shown in (D) of FIG. 5, the dividing unit 120 designates the inside of the smaller one of two concentric circles (ellipses) centered on the "center of the captured image PI" in the captured image PI as region AR(0). Further, the detection device 10 designates, in the captured image PI, the region inside the larger one of two concentric circles centered on the "center of the captured image PI" (ellipse) and outside the smaller one of the concentric circles (ellipse) as region AR(1). Furthermore, the detection device 10 divides, in the captured image PI, the region outside the above-described larger concentric circle (ellipse) into 6 regions AR from region AR(2) to region AR(7) by a part of a plurality of straight lines intersecting at 60 degrees with each other at the center of the captured image PI. Therefore, the captured image PI illustrated in (D) of FIG. 5 is divided by the dividing unit 120 into 8 regions AR, namely, region AR(0), region AR(1), region AR(2), ···, region AR(7).

[0184] As described with reference to FIG. 5, the captured image PI is divided by the dividing unit 120 into a plurality of regions AR by a predetermined dividing method. The "predetermined dividing method" is set and changed by a user, for example.

[0185] (An example of the learning process) FIG. 6 is a diagram showing an example of teacher data DT learned in the generation process (learning process) of FIG. 4. The teacher data DT is generated by the teacher data generation unit 210 and learned by the learning unit 220. In the teacher data DT illustrated in FIG. 6, the captured image PI is divided by the division unit 120 into seven regions AR from region AR(0) to region AR(6). Region AR(0) is a region inside one concentric circle (ellipse) centered on the "center of the captured image PI" that substantially corresponds to the optical axis of the fisheye lens provided in the ceiling camera 20. Also, the region outside the concentric circle (ellipse) is divided by a plurality of straight lines intersecting at a predetermined angle (60 degrees in the example of FIG. 6) with each other at the center of the captured image PI into six regions AR from region AR(1) to region AR(6).

[0186] In the teacher data DT illustrated in FIG. 6, a detection rectangle BB(1) (that is, rectangle information indicating the position, shape, and size of the detection rectangle BB(1)) surrounding the person captured in the captured image PI is assigned as a label to the captured image PI.

[0187] Also, in the teacher data DT illustrated in FIG. 6, the number "1" is assigned as a label to the captured image PI as the class of the person captured in the captured image PI. The class of the person captured in the captured image PI is information for identifying the region AR including the position PS where the person captured in the captured image PI exists from other regions AR.

[0188] In the teacher data DT illustrated in FIG. 6, the position PS where the person imaged in the captured image PI exists is included in the region AR(1) among the regions AR(0) to AR(6) obtained by dividing the captured image PI. As described above, the "position PS where the person imaged in the captured image PI exists" is, for example, "the position corresponding to the feet of the person imaged in the captured image PI (foot position)". In the example shown in FIG. 6, the position corresponding to the feet of the person imaged in the captured image PI (the person surrounded by the detection rectangle BB(1)) is included in the region AR(1). Therefore, in the teacher data DT illustrated in FIG. 6, the class of the person imaged in the captured image PI is set to "1".

[0189] The detection device 10 (particularly, the learning unit 220) learns the data set DS which is a set of the teacher data DT illustrated in FIG. 6 to construct the learned model 140. That is, the learning unit 220 learns the teacher data DT in which the rectangular information and the class are attached as labels to the captured image PI. Specifically, the learning unit 220 learns the teacher data DT in which "the detection rectangle BB surrounding the person imaged in the captured image PI" and "the identification information of the region AR including the position PS where the person imaged in the captured image PI exists" are attached as labels to the captured image PI.

[0190] (Overview of person detection process) The detection device 10 executes a person detection process for detecting "the person imaged in the captured image PI" from the captured image PI. The person detection process includes an estimation process for performing estimation using the learned model 140 constructed by the learning process, and a determination process for determining the correctness of the result of the estimation process (estimation result).

[0191] In the estimation process, the detection device 10 (specifically, the estimation unit 150 (the learned model 140)) outputs rectangle information and a class by inputting the captured image PI into the learned model 140. Specifically, the learned model 140 outputs, as the rectangle information, information indicating the position, shape, and size of the detection rectangle BB that encloses the detection target OB estimated to be "a person". Further, the learned model 140 outputs, as the class, identification information of the area AR including the position PS where the detection target OB surrounded by the detection rectangle BB and estimated to be "a person" exists, that is, for each of the plurality of areas AR, it outputs "the probability PR including the position PS". The detection device 10 can improve the detection speed (estimation speed) and the detection accuracy (estimation accuracy) by executing the estimation process using, for example, the learned model 140 constructed by deep learning.

[0192] In the determination process, the detection device 10 (specifically, the determination unit 190) determines whether the estimation result is correct. Specifically, it determines the correctness of the "rectangle information output from the learned model 140" using the "class output from the learned model 140". Specifically, the determination unit 190 executes a first determination process or a second determination process in the determination process.

[0193] In the first determination process, the determination unit 190 determines the correctness of the class output from the learned model 140. If it determines that the class is correct, it determines that the detection information output by the learned model 140 is also correct. That is, when the determination unit 190 determines that the class is correct, it also determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct. Further, when the determination unit 190 determines that the class output from the learned model 140 is incorrect, it determines that the detection information output by the learned model 140 is also incorrect. That is, when the determination unit 190 determines that the class is incorrect, it also determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is incorrect.

[0194] In the second determination process, the determination unit 190 determines, for example, whether a class confidence value CR calculated by multiplying the probability PR of the region AR where the detection rectangle BB exists (that is, the region AR corresponding to the specific region IA) by the object reliability OR of the detection rectangle BB is greater than a predetermined value TH. When the determination unit 190 determines that the class confidence value CR is greater than the predetermined value TH, it determines that the detection information output by the learned model 140 is correct. That is, it determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct. When the determination unit 190 determines that the class confidence value CR is less than or equal to the predetermined value TH, it determines that the detection information output by the learned model 140 is incorrect. That is, it determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is incorrect.

[0195] Hereinafter, the first determination process will be described with reference to FIGS. 7 to 13, and the second determination process will be described with reference to FIGS. 14 to 19.

[0196] (Example of human detection process including the first determination process) FIG. 7 is a flowchart for explaining an example of a human detection process that executes the first determination process on the output of the learned model 140 to detect a person. As shown in FIG. 7, first, the image acquisition unit 110 acquires a captured image PI, and the division unit 120 divides the captured image PI acquired by the image acquisition unit 110 into a plurality of regions AR (S210). That is, first, the image acquisition unit 110 acquires a captured image PI captured by the ceiling camera 20 from the ceiling camera 20. The image acquisition unit 110 outputs the acquired captured image PI to the division unit 120. When the division unit 120 receives the captured image PI from the image acquisition unit 110, it divides the received captured image PI into a plurality of regions AR by a predetermined division method. The division unit 120 stores the captured image PI divided into a plurality of regions AR in the storage unit 100 as division image information 130.

[0197] The estimation unit 150 acquires the learned model 140 with reference to the storage unit 100. Further, the estimation unit 150 acquires the captured image PI (the captured image PI divided into a plurality of regions AR) for which the estimation process is to be executed, with reference to the divided image information 130 of the storage unit 100. The estimation unit 150 performs person detection and class output on the captured image PI (the captured image PI divided into a plurality of regions AR) acquired with reference to the divided image information 130, using the learned model 140 (S220). The estimation unit 150 may perform person detection and class output on the captured image PI that has not been divided into a plurality of regions AR by the dividing unit 120, using the learned model 140.

[0198] The estimation unit 150 stores the result of the estimation process for the captured image PI using the learned model 140 in the estimation result information table 160 of the storage unit 100. Specifically, the estimation unit 150 stores the rectangular information, class, and object reliability OR output by the learned model 140 input with the captured image PI in the storage unit 100.

[0199] The rectangular information is information indicating the position, shape, and size of each of one or more detection rectangles BB arranged in the captured image PI by the estimation unit 150 (learned model 140), that is, information indicating the position, shape, and size for each detection rectangle BB. The rectangular information is stored in the rectangular information table 161 of the storage unit 100.

[0200] The class is the "probability PR of each of the plurality of regions AR, 'including the position PS of the detection object OB surrounded by the detection rectangle BB'" calculated for each detection object OB surrounded by the detection rectangle BB (that is, for each detection rectangle BB). The class is stored in the probability table 162 of the storage unit 100.

[0201] The object reliability OR is information indicating the magnitude of the "possibility that some object is surrounded by the detection rectangle BB" for each detection rectangle BB. The object reliability OR is stored in the object reliability table 163 of the storage unit 100.

[0202] The detection device 10 executes the processes from S230 to S270 for each detection rectangle BB, that is, for each detection target OB, by the number of people detected by the estimation unit 150 from the captured image PI, in other words, by the number of detection rectangles BB arranged by the estimation unit 150.

[0203] The determination unit 190 refers to the probability table 162 in the storage unit 100 and acquires the class (output class) for each detection target OB surrounded by the detection rectangle BB, that is, for each detection rectangle BB (S230), in other words, acquires the probability PR of each of the plurality of regions AR. Specifically, the determination unit 190 acquires, for each detection rectangle BB, the probability PR of each of the plurality of regions AR that includes the position PS where the detection target OB surrounded by the detection rectangle BB and estimated to be "a person" exists.

[0204] For example, the determination unit 190 acquires the probabilities PR(0-1), PR(1-1), PR(2-1), ···, PR(n-1) of each of the plurality of regions AR for the detection rectangle BB(1). Similarly, the determination unit 190 acquires the probabilities PR(0-2), PR(1-2), PR(2-2), ···, PR(n-2) of each of the plurality of regions AR for the detection rectangle BB(2).

[0205] The specifying unit 170 acquires, for each detection rectangle BB, the position PS (for example, the position corresponding to the feet position of the "person") of the detection target OB surrounded by the detection rectangle BB and estimated to be "a person" (S240). For example, the specifying unit 170 refers to the storage unit 100 to acquire the rectangle information table 161, and for each detection rectangle BB, calculates the "position PS (feet position) of the detection target OB" from the detection rectangle BB (the position, shape, and size of the detection rectangle BB). The specifying unit 170 may specify the center position (or centroid position) of the detection rectangle BB as the position PS (feet position) of the detection target OB. The specifying unit 170 sets the "position PS (feet position) of the detection target OB for each detection rectangle BB" calculated from the detection rectangle BB as the position PS of the detection target OB surrounded by the detection rectangle BB and estimated to be "a person" for each detection rectangle BB.

[0206] The determination unit 190 determines, for each detection target OB (that is, for each detection rectangle BB surrounding the detection target OB), whether the foot position (position PS) of the detection target OB is included in the output class (that is, the region AR with the largest probability PR) (S250).

[0207] Specifically, the determination unit 190 first refers to the probability table 162 in the storage unit 100 to obtain the "probability PR of each of the plurality of regions AR" for each detection target OB (that is, for each detection rectangle BB). Next, the determination unit 190 selects, for each detection target OB (that is, for each detection rectangle BB), the region AR with the highest probability PR. Then, the determination unit 190 determines whether the position PS (foot position) for each detection target OB is included in the "region AR with the highest probability PR" (that is, the output class) selected for each detection target OB (that is, for each detection rectangle BB).

[0208] When the determination unit 190 determines that "the position PS (foot position) of the detection target OB is included in the output class" (YES in S250), it adopts the detection rectangle BB as the correct detection result (S260). That is, the determination unit 190 determines that "the position PS (foot position) of the detection target OB is included in the output class" and that "the estimation that 'the detection target OB surrounded by the detection rectangle BB is a person' is correct".

[0209] When the determination unit 190 determines that "the position PS (foot position) of the detection target OB is not included in the output class" (NO in S250), it removes the detection rectangle BB as a false detection (S270). That is, the determination unit 190 determines that "the position PS (foot position) of the detection target OB is not included in the output class" and that "the estimation that 'the detection target OB surrounded by the detection rectangle BB is a person' is incorrect".

[0210] In the first determination process illustrated in FIG. 7, the detection device 10 (particularly, the determination unit 190) determines, for each detection target OB (that is, the detection rectangle BB), whether the position PS of the detection target OB is included in the “region AR with the highest probability PR”. However, it is not essential for the determination unit 190 to determine, in the first determination process, whether the position PS of the detection target OB is included in the “region AR with the highest probability PR”.

[0211] In the first determination process, it is sufficient for the determination unit 190 to be able to determine, for each detection target OB (that is, the detection rectangle BB), whether the “class of the detection target OB” output from the imaging image PI by the estimation unit 150 (trained model 140) is correct. In the first determination process, it is sufficient for the determination unit 190 to determine whether the position PS of the detection target OB estimated by the trained model 140 and the (actual) position PS of the detection target OB surrounded by the detection rectangle BB and estimated to be “a person” match at the level of the region AR. The (actual) position PS of the detection target OB surrounded by the detection rectangle BB and estimated to be “a person” is calculated from the detection rectangle BB (the position, shape, and size of the detection rectangle BB). Also, the position PS may be, for example, the center position (or the centroid position) of the detection rectangle BB.

[0212] That is, in the first determination process, it is sufficient for the determination unit 190 to be able to determine whether the class output from the trained model 140 (that is, the region AR with the highest probability PR) matches the region AR where the detection rectangle BB actually exists (that is, the specific region IA). Although details will be described later, in the first determination process, the determination unit 190 may determine whether the “region AR with the highest probability PR” or the “region AR adjacent to the region AR with the highest probability PR” matches the region AR where the detection rectangle BB actually exists (that is, the specific region IA).

[0213] For example, in S240, for each detection rectangle BB, the specifying unit 170 specifies the area AR where the detection rectangle BB exists as a specific area IA. The specifying unit 170 may specify, as the specific area IA, an area AR that "includes the position PS (foot position) of the detection target OB calculated from the detection rectangle BB (the position, shape, and size of the detection rectangle BB)". Further, the specifying unit 170 may specify an area AR including the center position (or centroid position) of the detection rectangle BB as the specific area IA.

[0214] Specifically, the specifying unit 170 refers to the storage unit 100 to obtain the rectangle information table 161 and the divided image information 130. The specifying unit 170 specifies, for each detection rectangle BB, a specific area IA that is "the area AR where the detection rectangle BB exists" from the obtained rectangle information table 161 and divided image information 130. The specifying unit 170 stores information for identifying the specific area IA specified for each detection rectangle BB in the specific area table 180 of the storage unit 100.

[0215] For example, in S250, the determination unit 190 may determine, for each detection rectangle BB (that is, for each detection target OB), whether the specific area IA of the detection rectangle BB matches "the area AR with the largest probability PR".

[0216] Specifically, the determination unit 190 refers to the specific area table 180 of the storage unit 100 to obtain the specific area IA for each detection rectangle BB. Further, the determination unit 190 refers to the probability table 162 of the storage unit 100 to obtain "the probability PR of each of the plurality of areas AR" for each detection rectangle BB (that is, for each detection target OB), and selects, for each detection rectangle BB, the area AR with the largest probability PR. Then, the determination unit 190 determines, for each detection rectangle BB (that is, for each detection target OB), whether the specific area IA of the detection rectangle BB matches "the area AR with the largest probability PR".

[0217] The determination unit 190 adopts the detection rectangle BB as a correct detection result on the grounds that "the specific area IA of the detection rectangle BB matches the area AR with the largest probability PR", that is, determines that "the estimation that 'the detection target OB surrounded by the detection rectangle BB is a person' is correct".

[0218] If the determination unit 190 determines that "the specific area IA of the detection rectangle BB does not match the area AR with the largest probability PR", it removes the detection rectangle BB as a false detection, that is, determines that "the estimation that 'the detection target OB surrounded by the detection rectangle BB is a person' is incorrect".

[0219] Although details will be described later, the determination unit 190 may determine whether the specific area IA matches the area AR with the largest probability PR or an area AR adjacent to the area AR with the largest probability PR.

[0220] Specifically, the determination unit 190 may select, for each detection rectangle BB (that is, for each detection target OB), the area AR with the largest probability PR and the area AR adjacent to the area AR with the largest probability PR from the "probability PR of each of the plurality of areas AR" for each detection rectangle BB. Then, the determination unit 190 may determine, for each detection rectangle BB (that is, for each detection target OB), whether the specific area IA of the detection rectangle BB matches the area AR with the largest probability PR or the area AR adjacent to the area AR with the largest probability PR.

[0221] In the first determination process, the determination unit 190 determines, for each detection target OB (that is, for each detection rectangle BB), whether the "class of the detection target OB" output by the estimation unit 150 (the learned model 140) from the captured image PI is correct, using the specific area IA of the detection rectangle BB.

[0222] (Output example of the learned model) FIG. 8 is a diagram showing an example of a class among the outputs of the learned model 140. Specifically, FIG. 8 is a diagram showing an example of a class (that is, a probability PR including a position PS where a detection target OB exists for each of a plurality of regions AR) output by the detection device 10 from the captured image PI using the learned model 140. In the example shown in FIG. 8, the detection device 10 (particularly, the division unit 120) divides the captured image PI into seven regions AR, namely, region AR(0), region AR(1), region AR(2), ···, region AR(6).

[0223] In the example shown in FIG. 8, two detection rectangles BB are arranged in the captured image PI. Specifically, a detection rectangle BB(1) and a detection rectangle BB(2) are arranged. The detection rectangle BB(1) and the detection rectangle BB(2) are each arranged so as to surround a detection target OB estimated to be a "person (human body).

[0224] Also, in the example shown in FIG. 8, the class of each of the two detection targets OB estimated to be "a person" is output (calculated) by the learned model 140. Specifically, the probability PR including the "position PS(1) where the detection target OB(1) surrounded by the detection rectangle BB(1) exists" is calculated for each of the seven regions AR from region AR(0) to region AR(6) by the learned model 140. Similarly, the probability PR including the "position PS(2) where the detection target OB(2) surrounded by the detection rectangle BB(2) exists" is calculated for each of the seven regions AR from region AR(0) to region AR(6) by the learned model 140.

[0225] In the example shown in FIG. 8, as the probability PR(0-1) which is the probability PR that "region AR(0) includes position PS(1)", "0.596" is calculated. As the probability PR(1-1) which is the probability PR that "region AR(1) includes position PS(1)", "0.357" is calculated. As the probability PR(2-1) which is the probability PR that "region AR(2) includes position PS(1)", "0.032" is calculated. As the probability PR(3-1) which is the probability PR that "region AR(3) includes position PS(1)", "0.004" is calculated. As the probability PR(4-1) which is the probability PR that "region AR(4) includes position PS(1)", "0.006" is calculated. As the probability PR(5-1) which is the probability PR that "region AR(5) includes position PS(1)", "0.002" is calculated. As the probability PR(6-1) which is the probability PR that "region AR(6) includes position PS(1)", "0.006" is calculated.

[0226] In the example shown in FIG. 8, as the probability PR(0-2) which is the probability PR that "region AR(0) includes position PS(2)", "0.642" is calculated. As the probability PR(1-2) which is the probability PR that "region AR(1) includes position PS(2)", "0.1" is calculated. As the probability PR(2-2) which is the probability PR that "region AR(2) includes position PS(2)", "0.179" is calculated. As the probability PR(3-2) which is the probability PR that "region AR(3) includes position PS(2)", "0.014" is calculated. As the probability PR(4-2) which is the probability PR that "region AR(4) includes position PS(2)", "0.013" is calculated. As the probability PR(5-2) which is the probability PR that "region AR(5) includes position PS(2)", "0.035" is calculated. As the probability PR(6-2) which is the probability PR that "region AR(6) includes position PS(2)", "0.017" is calculated.

[0227] (Regarding the specification of a specific region) FIG. 9 is a diagram showing an example of a specific area IA specified by the detection device 10 (particularly, the specific part 170). The detection device 10 (particularly, the division part 120) divides the captured image PI illustrated in FIG. 9 into seven areas AR from area AR(0) to area AR(6). Then, the detection device 10 (particularly, the specific part 170) specifies, for each detection rectangle BB arranged in the captured image PI, the area AR in which the detection rectangle BB exists as the specific area IA.

[0228] The specific part 170 may specify the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB, and specify the area AR including the position PS as the specific area IA of the detection rectangle BB. Specifically, the specific part 170 calculates the position PS (foot position) of the detection target OB from the detection rectangle BB (the position, shape, and size of the detection rectangle BB), and specifies the area AR including the calculated position PS as the specific area IA. The specific part 170 may specify the area AR including the center position (or the centroid position) of the detection rectangle BB as the specific area IA of the detection rectangle BB.

[0229] In the captured image PI illustrated in FIG. 9, two detection rectangles BB are arranged. Specifically, the detection device 10 (particularly, the estimation part 150 (the learned model 140)) arranges the detection rectangle BB(1) and the detection rectangle BB(2) with respect to the captured image PI illustrated in FIG. 9.

[0230] The specific part 170 specifies a specific area IA(1) which is an area AR where the detection rectangle BB(1) exists, and a specific area IA(2) which is an area AR where the detection rectangle BB(2) exists. For example, the specific part 170 specifies "the position PS(1) (foot position) of the detection target OB(1) surrounded by the detection rectangle BB(1)" and "the position PS(2) (foot position) of the detection target OB(2) surrounded by the detection rectangle BB(2)". Specifically, the specific part 170 calculates the position PS(1) from the detection rectangle BB(1), and also calculates the position PS(2) from the detection rectangle BB(2). For example, the specific part 170 specifies the center position (or centroid position) of the detection rectangle BB(1) as the position PS(1), and specifies the center position (or centroid position) of the detection rectangle BB(2) as the position PS(2). Then, the specific part 170 specifies the area AR including the position PS(1) (center position of the detection rectangle BB(1)) as the specific area IA(1), and specifies the area AR including the position PS(2) (center position of the detection rectangle BB(2)) as the specific area IA(2).

[0231] In the example of FIG. 9, the specific part 170 specifies that among the seven areas AR from the area AR(0) to the area AR(6), the area AR(1) includes the position PS(1) (center position of the detection rectangle BB(1)). Therefore, the specific part 170 specifies the area AR(1) as the specific area IA(1) of the detection rectangle BB(1).

[0232] Also, the specific part 170 specifies that among the seven areas AR from the area AR(0) to the area AR(6), the area AR(6) includes the position PS(2) (center position of the detection rectangle BB(2)). Therefore, the specific part 170 specifies the area AR(6) as the specific area IA(2) of the detection rectangle BB(2).

[0233] (Example of executing the first determination process and determining that the estimation result is correct) FIG. 10 is a diagram showing an example in which the detection device 10 (particularly, the determination unit 190) executes the first determination process and determines that the estimation result is correct. In FIG. 10, for the detection rectangle BB(1) among the two detection rectangles BB illustrated in FIG. 9, the determination unit 190 executes the first determination process and determines that the estimation that "the detection target OB(1) surrounded by the detection rectangle BB(1) is a person (human body)" is correct.

[0234] As described with reference to FIG. 9, the detection device 10 (particularly, the specifying unit 170) specifies that the specific area IA(1) of the detection rectangle BB(1) arranged in the captured image PI illustrated in FIG. 10 is the area AR(1).

[0235] In FIG. 10, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of "the detection target OB(1) surrounded by the detection rectangle BB(1)". That is, the learned model 140 calculates (outputs) the "probability PR including the position PS(1) where the detection target OB(1) exists" for each of the seven areas AR from the area AR(0) to the area AR(6).

[0236] In FIG. 10, the learned model 140 calculates that the probability PR(0-1) that the area AR(0) includes the "position PS(1) where the detection target OB(1) exists" is "0.3". The learned model 140 calculates that the probability PR(1-1) that the area AR(1) includes the "position PS(1) where the detection target OB(1) exists" is "0.7". The learned model 140 calculates that the probability PR(2-1) that the area AR(2) includes the "position PS(1) where the detection target OB(1) exists" is "0.05". The learned model 140 calculates that the probability PR(3-1) that the area AR(3) includes the "position PS(1) where the detection target OB(1) exists" is "0.05". The learned model 140 calculates that the probability PR(4-1) that the area AR(4) includes the "position PS(1) where the detection target OB(1) exists" is "0.02". The learned model 140 calculates that the probability PR(5-1) that the area AR(5) includes the "position PS(1) where the detection target OB(1) exists" is "0.075". The learned model 140 calculates that the probability PR(6-1) that the area AR(6) includes the "position PS(1) where the detection target OB(1) exists" is "0.075".

[0237] The determination unit 190 selects the area AR(1) as the area AR with the highest probability PR that includes the "position PS(1) where the detection target OB(1) exists" among the seven areas AR from the area AR(0) to the area AR(6). The determination unit 190 checks whether the "area AR(1) with the highest probability PR that includes the 'position PS(1) where the detection target OB(1) exists'" matches the "area AR(1) which is the specific area IA(1)". Then, when the determination unit 190 confirms that the area AR(1) with the highest probability PR matches the area AR(1) which is the specific area IA(1), it determines that the estimation that "the detection target OB(1) surrounded by the detection rectangle BB(1) is a person (human body)" is correct.

[0238] (Example of executing the first determination process and determining that the estimation result is incorrect) FIG. 11 is a diagram showing an example in which the detection device 10 (particularly, the determination unit 190) executes the first determination process and determines that the estimation result is incorrect. In FIG. 11, the determination unit 190 executes the first determination process for the detection rectangle BB(2) among the two detection rectangles BB illustrated in FIG. 9, and determines that the estimation that "the detection target OB(2) surrounded by the detection rectangle BB(2) is a person (human body)" is incorrect.

[0239] As described with reference to FIG. 9, the detection device 10 (particularly, the specifying unit 170) specifies that the specific area IA(2) of the detection rectangle BB(2) arranged in the captured image PI illustrated in FIG. 11 is the area AR(6).

[0240] In FIG. 11, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of "the detection target OB(2) surrounded by the detection rectangle BB(2)". That is, the learned model 140 calculates (outputs) the "probability PR including the position PS(2) where the detection target OB(2) exists" for each of the seven areas AR from the area AR(0) to the area AR(6).

[0241] In FIG. 11, the learned model 140 calculates that the probability PR(0-2) that the region AR(0) includes the position PS(2) where the detection target OB(2) exists is "0.1". The learned model 140 calculates that the probability PR(1-2) that the region AR(1) includes the position PS(2) where the detection target OB(2) exists is "0.2". The learned model 140 calculates that the probability PR(2-2) that the region AR(2) includes the position PS(2) where the detection target OB(2) exists is "0.3". The learned model 140 calculates that the probability PR(3-2) that the region AR(3) includes the position PS(2) where the detection target OB(2) exists is "0.1". The learned model 140 calculates that the probability PR(4-2) that the region AR(4) includes the position PS(2) where the detection target OB(2) exists is "0.05". The learned model 140 calculates that the probability PR(5-2) that the region AR(5) includes the position PS(2) where the detection target OB(2) exists is "0.05". The learned model 140 calculates that the probability PR(6-2) that the region AR(6) includes the position PS(2) where the detection target OB(2) exists is "0.2".

[0242] The determination unit 190 selects the region AR(2) as the region AR with the highest probability PR that includes the position PS(2) where the detection target OB(2) exists among the seven regions AR from the region AR(0) to the region AR(6). The determination unit 190 checks whether the region AR(2) with the highest probability PR that includes the position PS(2) where the detection target OB(2) exists (the region AR(2)) matches the region AR(6) which is the specific region IA(2). Then, when the determination unit 190 confirms that the region AR(2) with the highest probability PR does not match the region AR(6) which is the specific region IA(2), it determines that the estimation that "the detection target OB(2) surrounded by the detection rectangle BB(2) is a person (human body)" is incorrect.

[0243] As described with reference to FIGS. 10 and 11, in the detection device 10, when the specific region IA matches the region AR with the highest probability PR among the plurality of regions AR, the determination unit 190 determines that "the detection target OB is a person".

[0244] According to the above configuration, when the detection rectangle BB surrounding the detection target OB estimated to be a person exists in the area AR (i.e., the specific area IA) which is the "area AR among the plurality of areas AR where the probability PR is the highest", the detection device 10 determines that "the detection target OB is a person". That is, when the specific area IA coincides with the "area AR where the probability PR including the position PS where the detection target OB estimated to be a person exists is the highest", the detection device 10 determines that "the detection target OB is a person".

[0245] When the detection target OB estimated to be a person is actually a person, it is considered highly likely that the specific area IA which is the area AR where the detection rectangle BB surrounding the detection target OB exists coincides with the "area AR where the probability PR including the position PS where the detection target OB exists is the highest".

[0246] Therefore, when the specific area IA coincides with the "area AR among the plurality of areas AR where the probability PR is the highest", the detection device 10 determines that "the detection target OB is a person".

[0247] Therefore, the detection device 10 can achieve the effect of detecting a person with high accuracy from the image (captured image PI) captured using the fisheye lens.

[0248] When the specific part 170 specifies the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB and estimated to be "a person" as the specific area IA, the determination part 190 may determine whether "the foot position is included in the area AR where the probability PR is the highest". When the foot position is included in the area where the probability PR is the highest, the determination part 190 may determine that "the detection target OB is a person".

[0249] (Regarding a modification example of the first determination process) In the example described with reference to FIGS. 10 and 11, the determination unit 190 verified the correctness of the estimation that "the detection target OB surrounded by the detection rectangle BB is a person (human body)" by confirming the following two matches. That is, the determination unit 190 verified the correctness of the estimation based on the match / mismatch between the region AR with the highest probability PR including "the position PS where the detection target OB surrounded by the detection rectangle BB exists" and the specific region IA which is the region AR where the detection rectangle BB exists.

[0250] However, as illustrated in FIG. 12, when the detection target OB estimated to be "a person" exists near the boundary line of a plurality of regions AR, the region AR with the highest probability PR may not match the specific region IA even though the estimation is correct.

[0251] FIG. 12 is a diagram showing an example of the captured image PI in which the detection target OB estimated to be "a person" exists near the boundary line of a plurality of regions AR. The detection device 10 (particularly, the division unit 120) divides the captured image PI illustrated in FIG. 12 into seven regions AR from region AR(0) to region AR(6).

[0252] In FIG. 12, the estimation unit 150 (the learned model 140) places one detection rectangle BB surrounding the detection target OB estimated to be "a person (human body)" on the captured image PI divided into seven regions AR from region AR(0) to region AR(6). Specifically, the learned model 140 places the detection rectangle BB(3) on the captured image PI.

[0253] In FIG. 12, the detection device 10 (particularly, the specifying unit 170) specifies that the specific region IA(3), which is the region AR among the seven regions AR from region AR(0) to region AR(6) where the detection rectangle BB(3) exists, is region AR(2). For example, the specifying unit 170 calculates the position PS(3) (foot position) of the detection target OB(3) from the detection rectangle BB(3), and specifies the region AR(2), which is the region AR including the calculated position PS(3), as the specific region IA(3). The specifying unit 170 may specify the region AR(2), which is the region AR including the center position (or the centroid position) of the detection rectangle BB(3), as the specific region IA(3).

[0254] FIG. 13 is a diagram showing an example in which even when the region AR with the highest probability PR does not match the specific region IA, the detection device 10 executes the first determination process and determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person (human body)" is correct.

[0255] As described with reference to FIG. 12, the specifying unit 170 specifies that the specific region IA(3) of the detection rectangle BB(3) arranged in the captured image PI illustrated in FIG. 13 is region AR(2).

[0256] In FIG. 13, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of "the detection target OB(3) surrounded by the detection rectangle BB(3)". That is, the learned model 140 calculates (outputs) the "probability PR including the position PS(3) where the detection target OB(3) exists" for each of the seven regions AR from region AR(0) to region AR(6).

[0257] In FIG. 13, the learned model 140 calculates that the probability PR(0-3) that the region AR(0) includes the "position PS(3) where the detection target OB(3) surrounded by the detection rectangle BB(3) exists" is "0.2". The learned model 140 calculates that the probability PR(1-3) that the region AR(1) includes the "position PS(3) where the detection target OB(3) exists" is "0.45". The learned model 140 calculates that the probability PR(2-3) that the region AR(2) includes the "position PS(3) where the detection target OB(3) exists" is "0.4". The learned model 140 calculates that the probability PR(3-3) that the region AR(3) includes the "position PS(3) where the detection target OB(3) exists" is "0.05". The learned model 140 calculates that the probability PR(4-3) that the region AR(4) includes the "position PS(3) where the detection target OB(3) exists" is "0.02". The learned model 140 calculates that the probability PR(5-3) that the region AR(5) includes the "position PS(3) where the detection target OB(3) exists" is "0.075". The learned model 140 calculates that the probability PR(6-3) that the region AR(6) includes the "position PS(3) where the detection target OB(3) exists" is "0.075".

[0258] In FIG. 13, the region AR with the highest probability PR of including the "position PS(3) where the detection target OB(3) exists" is the region AR(1). The region AR with the second highest probability PR of including the "position PS(3) where the detection target OB(3) exists" is the region AR(2), and the region AR(2) is adjacent to the region AR(1). The probability PR(1-3) of the region AR(1) is "0.45", which is less than or equal to "0.5", and the probability PR(2-3) of the region AR(2) is "0.4", which is less than or equal to "0.5".

[0259] In FIG. 13, the determination unit 190 selects the region AR(1) as the region AR with the highest probability PR of including the "position PS(3) where the detection target OB(3) exists" among the seven regions AR from the region AR(0) to the region AR(6).

[0260] Then, the determination unit 190 confirms that the region AR(1) with the highest probability PR does not match the region AR(2) which is the specific region IA(3).

[0261] Therefore, in FIG. 13, the determination unit 190 further confirms that the region AR adjacent to the region AR(1) with the highest probability PR is the region AR(2).

[0262] Then, the determination unit 190 determines whether or not "the region AR(2) adjacent to the region AR(1) with the highest probability PR" matches "the region AR(2) which is the specific region IA(3)". When the determination unit 190 confirms that "the region AR(2) adjacent to the region AR(1) with the highest probability PR" matches "the region AR(2) which is the specific region IA(3)", it determines that the detection rectangle BB(3) is correct. That is, the determination unit 190 determines that the estimation that "the detection object OB(3) surrounded by the detection rectangle BB(3) is a person (human body)" is correct.

[0263] As described with reference to FIG. 13, when the specific region IA matches "the region AR with the highest probability PR" or "the region AR adjacent to the region AR with the highest probability PR", the determination unit 190 may determine that "the detection object OB surrounded by the detection rectangle BB" is a person.

[0264] In the example shown in FIG. 13, the region AR(1) and the region AR(2) are adjacent to each other. Among the seven regions AR from the region AR(0) to the region AR(6), the region AR with the highest probability PR including "the position PS(3) where the detection object OB(3) exists" is the region AR(1). Also, the position PS(3) where the detection object OB(3) surrounded by the detection rectangle BB(3) and estimated to be "a person (human body)" exists is near the boundary line between the region AR(1) and the region AR(2).

[0265] In such a case, the determination unit 190 determines whether the specific region IA(3) of the detection rectangle BB(3) matches with “the region AR(1) with the highest probability PR” or “the region AR(2) adjacent to the region AR(1)”. When the specific region IA(3) of the detection rectangle BB(3) matches with “the region AR(1) with the highest probability PR” or “the region AR(2) adjacent to the region AR(1)”, the determination unit 190 determines that the estimation that “the detection target OB(3) is a person” is correct.

[0266] For example, in the example shown in FIG. 13, the determination unit 190 determines whether the detection rectangle BB(3) is included in the portion indicated by the dashed-dotted line in FIG. 13. When the determination unit 190 confirms that “the detection rectangle BB(3) is included in the portion indicated by the dashed-dotted line in FIG. 13”, the determination unit 190 determines that the estimation that “the detection target OB(3) is a person” is correct.

[0267] The determination unit 190 may determine whether the position PS(3) where the detection target OB(3) surrounded by the detection rectangle BB(3) and estimated to be “a person (human body)” exists is included in the portion indicated by the dashed-dotted line in FIG. 13. Specifically, the determination unit 190 may determine whether “the position PS(3) calculated from the detection rectangle BB(3) by the specific unit 170” is included in the portion indicated by the dashed-dotted line in FIG. 13. The determination unit 190 may use the center position (or the centroid position) of the detection rectangle BB(3) as the position PS(3) and determine whether “the position PS(3) is included in the portion indicated by the dashed-dotted line in FIG. 13”. When the determination unit 190 confirms that “the position PS(3) is included in the portion indicated by the dashed-dotted line in FIG. 13”, the determination unit 190 may determine that the estimation that “the detection target OB(3) is a person” is correct.

[0268] In S250 of FIG. 7, the determination unit 190 determined whether the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB was included in the output class (that is, the region AR with the highest probability PR). However, in S250, the determination unit 190 may also determine whether the position PS (foot position) of the detection target OB surrounded by the detection rectangle BB is included in any of the following two regions AR. That is, the determination unit 190 may determine whether the position PS of the detection target OB is included in the region AR with the highest probability PR or the position adjacent to the region AR with the highest probability PR. In other words, the determination unit 190 may determine whether the specific region IA of the detection rectangle BB matches the region AR with the highest probability PR or the region AR adjacent to the region AR with the highest probability PR.

[0269] In the portion indicated by the dashed-dotted line in FIG. 13, in addition to the region AR(1), the following two areas are included. That is, when the region AR(2) adjacent to the region AR(1) is divided into a plurality of areas, the area adjacent to the region AR(1), and when the region AR(6) adjacent to the region AR(1) is divided into a plurality of areas, the area adjacent to the region AR(1) are included.

[0270] As described so far, for example, when the detection rectangle BB is arranged across the adjacent regions AR(x) and AR(y), the determination unit 190 uses the specific region IA of the detection rectangle BB to determine the correctness of the class as follows. That is, the determination unit 190 determines whether the specific region IA of the detection rectangle BB (that is, the region AR(x) or the region AR(y)) matches the region AR with the highest probability PR or the region AR adjacent to the region AR with the highest probability PR.

[0271] In other words, when the position PS of the detection target OB calculated from the detection rectangle BB is near the boundary line between the adjacent regions AR(x) and AR(y), the determination unit 190 determines the correctness of the class as follows using the specific region IA of the detection rectangle BB. That is, the determination unit 190 determines whether the specific region IA of the detection rectangle BB (that is, the region AR(x) or the region AR(y)) matches the "region AR with the highest probability PR" or the "region AR adjacent to the region AR with the highest probability PR". The "near the boundary line" may be the buffer region BA described later with reference to FIG. 18. The position PS is calculated from the detection rectangle BB (the position, shape, and size of the detection rectangle BB), and may be, for example, the center position (or the centroid position) of the detection rectangle BB.

[0272] When the detection rectangle BB is arranged across the adjacent regions AR(x) and AR(y), the specific region IA of the detection rectangle BB is considered to be the region AR(x) or the region AR(y). Also, the "region AR with the highest probability PR" is considered to be the region AR(x) or the region AR(y), and the "region AR adjacent to the region AR with the highest probability PR" is considered to be the region AR(y) or the region AR(x).

[0273] The process of the determination unit 190 described with reference to FIG. 13 can be organized as follows. That is, for example, when the specific region IA matches the "region AR with the highest probability PR among the plurality of regions AR" or the "region AR adjacent to the region AR with the highest probability PR among the plurality of regions AR", the determination unit 190 determines that the detection target OB is a person.

[0274] According to the above configuration, when the specific region IA matches the "region AR with the highest probability PR among the plurality of regions AR" or the "region AR adjacent to the region AR with the highest probability PR among the plurality of regions AR", the detection device 10 determines that the "detection target OB is a person".

[0275] Here, in the captured image PI, situations are assumed where a person is captured so as to straddle two of the plurality of regions AR, or the position where the person is captured is near the boundary between the two regions AR.

[0276] Under such circumstances, when the detection target OB estimated to be a person is actually a person, the specific region IA, which is the region AR where the detection rectangle BB surrounding the detection target OB exists, is likely to match any of the following regions AR. That is, the specific region IA is likely to match "the region AR among the plurality of regions AR with the highest probability PR" or "the region AR adjacent to the region AR among the plurality of regions AR with the highest probability PR".

[0277] Therefore, when the specific region IA matches "the region AR among the plurality of regions AR with the highest probability PR", the detection device 10 determines that "the detection target OB is a person". Also, when the specific region IA matches "the region AR adjacent to the region AR among the plurality of regions AR with the highest probability PR", the detection device 10 determines that "the detection target OB is a person".

[0278] Accordingly, the detection device 10 can effectively detect a person with high accuracy from the captured image PI even when a person is captured near the boundary between two of the plurality of regions AR in the captured image PI.

[0279] When the specific part 170 specifies the foot position as the specific region IA, the determination part 190 may determine whether the foot position is included in "the region AR with the highest probability PR" or "the region AR adjacent to the region AR with the highest probability PR". The determination part 190 may determine that "the detection target OB is a person" when the foot position is included in "the region AR with the highest probability PR" or "the region AR adjacent to the region AR with the highest probability PR".

[0280] (Example of human detection process including second determination process) FIG. 14 is a flowchart for explaining an example of a person detection process that executes a second determination process on the output of the learned model 140 to detect a person. The processes of S310 and S320 of the second determination process illustrated in FIG. 14 are the same as the processes of S210 and S220 of the first determination process illustrated in FIG. 7.

[0281] That is, the image acquisition unit 110 acquires the captured image PI, and the division unit 120 divides the captured image PI acquired by the image acquisition unit 110 into a plurality of regions AR (S310). The estimation unit 150 performs person detection and class output on the captured image PI (the captured image PI divided into a plurality of regions AR) acquired with reference to the divided image information 130 using the learned model 140 (S320).

[0282] The detection device 10 executes the processes from S330 to S370 for each detection rectangle BB (that is, for each detection target OB), that is, for the number of people detected by the estimation unit 150 from the captured image PI, in other words, for the number of detection rectangles BB arranged by the estimation unit 150.

[0283] The determination unit 190 refers to the object reliability table 163 and the probability table 162 to acquire the object reliability OR and the probability PR of the class to which the detection rectangle BB (that is, the detection target OB) belongs (that is, the region AR corresponding to the specific region IA) for each detection rectangle BB (S330).

[0284] Specifically, the determination unit 190 refers to the object reliability table 163 to acquire the object reliability OR for each detection rectangle BB. Further, the determination unit 190 refers to the specific region table 180 to acquire the specific region IA for each detection rectangle BB. Then, the determination unit 190 refers to the probability table 162 to acquire the probability PR of the region AR corresponding to the acquired specific region IA.

[0285] For each detection rectangle BB, the determination unit 190 calculates the class confidence value CR of the detection rectangle BB by multiplying the object confidence OR of the detection rectangle BB by the probability PR of the region AR corresponding to the specific region IA of the detection rectangle BB (S340). Then, the determination unit 190 determines whether the class confidence value CR of the detection rectangle BB is greater than a predetermined value TH (S350).

[0286] If the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH (YES in S350), the determination unit 190 adopts the detection rectangle BB as a correct detection result (S360). That is, the determination unit 190 determines that if the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH, the estimation that the detection target OB surrounded by the detection rectangle BB is a person is correct.

[0287] If the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH (NO in S350), the determination unit 190 removes the detection rectangle BB as a false detection (S370). That is, the determination unit 190 determines that if the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH, the estimation that the detection target OB surrounded by the detection rectangle BB is a person is incorrect.

[0288] (Example of executing the second determination process and determining that the estimation result is correct) FIG. 15 is a diagram showing an example in which the detection device 10 (particularly, the determination unit 190) executes the second determination process and determines that the estimation result is correct. In FIG. 15, for the detection rectangle BB(1) arranged in the captured image PI illustrated in FIG. 15, the determination unit 190 executes the second determination process and determines that the estimation that the detection target OB(1) surrounded by the detection rectangle BB(1) is a person (human body) is correct.

[0289] As described with reference to FIG. 9, the detection device 10 (particularly, the specifying unit 170) specifies that the specific region IA(1) of the detection rectangle BB(1) arranged in the captured image PI illustrated in FIG. 15 is the region AR(1).

[0290] In FIG. 15, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of the "detection target OB(1) surrounded by the detection rectangle BB(1)". That is, the learned model 140 calculates (outputs) the "probability PR including the position PS(1) where the detection target OB(1) exists" for each of the seven regions AR from the region AR(0) to the region AR(6).

[0291] In FIG. 15, the learned model 140 calculates that the probability PR(0-1) that the region AR(0) includes the "position PS(1) where the detection target OB(1) exists" is "0". The learned model 140 calculates that the probability PR(1-1) that the region AR(1) includes the "position PS(1) where the detection target OB(1) exists" is "0.4". The learned model 140 calculates that the probability PR(2-1) that the region AR(2) includes the "position PS(1) where the detection target OB(1) exists" is "0.005". The learned model 140 calculates that the probability PR(3-1) that the region AR(3) includes the "position PS(1) where the detection target OB(1) exists" is "0.45". The learned model 140 calculates that the probability PR(4-1) that the region AR(4) includes the "position PS(1) where the detection target OB(1) exists" is "0.005". The learned model 140 calculates that the probability PR(5-1) that the region AR(5) includes the "position PS(1) where the detection target OB(1) exists" is "0.14". The learned model 140 calculates that the probability PR(6-1) that the region AR(6) includes the "position PS(1) where the detection target OB(1) exists" is "0".

[0292] Also, in FIG. 15, the learned model 140 calculates the object reliability OR(1) of the detection rectangle BB(1) to be "920".

[0293] The determination unit 190 calculates the class confidence value CR(1) of the detection rectangle BB(1) by multiplying the probability PR(1-1) of the region AR(1) corresponding to the "specific region IA(1) specified for the detection rectangle BB(1)" by the object reliability OR(1) of the detection rectangle BB(1). That is, the determination unit 190 multiplies the "probability PR(1-1) of the region AR(1): 0.4" by the "object reliability OR(1) of the detection rectangle BB(1): 920" to calculate the "class confidence value CR(1) of the detection rectangle BB(1): 368".

[0294] The determination unit 190 checks whether the class confidence value CR of the detection rectangle BB is greater than a predetermined value TH. When it is confirmed that the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH, the determination unit 190 determines that the estimation that "the detection object OB surrounded by the detection rectangle BB is a person" is correct. For example, when the determination unit 190 confirms that the "class confidence value CR(1) of the detection rectangle BB(1): 368" is greater than the "predetermined value TH: 350", the determination unit 190 determines that the estimation that "the detection object OB(1) surrounded by the detection rectangle BB(1) is a person" is correct.

[0295] (Example of executing the second determination process and determining that the estimation result is incorrect) FIG. 16 is a diagram showing an example in which the detection device 10 (particularly, the determination unit 190) executes the second determination process and determines that the estimation result is incorrect. In FIG. 16, the determination unit 190 executes the second determination process for the detection rectangle BB(2) arranged in the captured image PI illustrated in FIG. 16, and determines that the estimation that "the detection object OB(2) surrounded by the detection rectangle BB(2) is a person (human body)" is incorrect.

[0296] As described with reference to FIG. 9, the detection device 10 (particularly, the specifying unit 170) specifies that the specific region IA(2) of the detection rectangle BB(2) arranged in the captured image PI illustrated in FIG. 16 is the region AR(6).

[0297] In FIG. 16, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of the “detection target OB(2) surrounded by the detection rectangle BB(2)”. That is, the learned model 140 calculates (outputs) the “probability PR including the position PS(2) where the detection target OB(2) exists” for each of the seven regions AR from the region AR(0) to the region AR(6).

[0298] In FIG. 16, the learned model 140 calculates that the probability PR(0-2) that the region AR(0) includes the “position PS(2) where the detection target OB(2) exists” is “0.1”. The learned model 140 calculates that the probability PR(1-2) that the region AR(1) includes the “position PS(2) where the detection target OB(2) exists” is “0.2”. The learned model 140 calculates that the probability PR(2-2) that the region AR(2) includes the “position PS(2) where the detection target OB(2) exists” is “0.3”. The learned model 140 calculates that the probability PR(3-2) that the region AR(3) includes the “position PS(2) where the detection target OB(2) exists” is “0.1”. The learned model 140 calculates that the probability PR(4-2) that the region AR(4) includes the “position PS(2) where the detection target OB(2) exists” is “0.05”. The learned model 140 calculates that the probability PR(5-2) that the region AR(5) includes the “position PS(2) where the detection target OB(2) exists” is “0.05”. The learned model 140 calculates that the probability PR(6-2) that the region AR(6) includes the “position PS(2) where the detection target OB(2) exists” is “0.2”.

[0299] Also, in FIG. 16, the learned model 140 calculates the object reliability OR(2) of the detection rectangle BB(2) to be “625”.

[0300] The determination unit 190 calculates the class confidence value CR(2) of the detection rectangle BB(2) by multiplying the probability PR(6-2) of the region AR(6) corresponding to the "specific region IA(2) specified for the detection rectangle BB(2)" by the object reliability OR(2) of the detection rectangle BB(2). That is, the determination unit 190 multiplies the "probability PR(6-2) of the region AR(6): 0.2" by the "object reliability OR(2) of the detection rectangle BB(2): 625" to calculate the "class confidence value CR(2) of the detection rectangle BB(2): 125".

[0301] The determination unit 190 checks whether the class confidence value CR of the detection rectangle BB is greater than a predetermined value TH. When it is confirmed that the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH, the determination unit 190 determines that the estimation that "the detection object OB surrounded by the detection rectangle BB is a person" is incorrect. For example, when the determination unit 190 confirms that the "class confidence value CR(2) of the detection rectangle BB(2): 125" is less than or equal to the "predetermined value TH: 350", the determination unit 190 determines that the estimation that "the detection object OB(2) surrounded by the detection rectangle BB(2) is a person" is incorrect.

[0302] The processing of the determination unit 190 described with reference to FIGS. 15 and 16 can be organized as follows. That is, for example, when the class confidence value CR calculated by multiplying the probability PR of the region AR corresponding to the specific region IA among the plurality of regions AR by the object reliability OR of the detection rectangle BB is greater than the predetermined value TH, the determination unit 190 determines that "the detection object OB is a person". The object reliability OR of the detection rectangle BB is a value indicating the likelihood that some object is surrounded by the detection rectangle BB.

[0303] According to the above configuration, when the class confidence value CR calculated by multiplying the probability PR of the region AR corresponding to the specific region IA among the plurality of regions AR by the object reliability OR of the detection rectangle BB is greater than the predetermined value TH, the detection device 10 determines that "the detection object OB is a person".

[0304] When the detection target OB estimated to be a person is actually a person, the probability PR that "the specific area IA, which is the area AR where the detection rectangle BB surrounding the detection target OB exists, includes the position PS where the detection target OB exists" is considered to be sufficiently high. Also, when the detection target OB estimated to be a person is actually a person, the object reliability OR of the detection rectangle BB, which is a value indicating the likelihood that some object is surrounded by the "detection rectangle BB surrounding the detection target OB", is also considered to be a sufficiently high value.

[0305] Therefore, when the class reliability value CR calculated by multiplying the probability PR of the area AR corresponding to the specific area IA among the plurality of areas AR by the object reliability OR of the detection rectangle BB is greater than the predetermined value TH, the detection device 10 determines that "the detection target OB is a person".

[0306] Therefore, the detection device 10 has the effect of being able to detect a person with high accuracy from the image (captured image PI) captured using the fisheye lens.

[0307] When the specific part 170 specifies the foot position as the specific area IA, the determination part 190 may determine that "the detection target OB is a person" when the class reliability value CR calculated by multiplying the probability PR of the area AR including the foot position by the object reliability OR is greater than the predetermined value TH.

[0308] (Regarding a modification example of the second determination process) In the example described with reference to FIGS. 15 and 16, the determination part 190 calculates the class reliability value CR of the detection rectangle BB by multiplying the probability PR of the area AR corresponding to the "specific area IA specified for the detection rectangle BB" by the object reliability OR of the detection rectangle BB. Then, the determination part 190 determines whether the calculated class reliability value CR is greater than the predetermined value TH. If the calculated class reliability value CR is greater than the predetermined value TH, the determination part 190 determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct. If the calculated class reliability value CR is less than or equal to the predetermined value TH, the determination part 190 determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is incorrect.

[0309] However, as illustrated in FIG. 17, when the detection target OB to be estimated as "being a person" exists near the boundary line of a plurality of regions AR, even though the estimation is correct, the class confidence value CR of the detection rectangle BB may be equal to or less than a predetermined value TH because the probability PR is small. That is, when the detection rectangle BB straddles two of the plurality of regions AR, even though the detection target OB surrounded by the detection rectangle BB is a person, the class confidence value CR of the detection rectangle BB may be equal to or less than a predetermined value TH because the probability PR is small.

[0310] FIG. 17 is a diagram showing an example in which the detection device 10 (particularly, the determination unit 190) executes a second determination process and determines that the estimation result is correct even when the detection rectangle BB (in other words, the position PS where the detection target OB exists) is near the boundary line of a plurality of regions AR. In the example shown in FIG. 17, the detection rectangle BB(3) is arranged so as to straddle the adjacent regions AR(1) and AR(2). In other words, the position PS(3) of the detection target OB(3) surrounded by the detection rectangle BB(3) is near the boundary line between the regions AR(1) and AR(2). The position PS(3) is calculated from the detection rectangle BB(3) and may be, for example, the center position or the centroid position of the detection rectangle BB(3).

[0311] In FIG. 17, the detection device 10 (particularly, the division unit 120) divides the captured image PI into seven regions AR from region AR(0) to region AR(6). Also, in FIG. 17, the estimation unit 150 (the learned model 140) places one detection rectangle BB that surrounds the detection target OB estimated to be "a person (human body)" with respect to the captured image PI divided into seven regions AR from region AR(0) to region AR(6). Specifically, the learned model 140 places the detection rectangle BB(3) with respect to the captured image PI. Further, in FIG. 17, the detection device 10 (particularly, the specification unit 170) specifies that the specific region IA(3), which is the region AR where the detection rectangle BB(3) exists among the seven regions AR from region AR(0) to region AR(6), is region AR(2). The specification unit 170 calculates the position PS(3) (foot position) of the detection target OB(3) surrounded by the detection rectangle BB(3) from the detection rectangle BB(3). Then, the specification unit 170 specifies the region AR(2), which is the region AR including the calculated position PS(3), as the specific region IA(3). For example, the specification unit 170 may specify the region AR(2), which is the region AR including the center position (or the centroid position) of the detection rectangle BB(3), as the specific region IA(3).

[0312] In FIG. 17, the detection device 10 (particularly, the estimation unit 150 (the learned model 140)) outputs the class of "the detection target OB(3) surrounded by the detection rectangle BB(3)". That is, the learned model 140 calculates (outputs) the "probability PR including the position PS(3) where the detection target OB(3) exists" for each of the seven regions AR from region AR(0) to region AR(6).

[0313] In FIG. 17, the learned model 140 calculates that the probability PR(0-3) that the region AR(0) includes the position PS(3) where the detection target OB(3) exists is "0.3". The learned model 140 calculates that the probability PR(1-3) that the region AR(1) includes the position PS(3) where the detection target OB(3) exists is "0.3". The learned model 140 calculates that the probability PR(2-3) that the region AR(2) includes the position PS(3) where the detection target OB(3) exists is "0.45". The learned model 140 calculates that the probability PR(3-3) that the region AR(3) includes the position PS(3) where the detection target OB(3) exists is "0.05". The learned model 140 calculates that the probability PR(4-3) that the region AR(4) includes the position PS(3) where the detection target OB(3) exists is "0.02". The learned model 140 calculates that the probability PR(5-3) that the region AR(5) includes the position PS(3) where the detection target OB(3) exists is "0.075". The learned model 140 calculates that the probability PR(6-3) that the region AR(6) includes the position PS(3) where the detection target OB(3) exists is "0.075".

[0314] Also, in FIG. 17, the learned model 140 calculates the object reliability OR(3) of the detection rectangle BB(3) as "800".

[0315] The determination unit 190 calculates the average value of the probability PR(2-3) of the region AR(2) which is the specified specific region IA(3) and the probability PR(1-3) of the region AR(1) adjacent to the region AR(2). That is, the determination unit 190 calculates "0.375" which is the average value of "probability PR(2-3) of region AR(2): 0.45" and "probability PR(1-3) of region AR(1): 0.3".

[0316] The determination unit 190 multiplies the calculated average value by the object reliability OR(3) of the detection rectangle BB(3) to calculate the class reliability value CR(3) of the detection rectangle BB(3). That is, the determination unit 190 multiplies "average value: 0.375" by "object reliability OR(3): 800" to calculate "class reliability value CR(3): 300".

[0317] The determination unit 190 checks whether the class confidence value CR of the detection rectangle BB is greater than a predetermined value TH. When it is confirmed that the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH, the determination unit 190 determines that the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct. For example, when the determination unit 190 confirms that "the class confidence value CR(3) of the detection rectangle BB(3): 300" is less than or equal to "the predetermined value TH: 250", the determination unit 190 determines that the estimation that "the detection target OB(3) surrounded by the detection rectangle BB(3) is a person" is correct.

[0318] In the example shown in FIG. 17, the detection rectangle BB(3) straddles the region AR(2) which is the specific region IA(3) and the region AR(1) adjacent to the region AR(2). That is, the detection target OB(3) exists near the boundary line between the region AR(2) and the region AR(1). In other words, the position PS(3) of the detection target OB(3) calculated from the detection rectangle BB(3) is near the boundary line between the region AR(2) which is the specific region IA(3) and the region AR(1) adjacent to the region AR(2). For example, the position PS(3) may be the center position or the centroid position of the detection rectangle BB(3).

[0319] Therefore, the determination unit 190 multiplies the average value of the probability PR(2 - 3) of the region AR(2) and the probability PR(1 - 3) of the region AR(1) by the object confidence level OR(3) of the detection rectangle BB(3) to calculate the class confidence value CR(3) of the detection rectangle BB(3). Then, the determination unit 190 determines the correctness of the estimation that "the detection target OB(3) surrounded by the detection rectangle BB(3) is a person" according to whether the class confidence value CR(3) of the detection rectangle BB(3) is greater than the predetermined value TH.

[0320] When the detection rectangle BB (that is, the position PS of the detection target OB surrounded by the detection rectangle BB) exists near the boundary line between the region AR(x) which is the specific region IA of the detection rectangle BB and the region AR(y) adjacent to the region AR(x), the determination unit 190 executes the following processing.

[0321] That is, the determination unit 190 calculates the probability PR(x) that "the area AR(x) includes the position PS of the detection target OB surrounded by the detection rectangle BB" and the probability PR(y) that "the area AR(y) includes the position PS of the detection target OB surrounded by the detection rectangle BB". The determination unit 190 multiplies the average value of the probability PR(x) and the probability PR(y) by the object reliability OR of the detection rectangle BB to calculate the class reliability value CR of the detection rectangle BB. Then, the determination unit 190 determines whether the estimation that "the detection target OB surrounded by the detection rectangle BB is a person" is correct or not according to whether the class reliability value CR of the detection rectangle BB is greater than a predetermined value TH.

[0322] In other words, in S330 illustrated in FIG. 14, the determination unit 190 calculates the average value of the following two probabilities PR of the area AR instead of "the probability PR of the class to which the detection rectangle BB belongs (that is, the probability PR of the area AR that is the specific area IA)". That is, the determination unit 190 calculates the average value of the probability PR(x) of the area AR(x) that is the specific area IA and the probability PR(y) of the area AR(y) adjacent to the area AR(x). Then, in S340 illustrated in FIG. 14, the determination unit 190 multiplies the "average value of the probability PR(x) and the probability PR(y)" instead of "the probability PR of the area AR that is the specific area IA" by the object reliability OR of the detection rectangle BB to calculate the class reliability value CR.

[0323] In the examples shown in FIGS. 15 and 16, the detection rectangle BB (that is, the position PS of the detection target OB surrounded by the detection rectangle BB) exists at a position away from the boundary lines of the plurality of areas AR, and the predetermined value TH was "350". On the other hand, in the example shown in FIG. 17, the detection rectangle BB (that is, the position PS of the detection target OB surrounded by the detection rectangle BB) is near the boundary lines of the plurality of areas AR, and the predetermined value TH is "250".

[0324] Thus, the determination unit 190 may change the value of the predetermined value TH according to "whether the detection rectangle BB (that is, the position PS of the detection target OB surrounded by the detection rectangle BB) exists near the boundary line of the plurality of regions AR". Further, the determination unit 190 may keep the value of the predetermined value TH constant regardless of "whether the detection rectangle BB (that is, the position PS of the detection target OB surrounded by the detection rectangle BB) is near the boundary line of the plurality of regions AR".

[0325] (Example of setting a buffer area) FIG. 18 is a diagram showing an example in which the detection device 10 (for example, the division unit 120) sets a buffer area BA near the boundary line of a plurality of regions AR for the captured image PI. As shown in FIG. 18, the division unit 120 may set a buffer area BA near the boundary line of the plurality of regions AR. Specifically, the division unit 120 may set a buffer area BA(x - y) near the boundary line between the region AR(x) and the region AR(y).

[0326] In the example shown in FIG. 18, the division unit 120 sets a buffer area BA(1 - 2) near the boundary line between the region AR(1) and the region AR(2). The division unit 120 sets a buffer area BA(2 - 3) near the boundary line between the region AR(2) and the region AR(3). The division unit 120 sets a buffer area BA(3 - 4) near the boundary line between the region AR(3) and the region AR(4). The division unit 120 sets a buffer area BA(4 - 5) near the boundary line between the region AR(4) and the region AR(5). The division unit 120 sets a buffer area BA(5 - 6) near the boundary line between the region AR(5) and the region AR(6). The division unit 120 sets a buffer area BA(6 - 1) near the boundary line between the region AR(6) and the region AR(1).

[0327] When the position PS of the detection target OB surrounded by the detection rectangle BB is included in the buffer area BA, the detection device 10 executes the process illustrated in FIG. 19 to determine whether the estimation that "the detection target OB is a person" is correct.

[0328] FIG. 19 is a flowchart for explaining an example of a person detection process that executes a second determination process and detects a person when a buffer area BA is set. The processes of S410, 420, and from S470 to S490 of the second determination process illustrated in FIG. 19 are the same as the processes of S310, S320, and from S350 to S370 of the second determination process illustrated in FIG. 14.

[0329] That is, the image acquisition unit 110 acquires the captured image PI, and the division unit 120 divides the captured image PI acquired by the image acquisition unit 110 into a plurality of regions AR (S410). The estimation unit 150 performs person detection and class output on the captured image PI (the captured image PI divided into a plurality of regions AR) acquired with reference to the divided image information 130 using the learned model 140 (S420).

[0330] The detection device 10 executes the processes from S430 to S490 for each detection rectangle BB, that is, for each detection target OB, by the number of people detected by the estimation unit 150 from the captured image PI, in other words, by the number of detection rectangles BB arranged by the estimation unit 150.

[0331] The determination unit 190 determines for each detection rectangle BB whether "the detection rectangle BB belongs to the inside of the buffer area BA (inside the buffer area BA)", that is, determines whether "the detection rectangle BB exists inside the buffer area BA". The determination unit 190 may determine for each detection rectangle BB whether "the position PS of the detection target OB surrounded by the detection rectangle BB" belongs to the inside of the buffer area BA. The position PS of the detection target OB is calculated from the detection rectangle BB (the position, shape, and size of the detection rectangle BB), and may be, for example, the center position (or the centroid position) of the detection rectangle BB.

[0332] If the detection rectangle BB belongs to the buffer area BA (YES in S430), the determination unit 190 calculates, as the probability PR (class probability) of the detection rectangle BB, the average value of the probabilities PR of the following two regions AR adjacent to each other in the buffer area BA. That is, the determination unit 190 calculates the average value of the probability PR of the region AR corresponding to the specific region IA and the probability PR of the region AR adjacent to the region AR corresponding to the specific region IA.

[0333] For example, when the detection rectangle BB is within the buffer area BA(x - y) of "the region AR(x) which is the specific region IA of the detection rectangle BB" and "the region AR(y) adjacent to the region AR(x)", the determination unit 190 sets the probability PR (class probability) of the detection rectangle BB to the following value. That is, the determination unit 190 sets the probability PR of the detection rectangle BB to the average value of "the probability PR(x) of the region AR(x)" and "the probability PR(y) of the region AR(y)".

[0334] If the detection rectangle BB does not belong to the buffer area BA (NO in S430), the determination unit 190 sets the probability PR (class probability) of the detection rectangle BB to the probability PR of the region AR to which the detection rectangle BB belongs, that is, the probability PR of "the region AR which is the specific region IA of the detection rectangle BB".

[0335] For example, when the detection rectangle BB belongs only to "the region AR which is the specific region IA of the detection rectangle BB", the determination unit 190 sets the probability PR (class probability) of the detection rectangle BB to the probability PR of "the region AR which is the specific region IA of the detection rectangle BB".

[0336] The determination unit 190 may determine whether "the detection rectangle BB belongs to the buffer area BA" based on whether the buffer area BA includes "the position PS of the detection target OB surrounded by the detection rectangle BB (the position PS calculated from the detection rectangle BB)". For example, the determination unit 190 may determine whether "the detection rectangle BB belongs to the buffer area BA" based on whether the buffer area BA includes the center position (or the centroid position) of the detection rectangle BB.

[0337] For each detection rectangle BB, the determination unit 190 calculates the class confidence value CR of the detection rectangle BB by multiplying the object confidence OR of the detection rectangle BB by the "class probability of the detection rectangle BB" (S460). Then, the determination unit 190 determines whether "the class confidence value CR of the detection rectangle BB is greater than a predetermined value TH" (S470).

[0338] If "the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH" (YES in S470), the determination unit 190 adopts the detection rectangle BB as a correct detection result (S480). That is, the determination unit 190 determines that "the class confidence value CR of the detection rectangle BB is greater than the predetermined value TH" and that "the estimation that the detection object OB surrounded by the detection rectangle BB is a person" is correct.

[0339] If "the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH" (NO in S470), the determination unit 190 removes the detection rectangle BB as a false detection (S490). That is, the determination unit 190 determines that "the class confidence value CR of the detection rectangle BB is less than or equal to the predetermined value TH" and that "the estimation that the detection object OB surrounded by the detection rectangle BB is a person" is incorrect.

[0340] As described with reference to FIGS. 17 to 19, the determination unit 190 may determine that "the detection object OB is a person" when the class confidence value CR calculated by multiplying the average value of the probabilities PR of the following two regions AR by the object confidence OR of the detection rectangle BB is greater than the predetermined value TH. That is, the determination unit 190 may calculate the class confidence value CR from the average value of the probability PR of the "region AR corresponding to the specific region IA" and the probability PR of the "region AR adjacent to the region AR corresponding to the specific region IA" among the plurality of regions AR.

[0341] According to the above configuration, the detection device 10 calculates the average value of the probability PR of "the region AR corresponding to the specific region IA" among the plurality of regions AR and the probability PR of "the region AR adjacent to the region AR corresponding to the specific region IA" among the plurality of regions AR. Then, when the class confidence value CR calculated by multiplying the calculated average value by the object confidence level OR of the detection rectangle BB is greater than a predetermined value TH, the detection device 10 determines that "the detection target OB is a person".

[0342] Here, in the captured image PI, situations are assumed where a person is captured so as to straddle two of the plurality of regions AR, or the position where the person is captured is near the boundary of the two regions AR.

[0343] Under such circumstances, when the detection target OB estimated to be a person is actually a person, it is considered that both the probability PR of "the region AR corresponding to the specific region IA" and the probability PR of "the region AR adjacent to the region AR corresponding to the specific region IA" become sufficiently high values. Similarly, when the detection target OB estimated to be a person is actually a person, the object confidence level OR of the detection rectangle BB, which is a value indicating the likelihood that some object is surrounded by the "detection rectangle BB surrounding the detection target OB", is also considered to be a sufficiently high value.

[0344] Therefore, the detection device 10 calculates the average value of the probability PR of "the region AR corresponding to the specific region IA" and the probability PR of "the region AR adjacent to the region AR corresponding to the specific region IA". Then, when the class confidence value CR of the detection rectangle BB calculated by multiplying the average value by the object confidence level OR of the detection rectangle BB is greater than a predetermined value TH, the detection device 10 determines that "the detection target OB is a person".

[0345] Therefore, the detection device 10 has the effect of being able to detect a person with high accuracy from the captured image PI even when a person is captured near the boundary of two of the plurality of regions AR in the captured image PI.

[0346] When the specific part 170 specifies the foot position as the specific area IA, the determination unit 190 may calculate the average value of the probability PR of the "area AR including the foot position" and the probability PR of the "area AR adjacent to the area AR including the foot position". Then, when the class confidence value CR of the detection rectangle BB calculated by multiplying the calculated average value by the object confidence level OR of the detection rectangle BB is greater than a predetermined value TH, the determination unit 190 may determine that "the detection target OB is a person".

[0347] (Summary of the control method executed by the detection device 10) The control method executed by the detection device 10, which has been described so far with reference to FIGS. 7 to 19, can be summarized as follows. That is, the control method executed by the detection device 10 is a control method of a detection device that detects a person being imaged from an imaging image PI captured by a ceiling camera using a fish-eye lens. The control method includes a division step, a region estimation step (estimation step), and a determination step.

[0348] In the division step, the imaging image PI is divided into a plurality of regions AR. In the region estimation step (estimation step), for each of the plurality of regions AR, the probability PR including the "position PS where the detection target OB estimated to be a person exists" is calculated. In the determination step, it is determined whether the "detection target OB is a person" using the probability PR of each of the plurality of regions AR.

[0349] For example, the division step corresponds to S210 in FIG. 7, S310 in FIG. 14, and S410 in FIG. 19. The region estimation step (estimation step) corresponds to S220 in FIG. 7, S320 in FIG. 14, and S420 in FIG. 19. The determination step corresponds to S250 in FIG. 7, S350 in FIG. 14, and S470 in FIG. 19.

[0350] According to the above configuration, the control method divides the captured image PI into a plurality of regions AR, and for each of the plurality of regions AR, calculates a probability PR including "the position PS where the detection target OB estimated to be a person exists". Then, the control method verifies the estimation result that "the detection target OB is a person" using the probability PR of each of the plurality of regions AR.

[0351] Here, generally, when trying to detect a person from an image using a dictionary showing human characteristics, analysis of the image is required for each dictionary. Therefore, when trying to improve the detection accuracy of a person using a plurality of dictionaries showing human characteristics, analysis of the image is also required multiple times, and the time required to detect a person becomes long.

[0352] In contrast, the control method calculates, for each of the plurality of regions AR obtained by dividing the captured image PI, "the probability PR including 'the position PS where the detection target OB estimated to be a person exists'". Then, the control method determines whether "the detection target OB is a person" using the calculated probability PR of each of the plurality of regions AR, thereby improving the detection accuracy of a person. That is, the control method does not detect a person from the captured image PI using a plurality of dictionaries showing human characteristics, but determines whether the detection target OB estimated to be a person is actually a person by using the probability PR of each of the plurality of regions AR, thereby improving the detection accuracy of a person.

[0353] While the method of using a plurality of dictionaries showing human characteristics improves the accuracy of the estimation itself, the control method improves the detection accuracy by verifying the estimated result (the estimation that the detection target OB is a person) (that is, removing an incorrect estimation result).

[0354] Therefore, the control method does not need to use a plurality of dictionaries showing human characteristics to detect a person from the captured image PI, and can shorten the time required to detect a person from the captured image PI compared to the case of detecting a person from the captured image PI using a plurality of dictionaries.

[0355] In addition, since the control method determines (verifies) whether the detection target OB estimated to be a person is actually a person by using the probability PR of each of the plurality of regions AR, a person can be detected with high accuracy from the captured image PI.

[0356] Therefore, the control method has the effect of being able to detect a person from an image (captured image PI) captured using a fisheye lens quickly and with high accuracy.

[0357] Also, as described above, the control method does not need to use a plurality of dictionaries indicating human features to detect a person from the captured image PI.

[0358] Therefore, the control method has the effect of reducing the labor required to prepare a dictionary (for example, a learned model) necessary for detecting a person from an image captured using a fisheye lens, and also reducing the memory capacity required to store the dictionary.

[0359] §4. Modification Example For the detection device 10, it is not essential to execute the learning process. That is, it is not essential that the entity that generates the learned model 140 is the detection device 10. For the detection device 10 that detects "a person captured in the captured image PI" from the captured image PI, it is sufficient to be able to detect "a person captured in the captured image PI" from the captured image PI using the learned model 140. Generating the learned model 140 is not essential. In other words, it is not essential for the detection device 10 to include the teacher data generation unit 210 and the learning unit 220.

[0360] A learned model generation device different from the detection device 10, which includes the teacher data generation unit 210 and the learning unit 220, may generate the learned model 140. Then, the detection device 10 may detect "a person captured in the captured image PI" from the captured image PI using the learned model 140 generated by the learned model generation device.

[0361] The specific unit 170 may calculate the position PS of the object OB to be detected from the detection rectangle BB, for example, from the position, shape, and size of the detection rectangle BB.

[0362] 〔Example of implementation by software〕 The functional blocks of the detection device 10 (specifically, the image acquisition unit 110, the segmentation unit 120, the estimation unit 150, the specific unit 170, the determination unit 190, the teacher data generation unit 210, and the learning unit 220) may be realized by a logic circuit (hardware) formed in an integrated circuit (IC chip) or the like. Further, these functional blocks may be realized by software using a CPU, GPU, DSP, or the like.

[0363] In the latter case, the detection device 10 includes a CPU, GPU, DSP, or the like that executes instructions of a program that is software for realizing each function, a ROM (Read Only Memory) or a storage device (collectively referred to as a "recording medium") in which the program and various data are recordable in a readable manner by a computer (or a CPU, GPU, DSP, etc.), and a RAM (Random Access Memory) or the like that expands the program. Then, when the computer (or a CPU, GPU, DSP, etc.) reads the program from the recording medium and executes it, the object of the present invention is achieved. As the recording medium, a "non-transitory tangible medium", such as a tape, a disk, a card, a semiconductor memory, a programmable logic circuit, or the like, can be used. Further, the program may be supplied to the computer via any transmission medium (such as a communication network or a broadcast wave) capable of transmitting the program. Note that the present invention can also be realized in the form of a data signal embedded in a carrier wave, in which the program is embodied by electronic transmission.

[0364] The present invention is not limited to the above-described embodiments, and various modifications are possible within the scope shown in the claims. Embodiments obtained by appropriately combining the technical means disclosed in different embodiments are also included in the technical scope of the present invention.

Explanation of reference numerals

[0365] 10 Detection device 20 Ceiling camera 120 Division unit 140 Learned model (area prediction model) 142 Area prediction model 150 Estimation unit (area estimation unit) 152 Area estimation unit 170 Identification unit 190 Judgment unit 220 Learning unit 222 Area learning unit (learning unit) AR area BB Detection rectangle (bounding box) CR Class confidence value DT Teacher data IA Specific area OB Detection target OR Object reliability PI Captured image PR Probability PS Position S120 Acquisition step S130 Learning step S210 Division step S220 Area estimation step S250 Judgment step S310 Division step S320 Area estimation step S350 Judgment step S410 Division step S420 Area estimation step S470 Judgment step TH Predetermined value

Claims

1. A detection device that detects a person being imaged from an imaging image captured by a ceiling camera using a fish-eye lens, comprising: a dividing unit that divides the imaging image into a plurality of regions; a region estimation unit that calculates the probability that the position where a detection target estimated to be a person exists belongs to each of the plurality of regions; a determination unit that determines whether the detection target is a person using the probabilities of each of the plurality of regions; and comprising: among the plurality of regions, a specifying unit that specifies a region in which a bounding box surrounding the detection target exists as a specific region; the determination unit determines whether the detection target is a person using the probabilities of each of the plurality of regions and the specific region; the determination unit determines that the detection target is a person when the specific region matches the region having the highest probability among the plurality of regions. A detection device.

2. The determination unit determines that the detection target is a person when the specific region matches the region having the highest probability among the plurality of regions or a region adjacent to the region having the highest probability among the plurality of regions. The detection device according to Claim 1. The detection device according to claim 1.

3. A detection device that detects a person being imaged from an imaging image captured by a ceiling camera using a fish-eye lens, comprising: a dividing unit that divides the imaging image into a plurality of regions; a region estimation unit that calculates the probability that the position where a detection target estimated to be a person exists belongs to each of the plurality of regions; a determination unit that determines whether the detection target is a person using the probabilities of each of the plurality of regions; and comprising: among the plurality of regions, a specifying unit that specifies a region in which a bounding box surrounding the detection target exists as a specific region; the determination unit determines whether the detection target is a person using the probabilities of each of the plurality of regions and the specific region; the determination unit determines that the detection target is a person when the specific region matches the region having the highest probability among the plurality of regions or a region adjacent to the region having the highest probability among the plurality of regions. A detection device.

4. A detection device that detects a person being imaged from an imaging image captured by a ceiling camera using a fish-eye lens, comprising: a dividing unit that divides the imaging image into a plurality of regions; An area estimation unit that calculates the probability that the position where a detection target estimated to be a person exists belongs to each of the plurality of areas; A determination unit that determines whether the detection target is a person using the probability of each of the plurality of areas; Comprising; Among the plurality of areas, further comprising a specifying unit that specifies, as a specific area, an area in which a bounding box surrounding the detection target exists; The determination unit determines whether the detection target is a person using the probability of each of the plurality of areas and the specific area; The determination unit multiplies the probability of the area corresponding to the specific area among the plurality of areas by an object reliability value, which is a value indicating the likelihood that some object is surrounded by the bounding box, and calculates the class reliability value. If the calculated class reliability value is greater than a predetermined value, the detection device determines that the detection target is a person.

5. The determination unit multiplies the average value of the probability of the area corresponding to the specific area among the plurality of areas and the probability of the area adjacent to the area corresponding to the specific area among the plurality of areas by an object reliability value, which is a value indicating the likelihood that some object is surrounded by the bounding box, and calculates the class reliability value. If the calculated class reliability value is greater than a predetermined value, the determination unit determines that the detection target is a person. The detection device according to claim 4.

6. A detection device that detects a person captured in a captured image captured by a ceiling camera using a fisheye lens, A division unit that divides the captured image into a plurality of areas; An area estimation unit that calculates the probability that the position where a detection target estimated to be a person exists belongs to each of the plurality of areas; A determination unit that determines whether the detection target is a person using the probability of each of the plurality of areas; Comprising; Among the plurality of areas, further comprising a specifying unit that specifies, as a specific area, an area in which a bounding box surrounding the detection target exists; The determination unit determines whether the detection target is a person using the probability of each of the plurality of areas and the specific area; The determination unit multiplies the average value of the probability of the region corresponding to the specific region among the plurality of regions and the probability of the region adjacent to the region corresponding to the specific region among the plurality of regions by an object reliability value indicating the likelihood that some object is surrounded by the bounding box, and if the class reliability value calculated thereby is greater than a predetermined value, determines that the detection target is a person.

7. The region estimation unit uses a region prediction model, which is a learned model that takes the captured image as input and outputs the probability of the presence of the detection target in each of the plurality of regions, to calculate the probability of each of the plurality of regions from the captured image. The detection device according to any one of claims 1 to 6.

8. The apparatus further includes a learning unit that constructs, by machine learning on teacher data in which information indicating a region including the position where a person captured in the captured image exists is given as a label with respect to the captured image, a region prediction model, which is a learned model that takes the captured image as input and outputs the probability of the presence of the detection target in each of the plurality of regions. The detection device according to any one of claims 1 to 7.

9. A control method for a detection device that detects a person captured in a captured image taken by a ceiling camera using a fish-eye lens, the method including: a dividing step of dividing the captured image into a plurality of regions; a region estimating step of calculating the probability that the position where the detection target estimated to be a person exists belongs to each of the plurality of regions; a determining step of determining whether the detection target is a person using the probability of each of the plurality of regions; and including: a specifying step of specifying, as a specific region, a region among the plurality of regions in which a bounding box surrounding the detection target exists; in the determining step, determining whether the detection target is a person using the probability of each of the plurality of regions and the specific region; and in the determining step, determining that the detection target is a person if the specific region matches the region having the highest probability among the plurality of regions.

10. A model generation method by a model generation device that generates a learned model, For a captured image divided into a plurality of regions, an acquisition step of acquiring teacher data to which (A) information indicating the position, shape, and size of a bounding box surrounding a person captured in the captured image and (B) information identifying a region including the position where the person captured in the captured image exists are attached as labels. A learning step of constructing a learned model that takes the captured image as an input and outputs (C) information indicating the position, shape, and size of the bounding box and (D) the probability of each of the plurality of regions including the position where the person captured in the captured image exists, by machine learning on the teacher data. A model generation method including the above.

11. An information processing program for causing a computer to function as the detection device according to any one of Claims 1, 3, 4, and 6, the information processing program for causing a computer to function as the division unit, the region estimation unit, the determination unit, and the specification unit.

12. An information processing program for causing a computer to execute the model generation method according to Claim 10, the information processing program for causing a computer to execute the acquisition step and the learning step.

13. A computer-readable recording medium on which the information processing program according to Claim 11 or 12 is recorded.

Citation Information

Patent Citations

  • Learning method, learning device, image recognition method, image recognition device and program

    JP2016099668A

  • Image sensor, person detection method, control system, control method, and computer program

    JP2016171526A

  • Person detection device and person detection method

    WO2020137160A1