Recognition processing apparatus, recognition processing method, and person detection model generation method

JP2024072345A5Active Publication Date: 2025-10-07JVC KENWOOD CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2022183073
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-11-16
Publication Date
2025-10-07
Estimated Expiration
2042-11-16

AI Technical Summary

Technical Problem

Existing image recognition systems struggle to stably detect individuals in crowded environments where multiple people overlap, leading to inconsistent detection results due to varying overlap patterns.

Method used

A recognition processing device that employs two detection models: one trained on single-person images and another on superimposed images, using position and map information to determine the appropriate model for the situation, and a stability determination mechanism to ensure consistent detection.

Benefits of technology

Enhances the accuracy and stability of person detection in crowded scenarios by adaptively selecting the most effective detection model based on environmental conditions and detection stability, reducing inconsistent results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To detect a person more appropriately in image recognition processing.SOLUTION: A recognition processing apparatus 10 includes: a video acquisition unit 20 which acquires a video captured by a camera 40 equipped in a mobile body; a person detection unit 14 which detects a person included in the acquired video using at least one of a first detection model trained by machine learning using a single person image as a ground truth image, and a second detection model trained by machine learning using a superimposed person image as a ground truth image; a condition determination unit 34 which determines whether a specific condition has occurred, the condition being estimated to be crowded with people, using position information indicating a position in which the acquired video was captured and map information indicating a location estimated to be crowded with people; and a valid model determination unit 36 which determines whether to validate the detection of the person using the first detection model or the second detection model, based on a result determined by the condition determination unit 34.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to a recognition processing device, a recognition processing method, and a person detection model generating method. [Background technology]

[0002] There is a known technology that detects objects such as pedestrians from images captured around a vehicle using image recognition technology such as pattern matching. For example, a technology has been proposed that improves detection accuracy by preparing multiple recognition dictionaries, including those for distant objects and those for nearby objects, and performing pattern matching using the multiple recognition dictionaries (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2022-17871 A Summary of the Invention [Problem to be solved by the invention]

[0004] When moving through a busy place such as a school route or a shopping street, the image may include multiple people who appear to overlap in the direction of the camera's image capture. When multiple people are captured walking and moving, the overlapping manner of the multiple people included in the image may change over time. When trying to detect such multiple people, depending on the overlapping manner of the multiple people, it may be impossible to detect the people stably over time.

[0005] The present invention has been made in consideration of the above circumstances, and has an object to provide a technique for more appropriately detecting a person in image recognition processing. [Means for solving the problem]

[0006] A recognition processing device of one embodiment of the present invention includes an image acquisition unit that acquires an image captured by a camera mounted on a moving body, a person detection unit that detects a person included in the acquired image using at least one of a first detection model trained by machine learning on a single person image as a correct image and a second detection model trained by machine learning on a superimposed person image as a correct image, a situation determination unit that determines whether or not a specific situation is present in which the number of people is estimated to be high, using location information indicating the capture location of the acquired image and map information indicating locations where the number of people is estimated to be high, and an effective model determination unit that determines whether the detection of people using the first detection model or the second detection model is effective, based on the determination result of the situation determination unit.

[0007] Another aspect of the present invention is a recognition processing method, which includes the steps of: acquiring a video captured by a camera installed on a moving object, detecting a person included in the acquired video using at least one of a first detection model trained by machine learning on a single person image as a correct answer image and a second detection model trained by machine learning on a superimposed person image as a correct answer image, determining whether or not the situation is a specific situation estimated to be high in foot traffic using location information indicating the imaging location of the acquired video and map information indicating a place estimated to be high in foot traffic, and determining whether or not to use the first detection model or the second detection model to detect a person based on the result of the determination on whether or not the situation is a specific situation.

[0008] Yet another aspect of the present invention is a method for generating a person detection model by machine learning, which generates a person detection model using a superimposed person image, which includes a full-body image of a person and includes at least a part of another person as a background of the person, as a ground-truth image, and does not use a single person image, which includes a full-body image of a person and does not include another person as a background of the person, as a ground-truth image. Effect of the Invention

[0009] According to the present invention, it is possible to provide a technique for more appropriately detecting people in image recognition processing. [Brief description of the drawings]

[0010] [Figure 1] 1 is a block diagram illustrating a functional configuration of a recognition processing device according to a first embodiment. [Diagram 2] 2(a) to (d) are diagrams showing examples of single person images. [Diagram 3] 3(a) to (d) are diagrams showing examples of superimposed human images. [Figure 4] FIG. 11 is a top view showing a schematic diagram of a case where a specific point is included in a partial area of ​​the angle of view of the camera. [Diagram 5] FIG. 1 is a diagram showing an example of an image including a first range that is not a specific situation and a second range that is a specific situation. [Figure 6] FIG. 13 is a diagram showing an example of a display image to which a person detection result has been added. [Figure 7] 5 is a flowchart showing an example of the flow of a recognition processing method according to the first embodiment. [Figure 8] FIG. 11 is a block diagram illustrating a functional configuration of a recognition processing device according to a second embodiment. [Figure 9] FIG. 13 is a diagram showing an example of a partial range of an image that is a target for determining stability. [Figure 10] 10 is a flowchart showing an example of the flow of a recognition processing method according to a second embodiment. [Figure 11] 13 is a flowchart showing another example of the flow of the recognition processing method according to the second embodiment. [Figure 12] FIG. 11 is a block diagram illustrating a functional configuration of a recognition processing device according to a third embodiment. [Figure 13] 13 is a flowchart showing an example of the flow of a recognition processing method according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0011] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Specific numerical values ​​and the like shown in the embodiment are merely examples for facilitating understanding of the invention, and do not limit the present invention unless otherwise specified. In the drawings, elements that are not directly related to the present invention are omitted.

[0012] (First embodiment) FIG. 1 is a block diagram showing a schematic functional configuration of a recognition processing device 10 according to a first embodiment. The recognition processing device 10 includes an acquisition unit 12, a person detection unit 14, and a determination unit 16. The recognition processing device 10 may further include a display control unit 18. The recognition processing device 10 is mounted on a moving object such as a vehicle, and detects people such as pedestrians around the vehicle. In this embodiment, a case where the recognition processing device 10 is mounted on a vehicle will be illustrated. The recognition processing device 10 may be mounted on an air vehicle such as a drone.

[0013] Each functional block shown in this embodiment can be realized, for example, by cooperation of hardware and software. The hardware of the recognition processing device 10 is realized by elements and mechanical devices such as a processor such as a central processing unit (CPU) or a graphics processing unit (GPU) of a computer, and memories such as a read only memory (ROM) or a random access memory (RAM). The software of the recognition processing device 10 is realized by a computer program or the like.

[0014] The acquisition unit 12 includes an image acquisition unit 20. The image acquisition unit 20 acquires an image captured by a camera 40. The camera 40 is mounted on a moving body and captures an image of the surroundings of the moving body. The camera 40 captures, for example, an image in front of the moving body. The camera 40 may capture an image behind the moving body, or may capture an image to the side of the moving body. The recognition processing device 10 may or may not include the camera 40.

[0015] The camera 40 is configured to capture infrared rays. The camera 40 is a so-called infrared thermography, and images the temperature distribution around the moving body, so that a heat source existing around the moving body can be identified. The camera 40 may be configured to detect mid-infrared rays with a wavelength of about 2 μm to 5 μm, or may be configured to detect far-infrared rays with a wavelength of about 8 μm to 14 μm. The camera 40 may be configured to capture visible light. The camera 40 may be configured to capture color images of red, green, and blue, or may be configured to capture monochrome images of visible light. In this embodiment, the camera 40 will be described as a camera that captures thermal images using far-infrared rays. The video captured by the camera 40 is a moving image, for example, at 30 frames per second.

[0016] The acquisition unit 12 may include a position information acquisition unit 22. The position information acquisition unit 22 acquires position information measured by a position sensor 42. The position sensor 42 is mounted on a moving object and measures the position of the moving object. The position sensor 42 is, for example, a Global Navigation Satellite System (GNSS) sensor. The position sensor 42 detects the imaging position of the camera 40. The recognition processing device 10 may or may not include the position sensor 42.

[0017] The acquisition unit 12 may include a map information acquisition unit 24. The map information acquisition unit 24 acquires map information from a map device 44. The map device 44 is a device that stores map information, and is, for example, a navigation device. The map information includes information on a specific point indicating a place where it is estimated that there is a lot of foot traffic. The place where it is estimated that there is a lot of foot traffic is, for example, around a station, a school route, a shopping district, around a commercial facility, around a tourist spot, etc. The map information may include information on a specific date and time that is a date and time when it is estimated that there is a lot of foot traffic. The map information may include information on a specific condition that indicates a combination of a place where it is estimated that there is a lot of foot traffic and a date and time. The recognition processing device 10 may or may not include the map device 44. The map information acquisition unit 24 may acquire map information from an external server or the like using a wireless communication function not shown.

[0018] The acquisition unit 12 may include a time information acquisition unit 26. The time information acquisition unit 26 acquires time information from a clock device 46. The clock device 46 is, for example, a clock device that generates current time information indicating the current date and time. The clock device 46 outputs the image capture date and time of the camera 40. The recognition processing device 10 may or may not include the clock device 46.

[0019] The acquisition unit 12 may include an orientation information acquisition unit 28. The orientation information acquisition unit 28 acquires orientation information measured by an orientation sensor 48. The orientation sensor 48 is mounted on a moving object and measures the orientation of the moving object. The orientation sensor 48 is, for example, an acceleration sensor or a gyro sensor, and detects the direction or orientation of the moving object. The orientation sensor 48 detects, for example, the imaging direction of the camera 40. The recognition processing device 10 may or may not include the orientation sensor 48.

[0020] The person detection unit 14 detects an area including a person in the video acquired by the video acquisition unit 20. The person detection unit 14 includes a first detection unit 30 that detects a person using a first detection model, and a second detection unit 32 that detects a person using a second detection model. The person detection unit 14 may be configured to operate the first detection unit 30 and the second detection unit 32 in parallel, or may be configured to selectively operate only one of the first detection unit 30 or the second detection unit 32. The person detection unit 14 may selectively operate only one of the first detection unit 30 or the second detection unit 32 by switching the detection model to be used.

[0021] The first detection unit 30 detects a person using a first detection model generated by machine learning using a single person image as a correct answer image. The single person image is an image including a full-body image of a person, and does not include another person in the background of the person.

[0022] 2(a) to (d) are diagrams showing examples of single person images. Each of the images in FIG. 2(a), (b), (c), and (d) includes a full-body image of each of the persons 52a, 52b, 52c, and 52d. Each of FIG. 2(a), (b), (c), and (d) does not include any person other than the persons 52a, 52b, 52c, and 52d. The single person image is generated, for example, by cutting out an area including the full-body image of each of the persons 52a to 52d. The single person image is cut out, for example, so as to be a vertically long rectangular image with a vertical and horizontal image size ratio of 2:1.

[0023] The second detection unit 32 detects a person using a second detection model generated by machine learning using the superimposed person image as a correct answer image. The superimposed person image is an image including a whole-body image of a person, and including at least a part (e.g., head, upper body, lower body, arms, legs) of another person in the background of the person. The superimposed person image differs from a single person image in that multiple people appear to be superimposed on each other.

[0024] 3(a)-(d) are diagrams showing examples of superimposed person images. Each of the images in FIG. 3(a), (b), (c), and (d) includes first persons 54a, 54b, 54c, and 54d shown by dashed lines, and second persons 56a, 56b, 56c, and 56d different from the first persons. The first persons 54a-54d are persons seen in the foreground, and their whole bodies are visible. The second persons 56a-56d are persons located on the back side of the first persons 54a-54d. At least a part of the whole body of the second persons 56a-56d is hidden by the first persons 54a-54d and cannot be seen. The superimposed person images are generated, for example, by cutting out an area including the whole body image of the first persons 54a-54d. The superimposed person images are cut out, for example, to be vertically long rectangular images with a vertical and horizontal image size of 2:1. In the examples of FIGS. 3(a) to (d), the superimposed person image includes only two people, but the superimposed person image may include three or more people.

[0025] The model used for machine learning may include an input corresponding to the image size (number of pixels) of the input image, an output that outputs a recognition score, and an intermediate layer that connects the input and the output. The intermediate layer may include a convolutional layer, a pooling layer, a fully connected layer, etc. The intermediate layer may have a multi-layer structure, and may be configured to enable so-called deep learning. The model used for machine learning may be constructed using a convolutional neural network (CNN). Note that the model used for machine learning is not limited to the above, and any machine learning model may be used.

[0026] The first detection model is generated using a single person image, and therefore has a high accuracy in detecting a person existing alone in a place with little foot traffic. The first detection model tends to have a low accuracy in detecting a person in a situation where multiple people appear to overlap, such as in a place with a lot of foot traffic. On the other hand, the second detection model is generated using an overlapping person image, and therefore has a high accuracy in detecting a person that appears in the foreground in a situation where multiple people appear to overlap, such as in a place with a lot of foot traffic. The second detection model tends to have a low accuracy in detecting a person existing alone in a place with little foot traffic.

[0027] The first detection model may be generated by machine learning that does not use a superimposed person image as a ground truth image, and the second detection model may be generated by machine learning that does not use a single person image as a ground truth image.

[0028] Returning to FIG. 1, the determination unit 16 includes a situation determination unit 34. The situation determination unit 34 determines whether or not the situation is estimated to be one in which there is a large number of people (also referred to as a specific situation). The situation determination unit 34 determines whether or not the situation is a specific situation using position information indicating the imaging position of the video acquired by the video acquisition unit 20. The situation determination unit 34 determines whether or not the situation is a specific situation using, for example, position information acquired by the position information acquisition unit 22.

[0029] The situation determination unit 34 may further use the map information acquired by the map information acquisition unit 24 to determine whether or not the situation is a specific situation. The situation determination unit 34 may determine that the situation is a specific situation when the image capture position of the video matches a place (i.e., a specific point) included in the map information that is estimated to have a lot of people. The situation determination unit 34 may determine that the situation is not a specific situation when the image capture position of the video does not match a specific point.

[0030] The situation determination unit 34 may further use the time information acquired by the time information acquisition unit 26 to determine whether or not it is a specific situation. The situation determination unit 34 may determine that it is a specific situation when the image capture position and image capture date and time match a combination of a place and date and time that is estimated to be a place with a lot of people, which is included in the map information (i.e., a specific condition). The situation determination unit 34 may determine that it is not a specific situation when the image capture position and image capture date and time do not match the specific condition.

[0031] The situation determination unit 34 may further use the orientation information acquired by the orientation information acquisition unit 28 to determine whether or not the situation is a specific situation. The situation determination unit 34 may identify a location included in the angle of view of the camera 40 from the image capture position and image capture direction of the image. That is, the situation determination unit 34 may identify a location included in the image from the image capture position and image capture direction of the image. The situation determination unit 34 may determine that the situation is a specific situation when the location included in the image matches a location (i.e., a specific point) that is estimated to have a large number of people and is included in the map information. The situation determination unit 34 may determine that the situation is not a specific situation when the location included in the angle of view of the camera 40 does not match a specific point.

[0032] The situation determination unit 34 may determine whether or not a specific situation exists by using any combination of location information, map information, time information, and direction information. The situation determination unit 34 may determine that a specific situation exists when a location and an image capture date and time included in an image match a specific condition. The situation determination unit 34 may determine that a specific situation does not exist when a location and an image capture date and time included in an image do not match a specific condition.

[0033] The situation determination unit 34 may determine whether or not the specific situation exists for the entire range of the acquired image, or may determine whether or not the specific situation exists for a partial range of the acquired image. For example, when a specific point is included in a first range of the image and a specific point is not included in a second range of the image, the situation determination unit 34 may determine that the first range is a specific situation and that the second range is not a specific situation.

[0034] Fig. 4 is a top view that typically shows a case where a specific location is included in a portion of the angle of view 62 of camera 40. In Fig. 4, a moving object 68 is moving near the boundary of a place 60 that is estimated to be heavily trafficked. In Fig. 4, a first range 64 that corresponds to the right side of the angle of view 62 of camera 40 does not match the specific location or specific conditions, and a second range 66 that corresponds to the left side of the angle of view 62 of camera 40 matches the specific location or specific conditions. In this case, the situation determination unit 34 determines that the first range 64 is not a specific situation, and determines that the second range 66 is a specific situation.

[0035] Fig. 5 is a diagram showing an example of a video including a first range 64 that is not a specific situation and a second range 66 that is a specific situation. In the first range 64 on the right side of Fig. 5, there are office buildings, and the number of people present in the first range 64 is relatively small. On the other hand, in the second range 66 on the left side of Fig. 5, there is a shopping street, and the number of people present in the second range 66 is relatively large.

[0036] 1, the effective model determination unit 36 ​​determines whether the detection of a person using the first detection model or the second detection model is to be valid, based on the determination result of the situation determination unit 34. In other words, the effective model determination unit 36 ​​determines whether the detection result of the first detection unit 30 or the second detection unit 32 is to be valid.

[0037] When the situation determination unit 34 determines that the situation is not a specific situation, the effective model determination unit 36 ​​enables the detection of a person using the first detection model. In other words, when the situation determination unit 34 determines that the situation is not a specific situation, the effective model determination unit 36 ​​enables the detection of a person by the first detection unit 30. When the situation determination unit 34 determines that the situation is not a specific situation, the effective model determination unit 36 ​​may disable the detection of a person using the second detection model (i.e., the detection of a person by the second detection unit 32).

[0038] When the situation determination unit 34 determines that the situation is a specific situation, the effective model determination unit 36 ​​enables the detection of a person using the second detection model. In other words, when the situation determination unit 34 determines that the situation is a specific situation, the effective model determination unit 36 ​​enables the detection of a person by the second detection unit 32. When the situation determination unit 34 determines that the situation is a specific situation, the effective model determination unit 36 ​​may disable the detection of a person using the first detection model (i.e., the detection of a person by the first detection unit 30).

[0039] When it is determined whether a certain range of the image is a specific situation, the effective model determination unit 36 ​​may determine whether the first detection model or the second detection model is to be effective for the certain range of the image. For example, the effective model determination unit 36 ​​may enable the detection of a person using the first detection model (i.e., the detection of a person by the first detection unit 30) for a first range of the image determined not to be a specific situation, and disable the detection of a person using the second detection model (i.e., the detection of a person by the second detection unit 32). For example, the effective model determination unit 36 ​​may enable the detection of a person using the second detection model (i.e., the detection of a person by the second detection unit 32) for a second range of the image determined to be a specific situation, and disable the detection of a person using the first detection model (i.e., the detection of a person by the first detection unit 30).

[0040] The person detection unit 14 may detect a person included in the video using a detection model that is made valid by the valid model determination unit 36. The person detection unit 14 may not operate a detection model that is made invalid by the valid model determination unit 36. When the situation determination unit 34 makes the first detection model valid, the person detection unit 14 may operate only the first detection unit 30 and stop the function of the second detection unit 32. When the situation determination unit 34 makes the second detection model valid, the person detection unit 14 may operate only the second detection unit 32 and stop the function of the first detection unit 30. The person detection unit 14 may operate the first detection unit 30 and the second detection unit 32 in parallel, regardless of the determination result of the situation determination unit 34. In this case, the first detection unit 30 and the second detection unit 32 may each detect a person included in the same frame of the acquired video.

[0041] The display control unit 18 generates a display image by adding the person detection result by the person detection unit 14 to the image acquired by the image acquisition unit 20, and causes the generated display image to be displayed on the display device 50. The display device 50 is provided on the moving object. The display device 50 includes an image display element such as a liquid crystal display (LCD) or an organic electroluminescence display (OLED). For example, when the moving object is a vehicle, the display device 50 is disposed at a position where it can be seen by the driver of the vehicle. The recognition processing device 10 may or may not include the display device 50.

[0042] The display control unit 18 generates a display image by superimposing an additional image, such as a frame image for indicating an area including a person detected by the person detection unit 14, on the image. The display control unit 18 adds a first additional image to the person detected by the first detection unit 30, and adds a second additional image to the person detected by the second detection unit 32. The display mode of the first additional image may be the same as the display mode of the second additional image. The display mode of the first additional image may be different from the display mode of the second additional image. For example, the first additional image may be a yellow frame, and the second additional image may be a red frame.

[0043] The display control unit 18 adds an additional image to a person detected by a model determined to be valid by the valid model determination unit 36. The display control unit 18 does not add an additional image to a person detected by a model determined to be invalid by the valid model determination unit 36. When no person is detected by the person detection unit 14, the display control unit 18 uses the acquired image as it is as an image to be displayed, and causes the display device 50 to display the acquired image as it is.

[0044] Fig. 6 is a diagram showing an example of a display image to which a person detection result is added. Fig. 6 is a display image generated by the display control unit 18 when the image shown in Fig. 5 is acquired. In a first range 64 that is not a specific situation, a first additional image 70 that is a frame image showing a person detected by the first detection unit 30 is superimposed. In a second range 66 that is a specific situation, second additional images 72a, 72b, 72c, 72d, and 72e that are frame images showing a person detected by the second detection unit 32 are superimposed.

[0045] FIG. 7 is a flowchart showing an example of the flow of the recognition processing method according to the first embodiment. The video acquisition unit 20 acquires a video captured by the camera 40 (step S10). The position information acquisition unit 22 acquires position information indicating the image capture position of the video from the position sensor 42 (step S12). The situation determination unit 34 determines whether or not the situation is a specific situation based on the image capture position (step S14). If the situation determination unit 34 determines that the situation is a specific situation (Yes in step S14), the effective model determination unit 36 ​​activates the second detection model, and the person detection unit 14 detects a person included in the video using the second detection model trained by machine learning using the superimposed person image as a correct answer image (step S16). If the situation determination unit 34 determines that the situation is not a specific situation (No in step S14), the effective model determination unit 36 ​​activates the first detection model, and the person detection unit 14 detects a person included in the video using the first detection model trained by machine learning using the single person image as a correct answer image (step S18). The display control unit 18 generates a display image by adding the person detection result by the person detection unit 14 to the image acquired by the image acquisition unit 20, and displays the generated display image on the display device 50 (step S20). The processes from step S10 to step S20 are repeatedly executed while the recognition processing device 10 is operating or while the image is being captured by the camera 40.

[0046] 7, the determination in step S14 may be whether or not there is a range that is a specific situation. In this case, the process in step S16 detects a person by validating the second detection model for the range that is the specific situation, and detects a person by validating the first detection model for the range that is not the specific situation.

[0047] According to this embodiment, when the specific situation is not one in which it is estimated that there are few people on the road, a person included in the video is detected by the first detection unit 30. As a result, a single person can be appropriately detected in a situation in which there is a high possibility that multiple people will appear individually without overlapping due to the low number of people on the road.

[0048] According to this embodiment, in a specific situation where it is estimated that there is a lot of people on the road, people included in the video are detected by the second detection unit 32. As a result, in a situation where there is a high possibility that multiple people will appear to overlap due to the high number of people on the road, people who appear to overlap can be appropriately detected.

[0049] If the first detection unit 30 is to detect people who appear to overlap, there is a possibility that an event will occur in which the person can be detected or not detected depending on the manner in which the multiple people overlap. For example, when the degree of overlap between the first person and the second person is large, the difference from the single person image is large, so the first detection unit 30 may not be able to detect both the first person and the second person. On the other hand, when the degree of overlap between the first person and the second person is small, the difference from the single person image is small, so the first detection unit 30 may be able to detect only the first person, and the first detection unit 30 may be able to detect both the first person and the second person. When multiple people are walking and moving, the manner in which the multiple people included in the video overlap may change over time. In this case, there is a possibility that, in the multiple frames constituting the video, a frame in which a person is detected by the first detection unit 30 and a frame in which a person is not detected by the first detection unit 30 are consecutive. When an additional image such as a frame image is added based on the person detection result of the first detection unit 30, there is a possibility that a frame in which the additional image is displayed and a frame in which the additional image is not displayed will be consecutive, and the additional image will be displayed fluctuating over time as if it is blinking. Even if a person is included in the image, it is not preferable for the additional image to be displayed fluctuatingly as if it is blinking in the display image. According to this embodiment, it is possible to reduce the possibility that such an inappropriate display image will be displayed.

[0050] Second embodiment 8 is a block diagram showing a schematic functional configuration of a recognition processing device 10A according to the second embodiment. The second embodiment differs from the first embodiment in that the determination unit 16 includes a stability determination unit 35 instead of the situation determination unit 34. The second embodiment will be described below, focusing on the differences from the first embodiment, and descriptions of commonalities will be omitted as appropriate.

[0051] The recognition processing device 10A includes an acquisition unit 12, a person detection unit 14, and a determination unit 16. The recognition processing device 10A may include a display control unit 18. The acquisition unit 12 includes a video acquisition unit 20. The person detection unit 14 includes a first detection unit 30 and a second detection unit 32. The determination unit 16 includes a stability determination unit 35 and an effective model determination unit 36A. The display control unit 18, the video acquisition unit 20, the first detection unit 30, and the second detection unit 32 are configured similarly to the first embodiment described above.

[0052] The stability determination unit 35 determines the stability of the person detection process by the person detection unit 14. The stability determination unit 35 determines the stability based on the variation in the number of people detected by the person detection unit 14 in multiple consecutive frames of the acquired video. If there is little variation in the number of people detected in multiple consecutive frames, the stability determination unit 35 determines that the stability of the detection process is high. If there is a large variation in the number of people detected in multiple consecutive frames, the stability determination unit 35 determines that the stability of the detection process is low.

[0053] Here, the state where the detection process by the person detection unit 14 is stable refers to a state where, when a video including a person is acquired, the person detection unit 14 can appropriately detect a person in a plurality of consecutive frames. When the detection process is stable, if there is no change in the actual number of people included in the video, the number of people detected in a plurality of consecutive frames is constant, so there is no variation in the number of detections. On the other hand, the state where the detection process by the person detection unit 14 is unstable refers to a state where, when a video including a person is acquired, there is frequent switching between a case where the person detection unit 14 can detect a person in a plurality of consecutive frames and a case where it cannot detect a person. When the detection process is unstable, the number of people detected in a plurality of consecutive frames fluctuates even though there is no change in the actual number of people included in the video, so there is variation in the number of detections.

[0054] For example, when the first detection unit 30 is to detect a person included in an image in which multiple people appear to overlap in a situation with a lot of people, an event may occur in which the person can be detected or not detected depending on the manner in which the multiple people overlap. In this case, the number of people detected by the first detection unit 30 varies, so the detection process by the first detection unit 30 can be said to be unstable. Also, when the second detection unit 32 is to detect a person included in an image in which only a single person appears in a situation with a few people, an event may occur in which the person can be detected or not detected depending on the state of the background of the single person. In this case, the number of people detected by the second detection unit 32 varies, so the detection process by the second detection unit 32 can be said to be unstable.

[0055] The stability determination unit 35 records the number of people detected by the person detection unit 14 for each frame constituting the acquired video. The stability determination unit 35 calculates the variance in the number of people detected using the number of people detected for multiple consecutive frames in a predetermined period (e.g., 1 second or more and 5 seconds or less). The variance in the number of detections can be represented by the variance or standard deviation of the number of people detected recorded for multiple consecutive frames. The variance in the number of detections may be represented by a value obtained by summing up the difference in the number of people detected between adjacent frames over a predetermined period and dividing the sum by the number of frames in the predetermined period.

[0056] The stability determination unit 35 may calculate a score indicating the stability of the detection of a person by the person detection unit 14. The stability determination unit 35 may calculate a first score indicating the stability of the detection of a person using the first model based on the variation in the number of people detected by the first detection unit 30. The stability determination unit 35 may calculate a second score indicating the stability of the detection of a person using the second model based on the variation in the number of people detected by the second detection unit 32. The first score and the second score may be values ​​indicating the variation in the number of people detected, or may be the variance or standard deviation of the number of people detected. In this case, the higher the stability, the lower the score, and the lower the stability, the higher the score.

[0057] The stability determination unit 35 may determine whether or not the detection of a person using the first model is stable based on the variation in the number of people detected by the first detection unit 30. When the variation in the number of people detected by the first detection unit 30 is less than a predetermined reference value, the stability determination unit 35 may determine that the detection of a person using the first model is stable and set the first score to "0". When the variation in the number of people detected by the first detection unit 30 is equal to or greater than a predetermined reference value, the stability determination unit 35 may determine that the detection of a person using the first model is stable and set the first score to "1". Similarly, the stability determination unit 35 may determine whether or not the detection of a person using the second model is stable based on the variation in the number of people detected by the second detection unit 32. When the variation in the number of people detected by the second detection unit 32 is less than a predetermined reference value, the stability determination unit 35 may determine that the detection of a person using the second model is stable and set the second score to "0". If the variation in the number of people detected by the second detection unit 32 is greater than or equal to a predetermined reference value, the stability determination unit 35 may determine that the detection of people using the second model is stable and set the second score to "1".

[0058] The stability determination unit 35 may determine the stability for the entire range of the acquired image, or may determine the stability for a part of the acquired image. The stability determination unit 35 may determine the stability for a range excluding the outer periphery of the acquired image. In this case, the stability determination unit 35 determines the stability based on the variation in the number of people detected in the part of the image excluding the outer periphery, and determines the stability by ignoring the number of people detected in the outer periphery of the image. This makes it possible to eliminate the influence of the variation in the number of people detected due to people entering and leaving the outer periphery of the image, and to more appropriately determine the stability.

[0059] The stability determination unit 35 may determine the stability of a range of the acquired video that includes multiple people. FIG. 9 is a diagram showing an example of a partial range 76 of the video that is the target of stability determination. In the example of FIG. 9, multiple people 74a to 74f shown in a dashed frame are detected by the person detection unit 14, and a rectangular partial range 76 is set so as to include all of the detected multiple people 74a to 74f. The stability determination unit 35 determines the stability based on the variation in the number of people detected in the partial range 76 that includes the multiple people 74a to 74f, and determines the stability by ignoring the number of people detected outside the partial range 76. This allows the stability of detection of detected people to be appropriately determined.

[0060] The stability determination unit 35 may determine the stability of the range adjacent to each of the multiple people in the acquired video. Here, the range adjacent to the person is a slightly larger area than the area including the person detected by the person detection unit 14, and corresponds to the area occupied by the multiple people who appear to overlap. For example, when a first person is detected by the person detection unit 14, the area in which a second person who appears to overlap the first person may exist corresponds to the range adjacent to the first person. The range adjacent to the detected person may include the area occupied by the detected person, and may be a range expanded in at least one of the vertical and horizontal directions with the area occupied by the detected person as the center. The range adjacent to the detected person may be set so that, for example, at least one of the vertical and horizontal sizes is 1.5 times or more and 3 times or less than the area occupied by the detected person.

[0061] When multiple people are detected by the person detection unit 14, the stability determination unit 35 may evaluate the stability for each range adjacent to each of the multiple detected people. For example, in the example of Fig. 9, the stability may be evaluated in the range adjacent to the first person 74a and the stability may be evaluated in the range adjacent to the second person 74b. In this case, the stability of person detection by the person detection unit 14 can be evaluated individually for each of the multiple detected people.

[0062] The effective model determination unit 36A has a configuration similar to that of the effective model determination unit 36 ​​according to the first embodiment, but differs from the first embodiment in that it uses the determination result of the stability determination unit 35. The effective model determination unit 36A determines, based on the determination result of the stability determination unit 35, whether the detection of a person using the first detection model or the second detection model is to be valid.

[0063] The effective model determination unit 36A may determine whether the detection of a person using the first detection model or the second detection model is effective, based on either the first score or the second score calculated by the stability determination unit 35.

[0064] For example, when a person is detected by the first detection unit 30 and a first score indicating the stability of the first detection unit 30 is calculated, the effective model determination unit 36A may determine whether to enable the first detection unit 30 or the second detection unit 32 based on the first score. When it is determined that the first detection unit 30 is stable based on the first score, the effective model determination unit 36A may enable the first detection unit 30. When it is determined that the first detection unit 30 is not stable based on the first score, the effective model determination unit 36A may enable the second detection unit 32.

[0065] For example, when a person is detected by the second detection unit 32 and a second score indicating the stability of the second detection unit 32 is calculated, the effective model determination unit 36A may determine whether to enable the first detection unit 30 or the second detection unit 32 based on the second score. When it is determined that the second detection unit 32 is stable based on the second score, the effective model determination unit 36A may enable the second detection unit 32. When it is determined that the second detection unit 32 is not stable based on the second score, the effective model determination unit 36A may enable the first detection unit 30.

[0066] The effective model determination unit 36A may determine whether to validate the detection of a person using the first detection model or the second detection model, based on a comparison result between the first score and the second score calculated by the stability determination unit 35. For example, when the first detection unit 30 and the second detection unit 32 function in parallel and both the first score and the second score are calculated, the effective model determination unit 36A may validate the model with higher stability by comparing the first score and the second score.

[0067] The valid model determination unit 36A may enable the second detection unit 32 when the stability of the second detection unit 32 is higher than that of the first detection unit 30. The valid model determination unit 36A may enable the first detection unit 30 when the stability of the first detection unit 30 is higher than that of the second detection unit 32. The valid model determination unit 36A may enable the first detection unit 30 when the first score and the second score are equivalent and the stability of both the first detection unit 30 and the second detection unit 32 is high. The valid model determination unit 36A may enable the second detection unit 32 when the first score and the second score are equivalent and the stability of both the first detection unit 30 and the second detection unit 32 is low.

[0068] When the stability of the first detection unit 30 or the second detection unit 32 is determined for a partial range of the image, the effective model determination unit 36A may determine which of the first detection model or the second detection model is to be enabled for the partial range of the image. For example, when the first person and the second person are detected, the effective model determination unit 36A may determine which of the first detection unit 30 or the second detection unit 32 is to be enabled in the adjacent range of the first person based on at least one of the first score and the second score calculated for the adjacent range of the first person. For example, when the first person and the second person are detected, the effective model determination unit 36A may determine which of the first detection unit 30 or the second detection unit 32 is to be enabled in the adjacent range of the second person based on at least one of the first score and the second score calculated for the adjacent range of the second person. For example, the detection by the first detection unit 30 may be enabled in the adjacent range of the first person, and the detection by the second detection unit 32 may be enabled in the adjacent range of the second person.

[0069] The person detection unit 14 may detect a person included in the video by using a detection model that is made valid by the valid model determination unit 36A. The person detection unit 14 may not operate a detection model that is made invalid by the valid model determination unit 36A. When the situation determination unit 34 makes the first detection model valid, the person detection unit 14 may operate only the first detection unit 30 and stop the function of the second detection unit 32. When the situation determination unit 34 makes the second detection model valid, the person detection unit 14 may operate only the second detection unit 32 and stop the function of the first detection unit 30. The person detection unit 14 may operate the first detection unit 30 and the second detection unit 32 in parallel, regardless of the determination result of the situation determination unit 34.

[0070] FIG. 10 is a flowchart showing an example of the flow of the recognition processing method according to the second embodiment. FIG. 10 shows a process flow in the case where the first detection unit 30 or the second detection unit 32 is selectively operated. The video acquisition unit 20 acquires a video captured by the camera 40 (step S30). The person detection unit 14 detects a person included in the video using a first detection model trained by machine learning on a single person image as a correct answer image, or a second detection model trained by machine learning on a superimposed person image as a correct answer image (step S32). The display control unit 18 generates a display video in which the person detection result by the person detection unit 14 is added to the video acquired by the video acquisition unit 20, and causes the display device 50 to display the generated display video (step S34).

[0071] The stability determination unit 35 determines the stability of the first detection model or the second detection model based on the variation in the number of people detected by the person detection unit 14 (step S36). When the person detection unit 14 uses the first detection model, the stability determination unit 35 determines the stability of the first detection model based on the variation in the number of people detected using the first detection model. When the person detection unit 14 uses the second detection model, the stability determination unit 35 determines the stability of the second detection model based on the variation in the number of people detected using the second detection model. When the stability determination unit 35 determines that the current detection model is not stable (No in step S38), the effective model determination unit 36A invalidates the current detection model and validates a detection model other than the current detection model, and the person detection unit 14 changes to the validated detection model (step S40). If it is determined that the first detection model is not stable, the effective model determination unit 36A validates the second detection model, and the human detection unit 14 changes from the first detection model to the second detection model. If it is determined that the second detection model is not stable, the effective model determination unit 36A validates the first detection model, and the human detection unit 14 changes from the second detection model to the first detection model. If it is determined that the current detection model is stable by the stability determination unit 35 (Yes in step S38), the process of step S40 is skipped. In this case, the effective model determination unit 36A validates the current detection model, and the human detection unit 14 continues to use the current detection model that has been validated.

[0072] The processes from step S30 to step S40 are repeatedly executed while the recognition processing device 10 is operating or while an image is being captured by the camera 40. If the detection model is changed in step S40, a person included in the image is detected using the changed detection model in step S32.

[0073] FIG. 11 is a flowchart showing another example of the flow of the recognition processing method according to the second embodiment. FIG. 11 shows a process flow in the case where the first detection unit 30 and the second detection unit 32 are made to function in parallel. The video acquisition unit 20 acquires a video captured by the camera 40 (step S50). The person detection unit 14 detects a person included in the video using a first detection model trained by machine learning on a single person image as a correct answer image (step S52), and detects a person included in the video using a second detection model trained by machine learning on a superimposed person image as a correct answer image (step S52). The stability determination unit 35 determines the stability of the first detection model based on the variation in the number of people detected by the first detection unit 30, and determines the stability of the second detection model based on the variation in the number of people detected by the second detection unit 32 (step S56). The effective model determination unit 36A determines whether the detection of a person using the first detection model or the second detection model is effective based on the stability of each of the first detection model and the second detection model (step S58). The display control unit 18 generates a display image by adding the person detection result by the enabled detection model to the image acquired by the image acquisition unit 20, and causes the display device 50 to display the generated display image (step S60). The processes from step S50 to step S60 are repeatedly executed while the recognition processing device 10 is operating or while the camera 40 is capturing an image.

[0074] According to this embodiment, by determining the stability of the first detection model or the second detection model, a person included in the video can be detected using a detection model with a more stable person detection process. For example, in a situation where there are many people, multiple people are seen overlapping each other, making the stability of detection by the first detection unit 30 low, the person included in the video can be detected by the second detection unit 32. Conversely, in a situation where there are few people, only a single person is seen, making the stability of detection by the second detection unit 32 low, the person included in the video can be detected by the first detection unit 30. As a result, an appropriate detection model can be adopted depending on the situation, and the person included in the video can be more appropriately detected.

[0075] Third embodiment 12 is a block diagram showing a schematic functional configuration of a recognition processing device 10B according to the third embodiment. The third embodiment differs from the first and second embodiments in that the determination unit 16 includes a situation determination unit 34B, a stability determination unit 35, an effective model determination unit 36B, and a history management unit 37. The following description of the third embodiment will focus on the differences from the first and second embodiments, and will omit a description of commonalities as appropriate.

[0076] The recognition processing device 10B includes an acquisition unit 12, a person detection unit 14, and a determination unit 16. The recognition processing device 10B may include a display control unit 18. The acquisition unit 12 includes a video acquisition unit 20 and a position information acquisition unit 22. The acquisition unit 12 may further include at least one of a map information acquisition unit 24, a time information acquisition unit 26, and an orientation information acquisition unit 28. The person detection unit 14 includes a first detection unit 30 and a second detection unit 32. The determination unit 16 includes a situation determination unit 34B, a stability determination unit 35, an effective model determination unit 36B, and a history management unit 37. The display control unit 18, the video acquisition unit 20, the position information acquisition unit 22, the map information acquisition unit 24, the time information acquisition unit 26, the orientation information acquisition unit 28, the first detection unit 30, the second detection unit 32, and the stability determination unit 35 are configured in the same manner as in the first or second embodiment described above.

[0077] The history management unit 37 manages history information of the stability determination result by the stability determination unit 35. The history management unit 37 records the stability determination result by the stability determination unit 35 in association with the imaging position. The history management unit 37 may record the stability determination result by the stability determination unit 35 in association with the imaging position and the imaging date and time. The history management unit 37 may record the stability determination result by the stability determination unit 35 in association with the imaging position and the imaging direction. The history management unit 37 may record the stability determination result by the stability determination unit 35 in association with the imaging position, the imaging date and time, and the imaging direction. In such a configuration, the determination unit 16 does not need to include the stability determination unit 35.

[0078] The history management unit 37 may manage history information of the person detection result by the person detection unit 14. The history management unit 37 may record history information on whether the first detection unit 30 or the second detection unit 32 detected a person in association with an imaging position. The history management unit 37 may record history information on whether the first detection unit 30 or the second detection unit 32 detected a person in association with an imaging position and an imaging date and time. The history management unit 37 may record history information on whether the first detection unit 30 or the second detection unit 32 detected a person in association with an imaging position and an imaging direction. The history management unit 37 may record history information on whether the first detection unit 30 or the second detection unit 32 detected a person in association with an imaging position, an imaging date and time, and an imaging direction.

[0079] The history management unit 37 records the determination result of the stability of the first detection model (e.g., a first score) when the stability of the first detection model is determined by the stability determination unit 35. The history management unit 37 records the determination result of the stability of the second detection model (e.g., a second score) when the stability determination unit 35 determines the stability of the second detection model.

[0080] The situation determination unit 34B determines whether or not the situation is a specific situation by using history information recorded in the history management unit 37. When history information matching the current imaging position is recorded in the history management unit 37, the situation determination unit 34B determines whether or not the situation is a specific situation based on past history information at the current imaging position. When the stability of the first detection model previously determined at the current imaging position is high, the situation determination unit 34B may determine that the situation is not a specific situation. When the stability of the first detection model previously determined at the current imaging position is low, the situation determination unit 34B may determine that the situation is a specific situation. When the stability of the second detection model previously determined at the current imaging position is high, the situation determination unit 34B may determine that the situation is not a specific situation. When the stability of the second detection model previously determined at the current imaging position is low, the situation determination unit 34B may determine that the situation is not a specific situation. When a person has been detected by the first detection model at the current imaging position in the past, the situation determination unit 34B may determine that the situation is not a specific situation. The situation determination section 34B may determine that the situation is a specific situation when a person has been detected in the current imaging position in the past by the second detection model.

[0081] The situation determination unit 34B may determine whether or not a specific situation exists based on past history information corresponding to the current imaging position and imaging time when history information matching the current imaging position and imaging time is recorded in the history management unit 37. The situation determination unit 34B may determine whether or not a specific situation exists based on past history information corresponding to the current imaging position and imaging direction when history information matching the current imaging position and imaging direction is recorded in the history management unit 37. The situation determination unit 34B may determine whether or not a specific situation exists based on past history information corresponding to the current imaging position, imaging time, and imaging direction when history information matching the current imaging position, imaging time, and imaging direction is recorded in the history management unit 37.

[0082] When history information that matches at least any one of the current imaging position, imaging time, and imaging direction is not recorded in the history management unit 37, the situation determination unit 34B may determine whether or not the situation is a specific situation depending on whether or not the specific condition included in the map device 44 is matched. In other words, when history information corresponding to the current situation is recorded in the history management unit 37, the situation determination unit 34B may determine whether or not the situation is a specific situation by preferentially using the history information recorded in the history management unit 37.

[0083] The effective model determination unit 36B determines whether to use the first detection model or the second detection model to detect a person based on the determination result of at least one of the situation determination unit 34B and the stability determination unit 35. When a person is not detected by the person detection unit 14, the effective model determination unit 36B may determine whether to use the first detection model or the second detection model to detect a person based on the determination result of the situation determination unit 34B. When a person is detected by the person detection unit 14, the effective model determination unit 36B may determine whether to use the first detection model or the second detection model to detect a person based on the determination result of the stability determination unit 35.

[0084] FIG. 13 is a flowchart showing an example of the flow of the recognition processing method according to the third embodiment. FIG. 13 shows a processing flow in the case where the first detection unit 30 or the second detection unit 32 is selectively operated. The processing of step S70 and steps S76 to S80 shown in FIG. 13 is the same as the processing of step S10 and steps S16 to S20 shown in FIG. 7, and therefore the description will be omitted. The position information acquisition unit 22 acquires position information indicating the image capture position from the position sensor 42. The situation determination unit 34B acquires history information corresponding to the position information (step S72). The situation determination unit 34B determines whether or not the position is a position that has been determined to be a specific situation in the past based on the image capture position and the history information (step S74). If the situation determination unit 34B determines that the history information corresponding to the image capture position is a specific situation (Yes in step S74), the processing of step S76 is executed. If the situation determination unit 34B determines that the history information corresponding to the image capture position is not a specific situation (No in step S74), the processing of step S78 is executed.

[0085] According to this embodiment, when there is past history information corresponding to the current imaging position, it is possible to determine whether to activate the first detection unit 30 or the second detection unit 32 based on the history information. An appropriate detection model can be adopted according to past performance, and people included in the video can be detected more appropriately.

[0086] The present invention has been described above with reference to the above-mentioned embodiment, but the present invention is not limited to the above-mentioned embodiment, and appropriate combinations or substitutions of the respective configurations shown in the embodiment are also included in the present invention.

[0087] Several aspects of the disclosure are described below.

[0088] A first aspect of the present disclosure is a recognition processing device including: an image acquisition unit that acquires an image captured by a camera mounted on a moving body; a person detection unit that detects a person included in the acquired image using at least one of a first detection model trained by machine learning on a single person image as a correct image and a second detection model trained by machine learning on a superimposed person image as a correct image; a situation determination unit that determines whether a specific situation is estimated to be high in foot traffic using location information indicating the capture location of the acquired image and map information indicating a location estimated to be high in foot traffic; and an effective model determination unit that determines whether person detection using the first detection model or the second detection model is effective based on a determination result of the situation determination unit.

[0089] In a first aspect, the situation determination unit may further use at least one of map information indicating a combination of a location and date and time that is estimated to be busy with people, time information indicating the date and time of image capture by the camera, and orientation information indicating the image capture direction of the image to determine whether or not the specific situation exists.

[0090] In the first aspect, the video camera may further include a history management unit that records history information that associates position information indicating an imaging position of the acquired video with a person detection result by the person detection unit. In the first aspect, the situation determination unit may further use the history information to determine whether or not the specific situation exists.

[0091] The first aspect may be provided as a recognition processing method. This method may include the steps of acquiring a video captured by a camera installed on a moving object, detecting a person included in the acquired video using at least one of a first detection model in which a single person image is machine-learned as a correct image and a second detection model in which a superimposed person image is machine-learned as a correct image, determining whether or not the situation is a specific situation in which it is estimated that there is a lot of people using location information indicating an imaging position of the acquired video and map information indicating a place in which it is estimated that there is a lot of people, and determining whether or not the situation is a specific situation using either the first detection model or the second detection model based on the determination result of whether or not the situation is a specific situation. This method may be configured to cause a computer to execute each step.

[0092] The first aspect may be provided as a program or a non-transitory recording medium storing the program. The program may be configured to cause a computer to realize a function of acquiring a video captured by a camera provided on a moving object, a function of detecting a person included in the acquired video using at least one of a first detection model in which a single person image is trained as a correct answer image and a second detection model in which a superimposed person image is trained as a correct answer image, a function of determining whether or not the specific situation is estimated to be a situation with a lot of people using location information indicating an imaging position of the acquired video and map information indicating a place estimated to be a place with a lot of people, and a function of determining whether or not the specific situation is detected using either the first detection model or the second detection model based on a result of the determination of whether or not the specific situation is detected.

[0093] A second aspect of the present disclosure is a person detection model generation method for generating a person detection model by machine learning using a superimposed person image that includes a full-body image of a person and includes at least a part of another person as a background of the person as a ground truth image. In the second aspect, the machine learning does not need to use a single person image that includes a full-body image of a person and does not include another person as a background of the person as a ground truth image.

[0094] A third aspect of the present disclosure is a recognition processing device including: an image acquisition unit that acquires an image captured by a camera mounted on a moving body; a person detection unit that detects a person included in the acquired image using at least one of a first detection model trained by machine learning on a single person image as a correct image and a second detection model trained by machine learning on a superimposed person image as a correct image; a stability determination unit that determines stability of person detection processing by the person detection unit based on variance in the number of people detected by the person detection unit in multiple consecutive frames of the acquired image; and an effective model determination unit that determines whether person detection using the first detection model or the second detection model is effective based on a result of the stability determination.

[0095] In a third aspect, the stability determination unit may determine the stability of person detection using the first detection model when the person detection unit uses the first detection model, and may determine the stability of person detection using the second detection model when the person detection unit uses the second detection model. In the third aspect, the effective model determination unit may determine whether person detection using the first detection model or the second detection model is to be effective based on a result of determining the stability of either the first detection model or the second detection model.

[0096] In a third aspect, the stability determination unit may determine the stability of person detection using the first detection model and the stability of person detection using the second detection model. In the third aspect, the effective model determination unit may determine whether person detection using the first detection model or the second detection model is effective based on a comparison result of the stabilities of the first detection model and the second detection model.

[0097] In a third aspect, the stability determination section may determine the stability based on a variation in the number of people detected by the person detection section in a range excluding an outer periphery of the acquired video.

[0098] In a third aspect, the stability determination section may determine the stability based on a variation in the number of people detected by the person detection section in an area in the acquired video that includes a plurality of people.

[0099] In a third aspect, the stability determination unit may determine the stability based on a variation in the number of people detected by the person detection unit in a range adjacent to each of a plurality of people in the acquired video.

[0100] The third aspect may be provided as a recognition processing method. This method may include the steps of acquiring a video captured by a camera installed on a moving object, detecting a person included in the acquired video using at least one of a first detection model trained by machine learning on a single person image as a correct answer image and a second detection model trained by machine learning on a superimposed person image as a correct answer image, determining the stability of the detection based on the variation in the number of people detected in a plurality of consecutive frames of the acquired video, and determining whether the detection of the person using the first detection model or the second detection model is valid based on the result of the determination of the stability. This method may be configured to cause a computer to execute each step.

[0101] The third aspect may be provided as a program or a non-transitory recording medium storing the program. The program may be configured to cause a computer to realize a function of acquiring a video captured by a camera provided on a moving object, a function of detecting a person included in the acquired video using at least one of a first detection model trained by machine learning on a single person image as a correct answer image and a second detection model trained by machine learning on a superimposed person image as a correct answer image, a function of determining stability of the detection based on a variance in the number of people detected in a plurality of consecutive frames of the acquired video, and a function of determining whether to use the first detection model or the second detection model to detect a person based on a result of the determination of the stability. [Explanation of symbols]

[0102] 10, 10A, 10B... recognition processing device, 12... acquisition unit, 14... person detection unit, 16... judgment unit, 18... display control unit, 20... image acquisition unit, 30... first detection unit, 32... second detection unit, 34, 34B... situation judgment unit, 35... stability judgment unit, 36, 36A, 36B... valid model judgment unit, 37... history management unit, 40... camera.

Claims

1. an image acquisition unit that acquires an image captured by a camera provided on the moving object; a person detection unit that detects a person included in the acquired video using at least one of a first detection model that has been machine-learned using a single person image as a correct image and a second detection model that has been machine-learned using a superimposed person image as a correct image; a situation determination unit that determines whether or not a specific situation is estimated to be a place with a lot of people using location information that indicates an imaging position of the acquired video and map information that indicates a place estimated to be a place with a lot of people; an effective model determination unit that determines whether to use the first detection model or the second detection model to detect a person based on a determination result of the situation determination unit.

2. 2. The recognition processing device according to claim 1, wherein the situation determination unit further uses at least one of map information indicating a combination of a location and date and time that is estimated to be busy with people, time information indicating a date and time when the image was captured by the camera, and orientation information indicating a direction in which the image was captured to determine whether the specific situation exists.

3. a history management unit that records history information that associates position information indicating an imaging position of the acquired video with a person detection result by the person detection unit, The recognition processing device according to claim 1 , wherein the situation determination unit further uses the history information to determine whether the specific situation exists.

4. acquiring an image captured by a camera provided on the moving object; Detecting a person included in the acquired video using at least one of a first detection model trained by machine learning using a single person image as a correct answer image and a second detection model trained by machine learning using a superimposed person image as a correct answer image; determining whether the specific situation is estimated to be a place with a lot of people using location information indicating an imaging position of the acquired video and map information indicating a place estimated to be a place with a lot of people; and determining whether to use the first detection model or the second detection model to detect a person based on the result of determining whether the specific situation exists.

5. An image acquisition unit that acquires an image captured by a camera installed on a moving object; a person detection unit that detects a person included in the acquired video using at least one of a first detection model that has been machine-learned using a single person image as a correct image and a second detection model that has been machine-learned using a superimposed person image as a correct image; a stability determination unit that determines the stability of a person detection process performed by the person detection unit based on a variation in the number of people detected by the person detection unit in a plurality of consecutive frames of the acquired video; and an effective model determination unit that determines whether to use the first detection model or the second detection model to detect a person based on the stability determination result.

6. The stability determination unit, when the person detection unit uses the first detection model, determines the stability of person detection using the first detection model, and when the person detection unit uses the second detection model, determines the stability of person detection using the second detection model; 6. The recognition processing device according to claim 5, wherein the effective model determination unit determines whether to use the first detection model or the second detection model to detect a person based on a determination result of stability of either the first detection model or the second detection model.

7. The stability determination unit determines the stability of person detection using the first detection model and the stability of person detection using the second detection model; 6. The recognition processing device according to claim 5, wherein the effective model determination unit determines whether to use the first detection model or the second detection model to detect a person based on a comparison result of the stability of each of the first detection model and the second detection model.

8. A step of acquiring an image captured by a camera installed on a moving object; Detecting a person included in the acquired video using at least one of a first detection model trained by machine learning using a single person image as a correct answer image and a second detection model trained by machine learning using a superimposed person image as a correct answer image; determining stability of the detection based on a variation in the number of people detected in a plurality of consecutive frames of the acquired video; and determining whether to use the first detection model or the second detection model to detect a person based on the stability determination result.