Image processing device, image processing method, and program

The image processing device uses congestion and movement analysis to detect vulnerable individuals in crowded scenes, addressing the challenge of obscured targets by focusing on crowd dynamics rather than visual features, thereby improving detection accuracy.

JP7861517B2Active Publication Date: 2026-05-19NEC CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
NEC CORP
Filing Date
2022-06-08
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing image analysis techniques struggle to accurately detect vulnerable individuals in crowded environments when they are obscured by others, as relying solely on visual features is insufficient.

Method used

An image processing device that identifies congestion levels, movement patterns, and movement speeds within a series of images to estimate the presence of targets, allowing detection even when visual features are obscured.

Benefits of technology

Enables accurate detection of vulnerable individuals in crowded scenes by utilizing congestion, movement patterns, and speed analysis, enhancing detection accuracy even when visual features are not visible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007861517000001
    Figure 0007861517000001
  • Figure 0007861517000002
    Figure 0007861517000002
  • Figure 0007861517000003
    Figure 0007861517000003
Patent Text Reader

Abstract

To detect a detection object using image analysis.SOLUTION: The present invention provides an image processing device 10 including an identification unit 11 configured to identify at least one of degree of congestion, line of flow, and movement speed of people included in a plurality of images that are continuous in time series; and an estimated presence location detection unit 12 configured to detect, based on an identification result, an estimated presence location that is a location where the presence of a detection object is estimated, from the images.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] A technique related to the present invention is disclosed in Patent Document 1. Patent Document 1 discloses a technique for detecting traffic vulnerable people such as infants, elderly people, wheelchair users, and white cane users from a crowd in image analysis.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The inventor of the present invention has found the following problem in the technique for detecting a detection target. When the detection target is hidden by another person, it may not be possible to detect the detection target.

[0005] Patent Document 1 discloses a technique for detecting traffic vulnerable people and the like from a crowd in image analysis, but does not disclose the above problems and their solutions.

[0006] An example of the object of the present invention is to provide an image processing apparatus, an image processing method, and a program that solve the problem of detecting a detection target by image analysis in view of the above-described problems.

Means for Solving the Problems

[0007] According to one aspect of the present invention, specification means for specifying at least one of the congestion degree, flow line, and moving speed of people included in a plurality of images continuous in time series; A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, is used to detect such locations within the image. An image processing device having the following is provided.

[0008] According to one aspect of the present invention, Computers Identify at least one of the following in a series of images: the degree of crowding, movement patterns, and movement speed of people. An image processing method is provided for detecting locations in the image where the presence of a target is estimated, based on the aforementioned specific results.

[0009] According to one aspect of the present invention, Computers, A means for identifying at least one of the degree of crowding, movement patterns, and movement speed of people included in a series of images in a time-series sequence, and A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, within the image. A program is provided to enable it to function as such. [Effects of the Invention]

[0010] According to one aspect of the present invention, an image processing apparatus, an image processing method, and a program are realized that solve the problem of detecting a target object by image analysis. [Brief explanation of the drawing]

[0011] The purposes mentioned above, as well as other purposes, features, and benefits, are described below. Suitable This will become even clearer from the following embodiments and accompanying drawings.

[0012] [Figure 1] This figure shows an example of a functional block diagram of an image processing device. [Figure 2] This figure shows an example of the hardware configuration of an image processing device. [Figure 3]It is a diagram showing an example of a functional block diagram of an image processing apparatus. [Figure 4] It is a diagram for explaining an example of a process of detecting a location where a person's presence is estimated based on the degree of crowding of people. [Figure 5] It is a diagram for explaining an example of a process of detecting a location where a person's presence is estimated based on the movement route of people. [Figure 6] It is a diagram for explaining another example of a process of detecting a location where a person's presence is estimated based on the movement route of people. [Figure 7] It is a diagram for explaining another example of a process of detecting a location where a person's presence is estimated based on the movement route of people. [Figure 8] It is a diagram showing an example of information output by the image processing apparatus. [Figure 9] It is a flowchart showing an example of the processing flow of the image processing apparatus. [Figure 10] It is a diagram showing an example of a functional block diagram of the image processing apparatus. [Figure 11] It is a diagram showing an example of information output by the image processing apparatus. [Figure 12] It is a flowchart showing an example of the processing flow of the image processing apparatus. [Figure 13] It is a diagram showing an example of a functional block diagram of the image processing apparatus. [Figure 14] It is a diagram showing an example of information output by the image processing apparatus. [Figure 15] It is a flowchart showing an example of the processing flow of the image processing apparatus. [Figure 16] It is a diagram for explaining an embodiment.

Mode for Carrying Out the Invention

[0013] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In all the drawings, the same components are denoted by the same reference numerals, and the description will be omitted as appropriate.

[0014] <First Embodiment> Figure 1 is a functional block diagram showing an overview of the image processing apparatus 10 according to the first embodiment. The image processing apparatus 10 includes a specific unit 11 and an existence estimation location detection unit 12.

[0015] The identification unit 11 identifies at least one of the crowd density, movement patterns, and movement speed of people included in a series of images that are consecutive in time. Based on the identification results from the identification unit 11, the presence estimation unit 12 detects presence estimation locations within the images, which are locations where the presence of the target is estimated.

[0016] The image processing device 10 with this configuration solves the problem of detecting objects through image analysis.

[0017] <Second Embodiment> "overview" The image processing apparatus 10 of the second embodiment is a more concrete example of the image processing apparatus 10 of the first embodiment.

[0018] Incidentally, in technologies for detecting targets from a crowd, there are challenges such as the following. If the visual features that identify the target are captured in the image, the target can be detected from the image by detecting those visual features. However, when the target is in a crowd, it may be hidden by other people, and the visual features that identify the target may not be captured in the image. For this reason, relying solely on means to detect the visual features that identify the target from the image is not sufficient to detect the target from a crowd with high accuracy. The image processing device 10 is equipped with means for accurately detecting the target from a crowd using image analysis.

[0019] The detected objects are those (including people and objects) that require some form of support or assistance. Examples include, but are not limited to, wheelchair users, white cane users, crutch users, sick people, injured people, lost children, other people requiring assistance, fallen objects, dropped objects, and other obstacles. Such detected objects may remain stationary or move more slowly than others. As a result, other people tend to move away from the location where the detected objects are present.

[0020] As a result, if the target is present in a crowd, (Feature 1) Within an area with high human congestion (crowd), there are areas with low human congestion (areas where the detected object exists). (Feature 2) Under normal circumstances (when no detection target is present), people pass through, but there are areas in the crowd that people avoid to pass (areas where detection target is present). (Feature 3) In a crowd, there is a person (detection target) whose movement speed is slower than the people around them. These characteristics appear.

[0021] The image processing device 10 detects the target from the crowd based on at least one of these features. In this process, the target can be detected from the image even if the visual features that identify the target are not visible in the image. The configuration of the image processing device 10 will be described in detail below.

[0022] "Hardware configuration" Next, an example of the hardware configuration of the image processing device 10 will be described. Each functional unit of the image processing device 10 is realized by any combination of hardware and software, centered around a CPU (Central Processing Unit) of any computer, memory, a program loaded into memory, a storage unit such as a hard disk that stores that program (which can store programs that are pre-installed at the time of shipment, as well as programs downloaded from recording media such as CDs (Compact Discs) or from servers on the Internet), and a network connection interface. It will be understood by those skilled in the art that there are various modifications to the implementation method and the device.

[0023] Figure 2 is a block diagram illustrating the hardware configuration of the image processing device 10. As shown in Figure 2, the image processing device 10 includes a processor 1A, memory 2A, input / output interface 3A, peripheral circuitry 4A, and bus 5A. Peripheral circuitry 4A includes various modules. The image processing device 10 does not necessarily have peripheral circuitry 4A. The image processing device 10 may also be composed of multiple physically and / or logically separated devices. In this case, each of the multiple devices may have the above hardware configuration.

[0024] Bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuits 4A, and input / output interface 3A to send and receive data to and from each other. Processor 1A is a processing unit such as a CPU or GPU (Graphics Processing Unit). Memory 2A is a memory such as RAM (Random Access Memory) or ROM (Read Only Memory). Input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. Input devices include, for example, keyboards, mice, microphones, physical buttons, touch panels, etc. Output devices include, for example, displays, speakers, printers, mailers, etc. Processor 1A can issue commands to each module and perform calculations based on the results of those calculations.

[0025] "Functional Configuration" Next, the functional configuration of the image processing apparatus 10 of the second embodiment will be described in detail. Figure 3 shows an example of a functional block diagram of the image processing apparatus 10. As shown in the figure, the image processing apparatus 10 has a specific unit 11, an existence estimation location detection unit 12, and an output unit 13.

[0026] The identification unit 11 identifies at least one of the crowding level, movement patterns, and movement speeds of people included in a series of images that are consecutive in time. The presence estimation location detection unit 12 then detects presence estimation locations in the images, which are locations where the presence of the target is estimated, based on the identification results from the identification unit 11. The presence estimation location detection unit 12 detects presence estimation locations in the images based, for example, on the relative positional relationship between crowded and uncrowded areas, the tendency of movement patterns of multiple people (tendency to avoid certain locations), and the trend between the movement speeds of multiple people.

[0027] "Multiple images in a time series" may, for example, be multiple frame images that make up a video. Alternatively, multiple images in a time series may be multiple still images taken in succession at relatively short time intervals. The images are, for example, images taken by a surveillance camera.

[0028] A "presumed location" is at least a portion of the image where the presence of the detected object is estimated based on at least one of the following factors: the degree of crowding, movement patterns, and movement speed of people within the image.

[0029] The following describes a specific example of a process for detecting locations where people are estimated to be present based on at least one of the following factors: the degree of crowding, movement patterns, and movement speed within an image. The image processing device 10 can perform one or more of the following processes.

[0030] "Process for detecting estimated locations of people based on the degree of crowding within an image (1)" In this process, based on the above-mentioned "(Feature 1) Within an area with high human congestion (crowd), there are areas with low human congestion (areas where the detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects "areas with low human congestion within an area with high human congestion" as estimated locations. There are various means for detecting "areas with low human congestion within an area with high human congestion" using image analysis, and any means can be employed in this embodiment. An example of this process is described below.

[0031] The identification unit 11 divides the image into multiple observation areas. The identification unit 11 divides the image into multiple observation areas D in a grid pattern, for example, as shown in Figure 5. The identification unit 11 then calculates the number of people present in each observation area D. The means for calculating the number of people present in each observation area D are not particularly limited. For example, the identification unit 11 detects people in the image using person detection technology (image analysis technology). The identification unit 11 then counts the number of people present in each observation area D. For example, the identification unit 11 may detect predetermined parts of a person's body present in the image (top of the head, nose, feet, etc.) and count the number of times such predetermined parts are present in each observation area D as the number of people present.

[0032] In this example, the identification unit 11 identifies the number of people (congestion level) present in each observation area D as described above. The presence estimation location detection unit 12 then detects locations as presence estimation locations where the number of people (congestion level) is surrounded by areas where the number of people (congestion level) is equal to or greater than the first congestion threshold value, and where the number of people (congestion level) is less than the second congestion threshold value, based on the identification results.

[0033] Here, observation area D where the number of people present (congestion level) is equal to or greater than the first congestion threshold is called the "congested observation area." Observation area D where the number of people present (congestion level) is less than the second congestion threshold is called the "uncongested observation area." In the example in Figure 5, the hatched observation area D is the congested observation area, and observation areas D1 to D6 are the uncongested observation areas.

[0034] An area where "the number of people present (congestion level) is equal to or greater than the first congestion threshold" is an area composed of multiple congestion observation areas. The presence estimation location detection unit 12 then detects one or more non-congested observation areas surrounded by such areas where "the number of people present (congestion level) is equal to or greater than the first congestion threshold" as presence estimation locations. In the example in Figure 5, the presence estimation location detection unit 12 detects observation areas D1 to D6, which are non-congested observation areas surrounded by observation area D filled with hatching, as presence estimation locations.

[0035] "Process for detecting locations where people are estimated to be present based on the degree of crowding within an image (2)" In this process, based on the above-mentioned "(Feature 1) Within an area with high human congestion (crowd), there are areas with low human congestion (areas where the detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects "areas where human congestion is low within an area with high human congestion, and where this state continues for a predetermined time or longer" as estimated locations. There are various means for detecting "areas where human congestion is low within an area with high human congestion, and where this state continues for a predetermined time or longer" using image analysis, and any means can be employed in this embodiment. Note that "predetermined time" may be defined as a time length or as a number of consecutive images. An example of this process is described below.

[0036] The identification unit 11 identifies the number of people (congestion level) present in each observation area D in the same manner as described in "Process for detecting estimated locations based on the degree of congestion of people in an image (1)". Then, the location detection unit 12 detects one or more non-congested observation areas from each of the multiple images, surrounded by "areas where the number of people (congestion level) present is equal to or greater than the first congestion criterion value", as candidates for estimated locations.

[0037] Next, the presence estimation location detection unit 12 detects candidates for presence estimation locations that have been detected at the same location for a predetermined period of time or longer (for example, in a predetermined number of consecutive images), and identifies them as presence estimation locations. This process will be explained using Figure 4.

[0038] Figure 4 shows three images in a time series. In the first image, area T1, where the number of people present (congestion level) is greater than or equal to the first congestion threshold, and candidate locations S1 and S2 within area T1 are detected. In the second image, area T2, where the number of people present (congestion level) is greater than or equal to the first congestion threshold, and candidate locations S3 and S4 within area T2 are detected. In the third image, area T3, where the number of people present (congestion level) is greater than or equal to the first congestion threshold, and candidate locations S5 and S6 within area T3 are detected. 。

[0039] In the example shown in Figure 4, the image locations of candidate locations S2, S4, and S6 coincide. That is, the same location is detected in three consecutive images. Therefore, the location detection unit 12 detects candidate locations S2, S4, and S6 as locations.

[0040] Furthermore, the "position of a candidate location for the estimated existence" can be any point (for example, the center) within the area occupied by the candidate location for the estimated existence. The "criterion for determining that the positions of two candidate locations for the estimated existence are the same" may be a perfect match, or it may be defined as a state where the difference is within a threshold. In addition, if two candidate locations for the estimated existence detected from two different images overlap in at least part, the two estimated locations for the estimated existence... Candidate The positions within the image can be considered to match.

[0041] "Process for detecting estimated locations based on movement patterns (1)" In this process, based on the above-mentioned "(Feature 2) There are places where people pass through under normal circumstances (when no detection target exists), but which people in a crowd avoid and pass through (places where a detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects "places where people in a crowd avoid and pass through" as estimated locations. There are various means for detecting "places where people in a crowd avoid and pass through" using image analysis, and any means can be employed in this embodiment. An example of this process is described below.

[0042] The identification unit 11 detects people in an image using human detection technology (image analysis technology), and then detects the movement path (travel trajectory) of each detected person. The detection of people's movement paths can be achieved using any technology.

[0043] The presence estimation location detection unit 12 then identifies areas in the image that people have not passed through as presence estimation locations, based on the calculated movement paths. For example, as shown in Figure 5, the presence estimation location detection unit 12 divides the image into multiple observation areas D. The presence estimation location detection unit 12 then determines that observation areas D through which movement paths calculated by the identification unit 11 pass are areas that people in the image pass through, and that observation areas D through which movement paths do not pass are areas that people in the image have not passed through (presence estimation locations). Alternatively, the presence estimation location detection unit 12 may determine that observation areas D through which movement paths for a predetermined number of people or more pass through are areas that people in the image pass through, and that observation areas D through which movement paths for less than a predetermined number of people pass through, and observation areas D through which no movement paths pass, are areas that people in the image have not passed through (presence estimation locations). In the example in Figure 5, observation areas D determined to be areas that people pass through are filled with hatching. Observation areas D1 to D6 are determined to be areas that people in the image have not passed through (presence estimation locations).

[0044] "Process for detecting estimated locations based on movement patterns (2)" In this process, based on the above-mentioned "(Feature 2) There are places where people pass through under normal circumstances (when no detection target exists), but where people in a crowd avoid passing (places where a detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects "places where people in a crowd avoid passing for a predetermined time or longer" as estimated locations. There are various means for detecting "places where people in a crowd avoid passing for a predetermined time or longer" using image analysis, and any means can be employed in this embodiment. Note that "predetermined time" may be defined as a time length or as a number of consecutive images. An example of this process is described below.

[0045] The presence estimation location detection unit 12 detects observation areas D in the image that have not been passed through, using the method described in "Process for detecting presence estimation locations based on movement paths (1)". The presence estimation location detection unit 12 then detects observation areas D in the image that have not been passed through as presence estimation locations, for a predetermined period of time or longer (for example, in a predetermined number of consecutive images).

[0046] "Process for detecting estimated locations based on movement patterns (3)" In this process, based on the above-mentioned "(Feature 2) There are places where people pass through under normal circumstances (when no detection target exists), but which people in a crowd avoid and pass through (places where a detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects "places where people pass through under normal circumstances (when no detection target exists), but which people in a crowd avoid and pass through" as estimated locations. There are various means of detecting "places where people pass through under normal circumstances (when no detection target exists), but which people in a crowd avoid and pass through" using image analysis, and any means can be employed in this embodiment. An example of this process is described below.

[0047] The identification unit 11 detects people in an image using human detection technology (image analysis technology), and then detects the movement path (travel trajectory) of each detected person. The detection of people's movement paths can be achieved using any technology.

[0048] The presence estimation location detection unit 12 then identifies areas in the image that people have not passed through, based on the calculated movement paths. For example, as shown in Figure 6, the presence estimation location detection unit 12 divides the image into multiple observation areas D. The presence estimation location detection unit 12 then determines that observation areas D through which movement paths calculated by the identification unit 11 pass are areas that people in the image pass through, and that observation areas D through which movement paths do not pass are areas that people in the image have not passed through. Alternatively, the presence estimation location detection unit 12 may determine that observation areas D through which movement paths for a predetermined number of people or more pass through are areas that people in the image pass through, and that observation areas D through which movement paths for less than a predetermined number of people pass through, and observation areas D through which no movement paths pass, are areas that people in the image have not passed through. In the example in Figure 6, observation areas D determined to be areas through which people pass are filled with hatching.

[0049] In this example, reference information showing the movement of people under normal circumstances (when no object to be detected is present) is generated in advance. Based on this reference information, the presence estimation location detection unit 12 determines whether or not a person normally passes through each observation area D, as shown in Figure 7. The criteria for this determination are the same as those used to determine whether or not a person passes through the image described above. In the example in Figure 7, the observation areas D that are determined to be passed through by people under normal circumstances are filled with hatching.

[0050] The presence estimation location detection unit 12 then detects, based on the data in Figures 6 and 7, observation areas D that people normally pass through but that are not shown in the image as presence estimation locations. In the example in Figures 6 and 7, observation areas D1 to D6 are determined to be areas that people normally pass through but that are not shown in the image (presumed presence locations).

[0051] "Process for detecting estimated locations based on movement patterns (4)" In this process, based on the above-mentioned "(Feature 2) There are places where people pass through under normal circumstances (when no detection target exists), but people in a crowd avoid passing through (places where a detection target exists)," estimated locations are detected. The estimated location detection unit 12 detects as estimated locations "places where people pass through under normal circumstances (when no detection target exists), but people in a crowd avoid passing through for a predetermined time or longer." There are various means for detecting "places where people pass through under normal circumstances (when no detection target exists), but people in a crowd avoid passing through for a predetermined time or longer" using image analysis, and any means can be employed in this embodiment. Note that "predetermined time" may be defined as a time length or as a number of consecutive images. An example of this process is described below.

[0052] The presence estimation location detection unit 12 uses the method described in "Processing to detect presence estimation locations based on movement paths (3)" to detect observation areas D in the image that people normally pass through but that people have not passed through. The presence estimation location detection unit 12 then detects as presence estimation locations observation areas D in the image that people normally pass through but that people have not passed through for a predetermined period of time or longer (for example, in a predetermined number of consecutive images or more).

[0053] "Process for detecting estimated locations based on movement speed" In this process, based on the above-mentioned "(Feature 3) In a crowd, there are people (detection targets) whose movement speed is slower than the people around them," locations where presence is estimated are detected. The presence estimation location detection unit 12 detects "locations where people whose movement speed is less than the speed standard value exist" as presence estimation locations. There are various means for detecting "locations where people whose movement speed is less than the speed standard value exist" using image analysis, and any means can be employed in this embodiment. An example of this process is described below.

[0054] The identification unit 11 detects people in an image using human detection technology (image analysis technology), and then detects the movement speed of each detected person. The detection of movement speed can be achieved using any technology.

[0055] The presence estimation location detection unit 12 then detects locations where a person whose movement speed is less than the speed standard value is present as a presence estimation location.

[0056] The speed reference value may be a predetermined value. By setting the speed reference value based on the average person's walking speed, it is possible to detect people whose walking speed is slower than the average person's speed. For example, the average person's walking speed may be used as the speed reference value, or any speed slower than the average person's walking speed may be used as the speed reference value.

[0057] In addition, the speed reference value may be a value calculated based on the movement speeds of multiple people included in the image. By setting the speed reference value based on the movement speeds of multiple people included in the image, it is possible to detect people whose movement speed is slower than the movement speed of the other people included in the image. For example, the movement speeds of multiple people included in the image may be used as the speed reference value, or any speed slower than the movement speeds of multiple people included in the image may be used as the speed reference value. The movement speeds of multiple people included in the image can be the statistical values ​​(mean, maximum, minimum, mode, median, etc.) of the movement speed of each person included in the image.

[0058] The output unit 13 outputs information indicating the location of the estimated existence detected by the location detection unit 12. For example, the output unit 13 causes an output device such as a display, projection device, or printer to output information indicating the detected location of the estimated existence.

[0059] For example, as shown in Figure 8, the output unit 13 may output an image in which information P indicating the detected estimated location is superimposed on the image processed by the identification unit 11 (an image processed to identify at least one of the degree of crowding, movement patterns, and movement speed of people). For example, the image processed by the identification unit 11 may be an image taken by a surveillance camera, and for surveillance purposes, the image may be displayed on a display in real time. The output unit 13 may then superimpose the information P indicating the detected estimated location onto this image displayed on the display in real time.

[0060] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 9. The image processing device 10 acquires multiple images that are consecutive in time series in the order they are generated, and repeats the processing of S10 to S12 each time an image is acquired.

[0061] First, the image processing device 10 identifies at least one of the following in the acquired image: the degree of crowding, movement patterns, and movement speed (S10). The movement patterns and movement speeds of the people are calculated based on the newly acquired image and one or more previously acquired images.

[0062] Next, the image processing device 10 detects locations in the image where the presence of the object to be detected is estimated, based on the specific result of S10 (S11). Then, the image processing device 10 outputs information indicating the locations where the presence was estimated in S11 (S12).

[0063] <Effects and Effects> The image processing device 10 of this embodiment identifies locations where the presence of a detection target is estimated (locations where presence is estimated) based not on external features that identify the detection target, but on at least one of the degree of crowding, movement patterns, and movement speed of people included in the image. Ba Even if the image does not contain any external features that identify the target, it can still detect the target from within the image.

[0064] Furthermore, the image processing device 10 identifies a location where the presence of a detection target is estimated, based on at least one of the above-described features 1 to 3 that appear when a detection target is present in a crowd. With such an image processing device 10, the location where the presence of a detection target is estimated can be identified with high accuracy.

[0065] <Third Embodiment> The image processing apparatus 10 of the third embodiment performs both a process to identify locations where the presence of a detection target is estimated (locations where presence is estimated) based on at least one of the degree of crowding, movement patterns, and movement speed of people included in the image, and a process to detect a detection target based on the characteristics of the appearance of the detection target. This will be described in detail below.

[0066] Figure 10 shows an example of a functional block diagram of the image processing apparatus 10 of this embodiment. As shown in the figure, the image processing apparatus 10 includes a specific unit 11, an estimated location detection unit 12, an output unit 13, and a detection target detection unit 14.

[0067] The detection target detection unit 14 detects the detection target from the image based on the characteristic features of the detection target's appearance. For example, the detection target detection unit 14 may detect the detection target by using object detection technology to detect an object used by the detection target (such as a wheelchair, white cane, or crutch). Alternatively, the detection target detection unit 14 may detect the detection target by using posture detection technology to detect a person who is in a posture specific to when using an object used by the detection target (such as a wheelchair, white cane, or crutch), or a person who is in a posture specific to a predetermined state, such as a sick or injured person (such as lying down or crouching). Furthermore, the detection target detection unit 14 may detect the detection target by using face recognition technology to detect a person whose face image has been pre-registered as a detection target.

[0068] In addition, the detection target detection unit 14 uses object detection technology to register objects in advance as detection targets. hand The target can also be detected by detecting existing objects (e.g., fallen signs, fallen trees, etc.).

[0069] The output unit 13 outputs information indicating the estimated location detected by the estimated location detection unit 12, as well as information indicating the detected target detected by the detected target detection unit 14. For example, the output unit 13 causes an output device such as a display, projection device, or printer to output information indicating the detected estimated location and information indicating the detected target.

[0070] For example, as shown in Figure 11, the output unit 13 may output an image in which information P indicating the detected location of presence and information R indicating the detected target are superimposed on the image processed by the identification unit 11 described above (an image processed to identify at least one of the degree of crowding, movement patterns, and movement speed of people). The information P indicating the detected location of presence and the information R indicating the detected target may be mutually identifiable information. For example, the shape, color, and intensity of the marks may be different, but the implementation means are not limited to these.

[0071] For example, the image processed by the specific unit 11 is an image captured by a surveillance camera, and for surveillance purposes, the image may be displayed on a screen in real time. The output unit 13 may then superimpose information P indicating the detected location of the estimated presence and information R indicating the detected object onto the image displayed on the screen in real time.

[0072] Furthermore, the outputted information R indicating the detected object may include not only information indicating the location of the detected object, but also information indicating details of the detection result. For example, the outputted information R indicating the detected object may include information indicating what was detected as the detected object (e.g., a person whose face image has been registered in advance, a wheelchair user, a white cane user, a crutch user, a sick person, an injured person, a lost person, other persons requiring assistance, a fallen object, a dropped object, or other obstacle). In addition, if a person whose face image has been registered in advance is detected, the outputted information R indicating the detected object may further include the face image that was registered in advance.

[0073] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 12. The image processing device 10 acquires multiple images that are consecutive in time series in the order they are generated, and repeats the processing from S20 to S23 each time an image is acquired.

[0074] First, the image processing device 10 identifies at least one of the following in the acquired image: the degree of crowding, movement patterns, and movement speed (S20). The movement patterns and movement speeds of people are calculated based on the newly acquired image and one or more previously acquired images. Next, based on the results identified in S20, the image processing device 10 detects locations in the image where the presence of the object to be detected is estimated (S21).

[0075] In addition, the image processing device 10 detects the object to be detected from the image based on the characteristic features of the object's appearance, in parallel with S20 and S21 (S22).

[0076] The image processing device 10 then outputs information indicating the location where existence was estimated in S21, and information indicating the detected object in S22 (S23).

[0077] The other configurations of the image processing apparatus 10 in this embodiment are the same as those of the image processing apparatus 10 in the first and second embodiments.

[0078] The image processing device 10 of this embodiment achieves the same effects as the image processing device 10 of the first and second embodiments. Furthermore, the image processing device 10 of this embodiment performs both a process to identify locations where the presence of a detection target is estimated (locations where presence is estimated) based on at least one of the degree of crowding, movement patterns, and movement speed of people included in the image, and a process to detect a detection target based on the external characteristics of the detection target. With such an image processing device 10, detection targets whose external characteristics are visible in the image are detected based on those external characteristics, and detection targets whose external characteristics are not visible in the image are detected based on those external characteristics. image Based on at least one of the following factors—the degree of human congestion, movement patterns, and movement speed—locations can be estimated to be present. As a result, it becomes possible to detect targets in any state with high accuracy.

[0079] <Fourth Embodiment> In previous embodiments, users could understand the existence and location of estimated locations where detection targets are presumed to exist, based on information output from the image processing device 10, but they could not understand what kind of detection targets existed at those estimated locations.

[0080] The image processing device 10 of the fourth embodiment generates and outputs information to understand what kind of detection target exists at the detected location, based on the "detection result of the detection target based on the external characteristics" described in the third embodiment.

[0081] Specifically, the image processing device 10 is • Detected object found near the estimated location of existence, • Detected objects detected at a time close to the detection timing of the estimated location of existence, and • Detected objects that are located near the estimated location of existence and detected at a time close to the time of detection of the estimated location of existence. At least one of these will be output as a candidate for the object to be detected at the estimated location. This will be explained in detail below.

[0082] Figure 13 shows an example of a functional block diagram of the image processing apparatus 10 of this embodiment. As shown in the figure, the image processing apparatus 10 includes a identification unit 11, an existence estimation location detection unit 12, an output unit 13, a detection target detection unit 14, and an extraction unit 15.

[0083] The extraction unit 15 extracts from the detection targets detected by the detection target detection unit 14 any detection targets that satisfy predetermined conditions between them and the estimated location detected by the location detection unit 12, and identifies them as candidates for detection targets located at that estimated location.

[0084] The specified conditions are, • The distance between the detection location of the detected object and the detection location of the estimated location is less than the distance threshold, and • The time difference between the detection timing of the target and the detection timing of the estimated location is less than the time threshold. It includes at least one of the following.

[0085] In addition, the specified conditions are as follows: • The estimated location and the detected object are not simultaneously detected within the same image. This may include the following. If the estimated location and the detected object are detected simultaneously in the same image, the detected object located at the estimated location cannot be that object. By adding this condition to the predetermined conditions, candidate objects located at the estimated location can be identified with high accuracy.

[0086] The extraction unit 15 may detect detection targets that satisfy predetermined conditions in relation to the estimated location from among the detection targets detected before the estimated location. Alternatively, the extraction unit 15 may detect detection targets that satisfy predetermined conditions in relation to the estimated location from among the detection targets detected after the estimated location. Furthermore, the extraction unit 15 may detect detection targets that satisfy predetermined conditions in relation to the estimated location from among the detection targets detected before and after the estimated location.

[0087] The output unit 13 outputs information linking the estimated location of existence with a detected object that satisfies predetermined conditions between that estimated location and the location. For example, the output unit 13 causes an output device such as a display, projection device, or printer to output information indicating the detected estimated location and information indicating the detected object.

[0088] For example, as shown in Figure 14, the output unit 13 may output an image in which information P indicating the detected location of presence and information Q indicating a detected object that satisfies predetermined conditions between the location of presence of presence and the image processed by the identification unit 11 described above (an image processed to identify at least one of the degree of crowding, movement patterns, and movement speed of people) are superimposed. Furthermore, information R indicating the detected object may also be displayed.

[0089] Furthermore, the outputted information Q indicating the detected object may include not only information indicating the location of the detected object, but also information indicating details of the detection result. For example, the outputted information Q indicating the detected object may include information indicating what was detected as the detected object (e.g., a person whose face image has been registered in advance, a wheelchair user, a white cane user, a crutch user, a sick person, an injured person, a lost person, other persons requiring assistance, a fallen object, a dropped object, or other obstacle). Also, if a person whose face image has been registered in advance is detected, the outputted information Q indicating the detected object may further include the face image that was registered in advance.

[0090] For example, the image processed by the specific unit 11 is an image captured by a surveillance camera, and for surveillance purposes, the image may be displayed on a screen in real time. The output unit 13 may then superimpose information P indicating the detected location of existence, and information Q indicating a detected object that satisfies predetermined conditions between itself and the location of existence of existence, onto the image displayed on the screen in real time.

[0091] Next, an example of the processing flow of the image processing device 10 will be explained using the flowchart in Figure 15. The image processing device 10 acquires multiple images that are consecutive in time series in the order they are generated, and repeats the processing from S30 to S34 each time an image is acquired.

[0092] First, the image processing device 10 identifies at least one of the following in the acquired image: the degree of crowding, movement patterns, and movement speed (S30). The movement patterns and movement speeds of people are calculated based on the newly acquired image and one or more previously acquired images. Next, based on the results identified in S30, the image processing device 10 detects locations in the image where the presence of the object to be detected is estimated (S31).

[0093] In addition, the image processing device 10 detects the object to be detected from the image based on the characteristic features of the object's appearance, in parallel with S30 and S31 (S32).

[0094] Then, in S32, the image processing device 10 extracts detection targets from among the detection targets previously detected that satisfy predetermined conditions in relation to the estimated existence location detected in S31 (S33). Next, the image processing device 10 outputs information linking the estimated existence location detected in S31 with the detection targets that satisfy predetermined conditions in relation to the estimated existence location (S34). In addition, in S34, information indicating the detection targets detected in S32 may be output.

[0095] The other configurations of the image processing apparatus 10 of this embodiment are the same as those of the image processing apparatus 10 of the first to third embodiments.

[0096] The image processing device 10 of this embodiment achieves the same effects as the image processing device 10 of the first to third embodiments. Furthermore, the image processing device 10 of this embodiment can output information (information Q in Figure 14) indicating candidate detection targets that are estimated to be present at the location where the detection target is estimated to be located. Based on the information output from the image processing device 10, the user can understand the existence and location of the location where the detection target is estimated to be present, as well as candidate detection targets present at that location.

[0097] <Examples> An example will be explained using Figure 16. The system of this embodiment can be used in facilities where large numbers of people gather, such as amusement parks, train stations, and sports facilities.

[0098] Visitors to the facility who require support from facility staff must access the transaction control server via a communication device such as a smartphone, tablet, mobile phone, or personal computer before their visit to register as a person requiring support, for example, by registering their facial image. This registration of the facial image may also be done using an application or website provided by the facility. If the registration is successful, the transaction control server will register the facial image in the facial data database.

[0099] Surveillance cameras are installed throughout the facility. Images captured by the surveillance cameras are transmitted in real time to a transaction control server by any means. The transaction control server sends the acquired images to a facial recognition server and requests the detection of people (targets for detection) registered in the facial data database. The facial recognition server performs the processing in response to the request and returns the result to the transaction control server.

[0100] Furthermore, the transaction control server sends the acquired images to the video analysis server and requests image analysis. The video analysis server uses all available image analysis technologies, such as human detection, object detection, posture detection, movement path detection, and movement speed detection, to detect targets other than people whose facial images have been registered in advance (e.g., wheelchair users, white cane users, crutch users, sick people, injured people, lost children, other people requiring assistance, fallen objects, dropped objects, and other obstacles), and to detect the estimated locations of such objects. The video analysis server then returns the results to the transaction control server.

[0101] The transaction control server generates output information based on the results received from the facial recognition server and the video analysis server, and displays this information on the detection result confirmation terminal.

[0102] In this embodiment, the image processing device 10 is realized by a transaction control server, a video analysis server, and a facial recognition server.

[0103] The embodiments of the present invention have been described above with reference to the drawings, but these are illustrative examples of the present invention, and various other configurations can be adopted. The configurations of the embodiments described above may be combined with each other, or some configurations may be replaced with other configurations. Furthermore, the configurations of the embodiments described above may be modified in various ways without departing from the spirit of the invention. In addition, the configurations and processes disclosed in each of the embodiments and modifications described above may be combined with each other.

[0104] Furthermore, while the flowcharts used in the above description show multiple steps (processes) in sequence, the execution order of the steps performed in each embodiment is not limited to the order in which they are described. In each embodiment, the order of the illustrated steps can be changed to the extent that it does not impede the content. Also, the above embodiments can be combined to the extent that their contents do not conflict.

[0105] Some or all of the above embodiments may also be described as follows, but are not limited to the following: 1. A means for identifying at least one of the degree of crowding, movement patterns, and movement speed of people included in a series of images in a timeline, A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, is used to detect such locations within the image. An image processing device having 2. The image processing apparatus according to claim 1, wherein the presence estimation location detection means detects as the presence estimation location a location surrounded by an area where the degree of human congestion is equal to or greater than a first congestion criterion value, and where the degree of human congestion is less than a second congestion criterion value. 3. The image processing apparatus according to 2, wherein the presence estimation location detection means detects as the presence estimation location a location that is surrounded by an area where the degree of human congestion is equal to or greater than a first congestion criterion value, and where the state in which the degree of human congestion is less than a second congestion criterion value continues for a predetermined period of time or longer. 4. The image processing apparatus according to any one of 1 to 3, wherein the presence estimation location detection means detects a location that is shown in the specific result to be a place where no person has passed as the presence estimation location. 5. The image processing apparatus according to 4, wherein the presence estimation location detection means detects as the presence estimation location a location where the reference information indicates that a person passes through, but the specific result indicates that a person does not pass through, based on reference information indicating the normal movement of people and the specific result. 6. The image processing apparatus according to 4 or 5, wherein the presence estimation location detection means detects a location where no person has passed for a predetermined period of time or longer as the presence estimation location. 7. The image processing apparatus according to any one of 1 to 6, wherein the presence estimation location detection means detects a location where a person whose movement speed is less than a speed reference value is present as the presence estimation location. 8. The image processing apparatus according to 7, wherein the presence estimation location detection means calculates the speed reference value based on the movement speed of multiple people included in the image. 9. A detection target detection means for detecting the target from the image based on the characteristic features of the appearance of the target to be detected, An extraction means for extracting from the detected detection targets that satisfy predetermined conditions in relation to the detected location of estimated existence, An output means that outputs information linking the estimated location of existence with the detection target that satisfies the predetermined conditions between the estimated location of existence and the output means, An image processing apparatus according to any one of 1 to 8, having the following characteristics. 10. The aforementioned conditions are: The distance between the detection location of the object to be detected and the detection location of the estimated location is less than the distance reference value, The image processing apparatus according to claim 9, which includes at least one of the time differences between the detection timing of the target to be detected and the detection timing of the location where existence is estimated to be located, where the time difference between these two is less than a time reference value. 11. The image processing apparatus according to 9 or 10, wherein the predetermined condition includes that the estimated location of existence and the detected object are not simultaneously detected in the same image. 12. Computers, Identify at least one of the following in a series of images: the degree of crowding, movement patterns, and movement speed of people. An image processing method for detecting locations within an image where the presence of a target is estimated, based on the aforementioned specific results. 13. Computers, A means for identifying at least one of the degree of crowding, movement patterns, and movement speed of people included in a series of images in a time-series sequence, and A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, within the image. A program that makes it function as such. [Explanation of symbols]

[0106] 10 Image Processing Device 11 Specific section 12. Detection unit for estimated location 13 Output section 14 Detection Target Detection Unit 15 Extraction part 1A Processor 2A Memory 3A input / output I / F 4A Peripheral Circuits 5A Bus

Claims

1. A means for identifying at least one of the degree of crowding, movement patterns, and movement speed of people included in a series of images in a time-series sequence, A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, is used to detect such locations within the image. A detection target detection means for detecting the target from the image based on the characteristic features of the appearance of the target to be detected, An extraction means for extracting from the detected detection targets that satisfy predetermined conditions in relation to the detected location of estimated existence, An output means that outputs information linking the estimated location of existence with the detection target that satisfies the predetermined conditions between the estimated location of existence and the output means, An image processing device having

2. The image processing apparatus according to claim 1, wherein the presence estimation location detection means detects as the presence estimation location a location surrounded by an area where the degree of human congestion is equal to or greater than a first congestion criterion value and where the degree of human congestion is less than a second congestion criterion value.

3. The image processing apparatus according to claim 2, wherein the presence estimation location detection means detects as the presence estimation location a location that is surrounded by an area where the degree of human congestion is equal to or greater than a first congestion threshold value, and where the state in which the degree of human congestion is less than a second congestion threshold value continues for a predetermined period of time or longer.

4. The image processing apparatus according to claim 1, wherein the presence estimation location detection means detects a location where it has been shown in the specific result that no person has passed through as the presence estimation location.

5. The image processing apparatus according to claim 4, wherein the presence estimation location detection means detects as the presence estimation location a location where the reference information indicates that a person passes through, but the specific result indicates that a person does not pass through, based on reference information indicating the normal movement of people and the specific result.

6. The image processing apparatus according to claim 4 or 5, wherein the presence estimation location detection means detects a location where no person has passed for a predetermined period of time or longer as the presence estimation location.

7. The image processing apparatus according to claim 1, wherein the presence estimation location detection means detects locations where a person whose movement speed is less than a speed reference value is present as the presence estimation location.

8. The image processing apparatus according to claim 7, wherein the presence estimation location detection means calculates the speed reference value based on the movement speed of a plurality of people included in the image.

9. The aforementioned predetermined conditions are: The distance between the detection location of the object to be detected and the detection location of the estimated location is less than the distance reference value, The image processing apparatus according to claim 1, wherein the time difference between the detection timing of the target to be detected and the detection timing of the location where existence is estimated is less than a time reference value, at least one of these.

10. The image processing apparatus according to claim 1 or 9, wherein the predetermined condition includes the fact that the estimated location of existence and the detected object are not simultaneously detected in the same image.

11. Computers Identify at least one of the crowd density, movement patterns, and movement speed of people in multiple images that appear consecutively in time. Based on the aforementioned specific results, locations where the presence of the target is estimated are detected within the image. Based on the characteristic features of the appearance of the target to be detected, the target to be detected is detected from the image. From among the detected objects, the objects that satisfy predetermined conditions in relation to the detected estimated locations are extracted. An image processing method that outputs information linking the estimated location of existence with the detection target that satisfies the predetermined conditions between the estimated location of existence and the location of existence.

12. Computers, A means for identifying at least one of the degree of crowding, movement patterns, and movement speed of people included in a series of images in a time-series sequence, and A means for detecting locations where the presence of a target is estimated to be present, based on the aforementioned specific results, within the image. A detection means for detecting the target from the image based on the characteristic features of the appearance of the target to be detected. Extraction means for extracting from the detected detection targets those detection targets that satisfy predetermined conditions in relation to the detected location of estimated existence, Output means that outputs information linking the estimated location of existence with the detection target that satisfies the predetermined conditions between the estimated location of existence, A program that makes it function as such.