Person detection system
The person detection system uses multiple cameras and AI to enhance fall detection on escalators, addressing misinterpretation issues by luggage, ensuring accurate and timely responses to passenger falls.
Patent Information
- Application Number
- JP2024060313
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-03
- Publication Date
- 2025-10-16
AI Technical Summary
Existing systems struggle to accurately detect a passenger's fall on an escalator due to potential misinterpretation from luggage, leading to incorrect detection or failure to detect falls.
A person detection system utilizing multiple cameras and an AI model to analyze images from different angles, prioritizing images near entrances, and employing transfer learning to enhance fall detection accuracy, distinguishing between passengers and luggage.
Accurately detects passenger falls on escalators, enabling timely notifications and appropriate escalator control to prevent accidents.
Smart Images

Figure 2025157942000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a person detection system. [Background technology]
[0002] Patent Document 1 discloses a human detection system that detects people on a passenger conveyor based on distance information from a captured image of the passenger conveyor. The human detection system in Patent Document 1 also uses the distance information to determine whether a person has fallen. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 4884154 Summary of the Invention [Problem to be solved by the invention]
[0004] If a passenger's fall on an escalator could be accurately detected, the escalator could be appropriately controlled, such as by slowing down or stopping. However, accurately detecting a passenger's fall on an escalator can be difficult. For example, when detecting a passenger's fall using distance information, as in the human detection system disclosed in Patent Document 1, there is a possibility that a passenger's fall will not be detected or will be mistakenly detected due to luggage carried by another passenger.
[0005] The present disclosure has been devised in view of the above-described conventional situation, and aims to accurately detect a fall of a passenger on an escalator. [Means for solving the problem]
[0006] The present disclosure provides a person detection system including a plurality of cameras that capture images of an escalator, and an image processing device that acquires images captured by each of the plurality of cameras, detects a person from the captured images, determines whether the person has fallen, and, if it determines that the person has fallen, notifies the user that the person's fall has been detected.
[0007] Any combination of the above components, and conversion of the present disclosure into a method, device, system, storage medium, computer program, etc., are also valid aspects of the present disclosure. [Effects of the Invention]
[0008] According to the present disclosure, a fall of a passenger on an escalator can be detected with high accuracy. [Brief explanation of the drawings]
[0009] [Figure 1] FIG. 1 is a block diagram showing a configuration example of a human detection system according to a first embodiment. [Figure 2] FIG. 1 is a schematic diagram illustrating an example of an imaging range of each camera according to the first embodiment; [Figure 3] Schematic diagrams showing examples of images captured by each camera according to the first embodiment. [Figure 4] Schematic diagram for explaining generation of an AI model according to the first embodiment. [Figure 5] 1 is a flowchart showing a processing flow of an image processing device according to a first embodiment. [Figure 6] FIG. 10 is a table showing an example of determining the continuity of falls according to the first embodiment. [Figure 7] FIG. 1 is a schematic diagram illustrating an example of determining the continuity of a fall according to the first embodiment; DETAILED DESCRIPTION OF THE INVENTION
[0010] Hereinafter, with reference to the drawings as appropriate, detailed descriptions of embodiments specifically disclosing a human detection system according to the present disclosure will be provided. However, unnecessary detailed descriptions may be omitted. For example, detailed descriptions of well-known matters and redundant descriptions of substantially identical configurations may be omitted. This is to avoid unnecessary redundancy in the following description and to facilitate understanding by those skilled in the art. Note that the accompanying drawings and the following description are provided to enable those skilled in the art to fully understand the present disclosure, and are not intended to limit the subject matter recited in the claims.
[0011] (Embodiment 1) First, a configuration example of a human detection system 1 according to the first embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing a configuration example of a human detection system 1 according to the first embodiment.
[0012] The human detection system 1 is a system that, when an image processing device 10 detects a fall of a passenger on the escalator ESC1 based on an image of the escalator ESC1 captured by a camera such as the lower camera C1, supports appropriate response to the passenger's fall by controlling the speed of the escalator ESC1 as necessary. In this specification, a passenger on the escalator ESC1 may be referred to as a person.
[0013] The person detection system 1 includes a lower camera C1, a middle camera C2, an upper camera C3, an escalator ESC1, an image processing device 10, a recorder 20, a Power over Ethernet (hereinafter referred to as "PoE") hub 30, a broadcasting facility 40, a notification terminal 50, and a notification terminal 60.
[0014] The lower camera C1, the middle camera C2, and the upper camera C3 are each installed in the station and capture images of the escalator ESC1, which is also installed in the station. The lower camera C1, the middle camera C2, and the upper camera C3 may each be used as a surveillance camera. As will be described in detail later with reference to Figures 2 and 3, the lower camera C1 captures images of the lower entrance E1 of the escalator ESC1 and its vicinity. The upper camera C3 captures images of the upper entrance E2 of the escalator ESC1 and its vicinity. The middle camera captures images of the escalator ESC1 from above. For ease of explanation, this embodiment assumes that the escalator ESC1 is an ascending escalator. Therefore, the lower entrance E1 is the entrance to the escalator ESC1, and the upper entrance E2 is the exit to the escalator ESC1. In this specification, for the sake of simplicity, the lower camera C1, the middle camera C2, and the upper camera C3 may be referred to as the individual cameras.
[0015] The lower camera C1, the middle camera C2, and the upper camera C3 each include at least a lens (not shown) and an image sensor (not shown) as optical elements for generating captured images. The lens receives light reflected by an object within the field of view of the area captured by each camera and forms an optical image of the object on the light-receiving surface of the image sensor, i.e., the imaging surface. The image sensor is a solid-state imaging device such as a Charged Coupled Device (hereinafter referred to as "CCD") or a Complementary Metal Oxide Semiconductor (hereinafter referred to as "CMOS"). The image sensor converts the optical image formed on the imaging surface via the lens into an electrical signal at predetermined intervals. For example, if the predetermined interval is 1 / 30 of a second, the frame rate of each camera is 30 fps. Each camera may also generate captured images by performing predetermined signal processing on the electrical signal at the predetermined intervals. The captured images generated by each camera include still images and videos. The frame rates of the lower camera C1, the middle camera C2, and the upper camera C3 may be different from each other.
[0016] A line F1 is drawn at the end of each step of the escalator ESC1 as a caution sign. The line F1 is, for example, a yellow line, and may be drawn on all four sides or on three sides of the step.
[0017] The image processing device 10, the recorder 20, and the PoE hub 30 are installed in a station machine room.
[0018] The image processing device 10 is configured using a general-purpose computer device, for example, a personal computer, a server computer, etc. The image processing device 10 includes at least a processor 11, a memory 12, and an interface device 13.
[0019] The processor 11 is configured using, for example, a central processing unit (hereinafter referred to as "CPU"), a graphics processing unit (hereinafter referred to as "GPU"), a micro processing unit (hereinafter referred to as "MPU"), a digital signal processor (hereinafter referred to as "DSP"), or a field programmable gate array (hereinafter referred to as "FPGA"), etc. The processor 11 realizes the functions of the image processing device 10 by reading and executing various data and programs stored and held in the memory 12.
[0020] The memory 12 is a storage area for storing and holding various data, programs, etc. The memory 12 is composed of, for example, a read only memory (hereinafter referred to as "ROM"), which is a non-volatile storage area, a hard disk drive (hereinafter referred to as "HDD"), and a random access memory (hereinafter referred to as "RAM"), which is a volatile storage area. The RAM is, for example, a work memory used during operation of the image processing device 10. The ROM stores and holds, for example, programs for controlling the image processing device 10 in advance.
[0021] The interface device 13 is an interface for transmitting and receiving data or signals to and from an external device. The interface device 13 may support either wired communication or wireless communication. The communication method used by the interface device 13 may be, for example, a Wide Area Network (hereinafter referred to as "WAN"), a Local Area Network (hereinafter referred to as "LAN"), Long Term Evolution (hereinafter referred to as "LTE"), mobile communication such as 5G, power line communication, short-range wireless communication such as Wi-Fi (registered trademark) and Bluetooth (registered trademark), or a combination of these. In this embodiment, the PoE hub 30 and the interface device 13 of the image processing device 10 are connected by a network cable such as an Ethernet cable.
[0022] The processor 11 of the image processing device 10 transmits a signal to the broadcasting equipment 40 via the interface device 13, causing the broadcasting equipment 40 to convey information to people in the station, such as passengers on the escalator ESC1.
[0023] The processor 11 of the image processing device 10 controls the speed of the escalator ESC1 via the interface device 13. Note that the image processing device 10 may also control the speed of the escalator ESC1 via a control device (not shown).
[0024] The image processing device 10 may further include components other than the processor 11, the memory 12, and the interface device 13. For example, the image processing device 10 may further include a display device, an input device, and the like.
[0025] The recorder 20 records images captured by the lower camera C1, the middle camera C2, and the upper camera C3. The recorder 20 may be equipped with a display device. For example, a user operating the recorder 20 may be able to view the images captured by each camera and recorded in the recorder 20 in real time on a display device equipped in the recorder 20. The recorder 20 is connected to the PoE hub 30 by a network cable such as an Ethernet cable. An example of the recorder 20 is a network video recorder.
[0026] The PoE hub 30 supplies power to the lower camera C1, middle camera C2, upper camera C3, recorder 20, and image processing device 10 via a network cable such as an Ethernet cable. The PoE hub 30 obtains power from a power supply (not shown). By connecting the PoE hub 30 to each of the lower camera C1, middle camera C2, upper camera C3, recorder 20, and image processing device 10 via a network cable, each device can perform data communication via the PoE hub 30. Note that the recorder 20 and image processing device 10 may obtain power from a power supply (not shown).
[0027] The broadcasting equipment 40 is installed within the station. When the broadcasting equipment 40 receives a signal from the image processing device 10, it conveys information to people within the station, such as passengers on the escalator ESC1. The broadcasting equipment 40 may broadcast, for example, pre-recorded audio. An example of the broadcasting equipment 40 is a speaker.
[0028] The notification terminal 50 and the notification terminal 60 are configured using a general-purpose computer device, for example, a personal computer, a server computer, or the like. The notification terminal 50 and the notification terminal 60 receive a notification that a passenger has fallen, transmitted from the image processing device 10 that detected the passenger's fall on the escalator ESC1. This allows a user operating the notification terminal 50 and the notification terminal 60 to confirm that the passenger on the escalator ESC1 has fallen. The notification terminal 50 is installed in a station office. This allows station staff in the station office to confirm that the passenger has fallen and take subsequent action. The notification terminal 60 is installed at a monitoring base. The monitoring base is, for example, a base that collectively manages footage from monitoring cameras at each station. This allows a user at the monitoring base to know, for example, at which station the passenger fell.
[0029] Next, examples of the imaging ranges of the lower camera C1, the middle camera C2, and the upper camera C3 according to embodiment 1 will be described with reference to Fig. 2. Fig. 2 is a schematic diagram for explaining examples of the imaging ranges of the cameras according to embodiment 1.
[0030] The imaging range R1 of the lower camera C1 includes the lower entrance E1 of the escalator ESC1 and its vicinity. In this embodiment, since the escalator ESC1 is an upward escalator, the lower camera C1 is provided so as to be able to capture an image of the entrance of the escalator ESC1 from behind in the direction of travel of the escalator ESC1 as it ascends and descends.
[0031] The intermediate camera C2 is provided so as to be able to capture images of the escalator ESC1 from the lower entrance to the upper entrance from above the escalator ESC1. Simply put, the imaging range R2 of the intermediate camera C2 includes the entire escalator ESC1 as seen from above.
[0032] The imaging range R3 of the upper camera C3 includes the upper entrance E2 of the escalator ESC1 and its vicinity. In this embodiment, since the escalator ESC1 is an upward escalator, the upper camera C3 is provided so as to be able to capture an image of the entrance of the escalator ESC1 from ahead in the direction of travel when the escalator ESC1 is ascending or descending.
[0033] Hereinafter, a camera (e.g., lower camera C1) that is provided so as to be able to capture an image of the entrance to the escalator ESC1 from behind in the direction of travel when the escalator ESC1 is ascending or descending may be referred to as a first camera. Also, hereinafter, a camera (e.g., middle camera C2) that is provided so as to be able to capture an image of the escalator ESC1 from above may be referred to as a second camera. Note that the first camera and the second camera may each include multiple cameras.
[0034] The lower camera C1 and the upper camera C3 may be, for example, Panoramic Tilt Zoom (hereinafter referred to as "PTZ") cameras. The middle camera C2 may be, for example, a wide-angle camera equipped with a wide-angle lens. However, the lower camera C1, the middle camera C2, and the upper camera C3 are not limited to these.
[0035] Next, examples of images captured by the lower camera C1, the middle camera C2, and the upper camera C3 according to the first embodiment will be described with reference to Fig. 3. Fig. 3 is a schematic diagram showing examples of images captured by each camera according to the first embodiment.
[0036] In the example of Figure 3, person P1 is standing at the lower entrance E1 of escalator ESC1 with baggage G1. Also, person P2 has fallen on a step of escalator ESC1. At this time, person P1 appears in captured image I1 generated by lower camera C1 capturing an image of escalator ESC1. Also, person P2 appears in captured image I1 as if half of his body is obscured by person P1's baggage G1.
[0037] A captured image I2 generated by the intermediate camera C2 capturing an image of the escalator ESC1 shows people P1 and P2. Unlike the captured image I1, the captured image I2 shows person P2 without being obstructed by luggage G1 of person P1.
[0038] A captured image I3 generated by capturing an image of the escalator ESC1 with the upper camera C3 includes the upper entrance E2 and its vicinity, but in the example of FIG. 3, neither person P1 nor person P2 is captured.
[0039] In this embodiment, an AI model 70 (see FIG. 4 ), which will be described later, detects a person in the captured images generated by each camera and detects the person's fall based on the person's falling posture. As shown in the example of FIG. 3 , falls on an upward escalator are more likely to occur near the entrance. Therefore, images captured by a camera installed to capture the area near the entrance have a higher priority for fall detection than images captured by a camera installed to capture the area near the exit. However, as shown in the example of FIG. 3 , the fallen person P2 may be partially visible in the image I1 captured by the lower camera C1, which captures the lower entrance E1 and its vicinity. In this case, the AI model 70 can detect the person P2's fall based on the image I2 captured by the middle camera C2. On the other hand, in an image captured by a camera such as the middle camera C2 that captures the escalator ESC1 from above, it may be difficult to determine whether the person has fallen, for example, if the person has fallen while crouching. Therefore, it is preferable that the cameras for capturing images of the escalator ESC1 are installed at a plurality of locations so that the blind spots of the cameras that are different from each other can be covered.
[0040] Next, the generation of the AI model 70 according to the first embodiment will be described with reference to Fig. 4. Fig. 4 is a schematic diagram for explaining the generation of the AI model 70 according to the first embodiment.
[0041] In this embodiment, a trained AI model 70 is used to detect a passenger on escalator ESC1, i.e., a person falling. The AI model 70 may be generated by an arbitrary learning algorithm so that, when a captured image is input, the AI model 70 calculates a score corresponding to the falling posture of the person appearing in the captured image, determines whether the person has fallen based on the score, and outputs the score and the determination result. Note that the learning algorithm for generating the trained AI model 70 according to this embodiment is not particularly limited, and machine learning including deep learning techniques such as a convolutional neural network (hereinafter referred to as "CNN") may be used, for example.
[0042] For ease of explanation, it is assumed that the AI model 70 is generated by the image processing device 10. First, the image processing device 10 uses a large number of images of a person in a posture in which they have fallen (falling posture) and a large number of images of a person in a posture in which they have not fallen (non-falling posture) to train the model so that it can detect a person's fall. Then, the image processing device 10 generates the AI model 70 as a fall detection model.
[0043] The image processing device 10 retrains the generated AI model 70 as a fall detection model specialized for detecting falls on escalators. To retrain the AI model 70, the image processing device 10 prepares a large number of images of a person falling on an escalator step, images of the person falling on the escalator step wearing clothing that is similar in color to the escalator step, and images of luggage. The images of the person falling on the escalator step wearing clothing that is similar in color to the escalator step are images of factors that may cause an undetected fall. This is because, if the clothing of the person falling on the escalator step is similar in color to the escalator step, the AI model 70 may not be able to accurately detect the person's falling posture. The images of luggage are images of factors that may cause an erroneous detection of a fall. This is because, for example, the AI model 70 may mistakenly detect a large suitcase in the image as a person. The image processing device 10 retrains the AI model 70 so that it can deal with issues specific to escalators, such as falls on escalator steps, failure to detect falls, and false detection of falls.
[0044] The image processing device 10 retrains the AI model 70, for example, by transfer learning. As a result, the AI model 70 becomes a fall detection model specialized for falls on escalators. The AI model 70 can detect, for example, a person's fall on an escalator step. Furthermore, the AI model 70 can detect the fall of a person even if the person is wearing clothing of the same color as the escalator step and falls on the escalator step. Furthermore, the AI model 70 can distinguish, for example, between large luggage and a person without misidentifying them.
[0045] When a captured image is input, the AI model 70 detects a person appearing in the captured image. The AI model 70 then calculates a score corresponding to the person's falling posture. The score is calculated, for example, as a numerical value. The AI model 70 then compares the calculated score with a predetermined threshold to determine whether the person has fallen. For example, the AI model 70 may determine that the person has fallen if the calculated score is equal to or greater than a predetermined threshold. The threshold may be specified in advance. The AI model 70 outputs the calculated score and a determination result. The image processing device 10 detects that the person has fallen based on the output determination result.
[0046] Hereinafter, the determination of a person's fall by the AI model 70 may be referred to as a first determination. Furthermore, the determination of the continuity of a person's fall may be referred to as a second determination. The determination of the continuity of falls includes a determination of whether one or more people have been falling for a predetermined period of time or more. The determination of the continuity of falls may further include a determination of whether there is a possibility of a cause of a fall. The continuity of falls may be determined by the AI model 70 or by the processor 11 of the image processing device 10. When the AI model 70 determines the continuity of falls, the AI model 70 may determine the continuity of falls by detecting a person, calculating a score, and performing a first determination for each of a plurality of captured images captured within a predetermined period of time. When the processor 11 of the image processing device 10 determines the continuity of falls, the processor 11 may determine the continuity of falls based on the output of the AI model 70 in response to each input of a plurality of captured images captured within a predetermined period of time (in other words, the fall determination result). For the sake of convenience, it is assumed below that the continuity of falls is determined by the AI model 70. The determination of the continuity of falls (second determination) will be described later with reference to FIGS.
[0047] For ease of explanation, in this embodiment, the generated AI model 70 is stored in the memory 12 of the image processing device 10. However, this is not limiting, and the AI model 70 may be stored, for example, in an external device (not shown) accessible to the image processing device 10. Since the learning process generally imposes a high processing load, it is preferable that the process of generating the AI model 70 and the fall determination by the AI model 70 be performed at different times. Furthermore, from the perspective of processing distribution, the process of generating the AI model 70 and the fall determination by the AI model 70 may be configured to be performed by separate devices.
[0048] Next, a flow will be described with reference to Fig. 5, from when the image processing device 10 acquires a captured image of the escalator ESC1 to when it detects a fall and executes a notification in response to the fall detection (for example, a notification to the notification terminal 50). Fig. 5 is a flowchart showing the flow of processing by the image processing device 10 according to the first embodiment.
[0049] The processor 11 of the image processing device 10 acquires captured images of the escalator ESC1 from the lower camera C1, the middle camera C2, and the upper camera C3 (step St80). The processor 11 inputs the captured images acquired in step St80 to the AI model 70 stored in the memory 12.
[0050] The AI model 70 held in the image processing device 10 detects a person appearing in the captured image based on the input captured image (step St81). Note that the AI model 70 may detect multiple people in step St81.
[0051] When detecting a person in step St81, the AI model 70 may selectively use multiple cameras (e.g., the lower camera C1, the middle camera C2, and the upper camera C3) and captured images (e.g., captured image I1, captured image I2, and captured image I3) captured by each of the multiple cameras, or may preferentially use the image captured by the lower camera C1, for example. Preferentially using the image captured by the lower camera C1 may mean, for example, that when a person detection result based on an image captured by the lower camera C1 at a specific time differs from a person detection result based on an image captured by the middle camera C2 at the same specific time, the person detection result based on the image captured by the lower camera C1 is prioritized. This is because the escalator ESC1 is an upward escalator, and a fall accident is more likely to occur at the lower entrance E1 (i.e., the boarding entrance). Alternatively, the processor 11 of the image processing device 10 may preferentially input the image captured by the lower camera C1 to the AI model 70.
[0052] The AI model 70 scores the falling posture of the person detected in step St81 (step St82). Note that scoring is synonymous with calculating a score. Furthermore, in the above description of step St81, preferentially using the image captured by the lower camera C1 may mean, for example, the following. That is, when the score of the falling posture of the person captured in the image captured by the lower camera C1 at a specific time differs from the score of the falling posture of the person captured in the image captured by the middle camera C2 at the same specific time, the score of the falling posture of the person captured in the image captured by the lower camera C1 may be given priority.
[0053] The AI model 70 performs a first determination of whether the person detected in step St81 has fallen, based on the score calculated in step St82 (step St83). That is, the AI model 70 determines whether the person detected in step St81 has fallen.
[0054] If the first determination result is not "has fallen" (step St84; NO), that is, if the AI model 70 determines that the person detected in step St81 has not fallen, it outputs the determination result and the score calculated in step St82. The output result may be temporarily stored and held in the memory 12, for example. In this case, the image processing device 10 does not detect that the person has fallen based on the first determination result output from the AI model 70. Then, the image processing device 10 returns to step St80 and repeats the process.
[0055] If the first determination result is "has fallen" (step St84; YES), that is, if the AI model 70 determines that the person detected in step St81 has fallen, it executes a second determination of the continuity of falls (step St85). The AI model 70 performs the second determination of the continuity of falls based on a plurality of captured images taken within a predetermined time after the person fell. The AI model 70 may perform each of the processes from step St81 to step St84 on the plurality of captured images. A specific example of determining the continuity of falls will be described later with reference to FIGS. 6 and 7. At this time, the processor 11 of the image processing device 10 detects a fall based on the first determination result.
[0056] If the second determination result is not "a person has fallen or there is a cause for a fall" (step St86; NO), that is, if it is determined that none of the people detected in step St81 have fallen and there is no cause for a fall, the AI model 70 outputs the determination result. The processor 11 of the image processing device 10 notifies a station staff member that a fall has been detected based on the first determination result (step St87), and alerts passengers on the escalator ESC1, etc., via the broadcasting equipment 40 (step St88). At this time, for example, the station staff member is notified that a person has fallen. Since the image processing device 10 has not detected a fall based on the second determination result but has detected a person's fall based on the first determination result, the image processing device 10 notifies the station staff member of this fact. This allows the station staff member to check, for example, whether or not the person who fell is injured. The content of the notification to the station staff member may be set in advance. The image processing device 10 may change the content of the notification to the station staff based on, for example, the number of cameras that captured an image of the fallen person among the cameras. Furthermore, the image processing device 10 may change the content of the notification to the station staff based on the traveling direction of the escalator ESC1.
[0057] Furthermore, the broadcasting equipment 40 may, for example, warn passengers on the escalator ESC1 to be careful of accidents involving falls. The content of the broadcasting equipment 40 may be set in advance.
[0058] If the second determination result is "a person has fallen or there is a factor that causes a fall" (step St86; YES), that is, if it is determined that any person among the people detected in step St81 has fallen or that there is a factor that causes a fall, the AI model 70 outputs the determination result. Based on the determination result, the processor 11 of the image processing device 10 notifies a station staff member (step St87), alerts passengers on the escalator ESC1 via the broadcasting equipment 40 (step St88), and controls the speed of the escalator ESC1 (step St89). At this time, for example, the station staff member is notified that a person has fallen. Furthermore, for example, the broadcasting equipment 40 alerts passengers on the escalator ESC1 to the risk of a fall. Furthermore, for example, the image processing device 10 reduces the speed of the escalator ESC1. By reducing the speed of the escalator ESC1, the person detection system 1 can, for example, assist in rescuing the person who has fallen or in investigating and eliminating the factor that caused the fall.
[0059] The notification to the station staff is performed, for example, by the image processing device 10 sending a notification to the notification terminal 50. When the notification terminal 50 receives a notification from the image processing device 10, it may emit a sound, an alarm, a light, or the like to promptly notify the station staff of the occurrence of a fall accident. Furthermore, when the notification terminal 50 receives a notification from the image processing device 10, it may display, on a display device (not shown), an image captured at the time of the fall accident and real-time video of the scene of the fall accident. The image captured at the time of the fall accident may be an image captured before and after the fall accident. As a measure for when a station staff member is absent from the station office, the image processing device 10 may also send a notification that a fall has been detected to a notification terminal 60 installed at a monitoring station.
[0060] Next, examples of determining the continuity of falls and controlling the speed of an escalator will be described with reference to Figures 6 and 7. Figure 6 is a table showing an example of determining the continuity of falls according to the first embodiment.
[0061] First, a case where there is one passenger on the escalator ESC1 will be described. As a premise, the AI model 70 detects one person riding on the escalator ESC1. Furthermore, the AI model 70 determines in the first determination that the person has fallen.
[0062] If the person detected by the AI model 70 recovers within a predetermined time after falling, the AI model 70 determines in the determination of the continuity of the fall (second determination) that the person has not fallen. In this specification, recovery means returning from a fallen posture to a non-falling posture by standing up, etc. In this case, the processor 11 of the image processing device 10 does not control the speed of the escalator.
[0063] If a person detected by the AI model 70 does not recover within a predetermined time after falling, the AI model 70 determines that the person has fallen in determining the continuity of the fall. In this case, the processor 11 of the image processing device 10 stops the escalator ESC1 if the position of the person determined by the AI model 70 to have fallen is at or near the entrance to the escalator ESC1, or at or near the exit. This is to prevent the occurrence of an entrapment accident and a chain reaction of falling accidents. The position of the passenger on the escalator ESC1 may be determined by the AI model 70 or the processor 11 based on images captured by each camera.
[0064] If the position of the person determined by the AI model 70 to have fallen is neither at or near the entrance to the escalator ESC1 nor at or near the exit, i.e., if the person is in the middle of the escalator ESC1, the processor 11 of the image processing device 10 slows down the speed of the escalator ESC1. This is because if the person who has fallen is in the middle of the escalator ESC1, there is a possibility that the person can return to the starting position before reaching the exit without stopping the escalator ESC1.
[0065] Next, a case where there are multiple passengers on the escalator ESC1 will be described. As a premise, the AI model 70 detects a first person riding on the escalator ESC1, a second person following the first person, and a third person following the second person. Furthermore, in the first determination, the AI model 70 determines that the first person has fallen. Note that in the example of FIG. 6, when there are multiple passengers on the escalator ESC1, the control of the escalator speed does not depend on the positions of the passengers.
[0066] If the first person detected by the AI model 70 recovers within a predetermined time after falling, the AI model 70 determines in determining the continuity of the fall that the passengers on the escalator (here, the first person, the second person, and the third person) have not fallen. Note that it is assumed that the second person and the third person have not fallen. In this case, the processor 11 of the image processing device 10 does not control the speed of the escalator. For ease of explanation, the situation in which the first person recovers within a predetermined time after falling will be referred to as situation A.
[0067] If the second person detected by the AI model 70 recovers within a predetermined time after falling during or after situation A, the AI model 70 determines in determining the continuity of the fall that the passenger on the escalator has not fallen. Note that it is assumed that the third person has not fallen. Also, it is assumed that the first person does not fall again after recovering. Here, "during situation A" refers to the period from when the first person falls until when they recover. Also, "after situation A" refers to the period from when the first person recovers until when they get off the escalator ESC1. In this case, the processor 11 of the image processing device 10 does not control the speed of the escalator ESC1. For convenience of explanation, a situation in which the second person recovers within a predetermined time after falling during or after situation A is referred to as situation B.
[0068] If the third person detected by the AI model 70 recovers within a predetermined time after falling during or after situation B, the AI model 70 determines in determining the continuity of falls that there may be a factor that caused the fall. It is assumed that the first and second persons do not fall again after recovering. Here, "during situation B" refers to the period from when the second person falls until when they recover. Furthermore, "after situation B" refers to the period from when the second person recovers until when they get off the escalator ESC1. In this case, the processor 11 of the image processing device 10 stops the escalator ESC1. In this way, if the AI model 70 detects falls of multiple passengers on the escalator ESC1 in succession, it may determine that there may be a factor that caused the fall. This causes the processor 11 of the image processing device 10 to stop the escalator ESC1. This allows, for example, station staff to investigate the cause and take measures.
[0069] If the second person detected by the AI model 70 does not recover within a predetermined time after falling during or after situation A, the AI model 70 determines that the passenger on the escalator has fallen in determining the continuity of the fall. In this case, the processor 11 of the image processing device 10 stops the escalator ESC1.
[0070] Note that the determination of the continuity of falls and the control of the escalator speed are not limited to the example shown in FIG. 6. For example, when stopping the escalator ESC1, the image processing device 10 may control the speed of the escalator ESC1 so as not to suddenly stop the escalator ESC1. For example, the escalator ESC1 may have a function for gradually stopping, and the image processing device 10 may stop the escalator ESC1 gradually. This makes it possible to prevent a new fall accident from occurring as a reaction to the stopping of the escalator ESC1, for example. Alternatively, when stopping the escalator ESC1, the image processing device 10 may stop the escalator ESC1 after issuing a warning via the broadcasting equipment 40.
[0071] FIG. 7 is a schematic diagram for explaining an example of the continuity determination of a fall according to Embodiment 1. FIG. 7 shows the passage of time and events at specific times. The magnitude relationship of each specific time shown in FIG. 7 is t1 < t2 < t3 < t4 < t5 < t6 < t7 < t8 < t9. First, the fall of the person R riding on the escalator ESC1 will be described.
[0072] The person R falls at time t1. The image processing device 10 detects the fall of the person R at time t2. When the image processing device 10 detects the fall of the person R, it can grasp that the person R fell at time t1. This is because the fall of the person R is detected based on the captured image generated at time t1.
[0073] The person R returns from the fall at time t3. The image processing device 10 confirms at time t4 that the person R has returned. That is, the image processing device 10 confirms that the person R has returned before a predetermined time T has elapsed since the person R fell. More precisely, the AI model 70 determines that the person is not falling in the determination of the continuity of the fall, and the processor 11 of the image processing device 10 confirms the return of the person R based on the determination result. In this case, the image processing device 10 does not control the speed of the escalator ESC1.
[0074] Next, the fall of the person S riding on the escalator ESC1 will be described. The person S falls at time t6. The image processing device 10 detects the fall of the person S at time t7. The image processing device 10 cannot confirm the return of the person S at time t8. In other words, the image processing device 10 confirms that the person S has not returned before a predetermined time T has elapsed since the person S fell. More precisely, the AI model 70 determines that the person is falling in the determination of the continuity of the fall, and the processor 11 of the image processing device 10 confirms the fall of the person R during the predetermined time T based on the determination result. The person S returns at time t9, but since the image processing device 10 could not confirm the return of the person S within the predetermined time T, the speed of the escalator ESC1 is reduced.
[0075] (Modification of the first embodiment) In the first embodiment described above, the escalator ESC1 is an upward escalator. However, this is not limiting, and the escalator ESC1 may be a downward escalator. In this case, the upper entrance E2 of the escalator ESC1 is the entrance for the escalator ESC1, and the lower entrance E1 is the exit for the escalator ESC1. The lower camera C1 is provided so as to be able to capture an image of the exit for the escalator ESC1 from the front in the direction of travel when the escalator ESC1 ascends or descends. The upper camera C3 is provided so as to be able to capture an image of the entrance for the escalator ESC1 from the rear in the direction of travel when the escalator ESC1 ascends or descends. In this case, when detecting a person, the AI model 70 may selectively use multiple cameras (e.g., the lower camera C1, the middle camera C2, and the upper camera C3) and captured images (e.g., the captured image I1, the captured image I2, and the captured image I3) captured by each of the multiple cameras. Alternatively, the AI model 70 may preferentially use images captured by the upper camera C3, for example, to detect falls, because falls on a descending escalator are more likely to occur at the upper entrance E2.
[0076] Furthermore, in the above-described first embodiment, an example was shown in which the AI model 70 held in the image processing device 10 detects a person from a captured image, calculates a score corresponding to the person's falling posture, and determines whether the person has fallen based on the calculated score and a predetermined threshold. However, this is not limited to this, and if there is a location in the captured image where a caution sign for a step of the escalator ESC1 cannot be detected, the AI model 70 may determine that a person has fallen at that location.
[0077] Furthermore, the processor 11 of the image processing device 10 may make a final determination, for example, as to whether or not a person has fallen, based on multiple determination results by the AI model 70. A specific example will be described below. For example, assume that five captured images are generated per second by the lower camera C1. The AI model 70 then detects a person and determines whether or not the person has fallen for each of the five captured images. For example, the AI model 70 determines that the person has fallen for three of the five captured images. Then, the AI model 70 determines that the person has not fallen for two of the five captured images. In this case, the processor 11 may make a final determination as to whether or not the person has fallen based on multiple determination results by the AI model 70 and either an AND condition or an OR condition.
[0078] When the processor 11 of the image processing device 10 makes a final determination of whether a person has fallen based on an AND condition, if all of the multiple determination results by the AI model 70 are "has fallen," the processor 11 determines that the person has fallen. When the processor 11 makes a final determination of whether a person has fallen based on an AND condition, the processor 11 determines that the person has not fallen because the AI model 70 has determined that two of the five captured images are "has not fallen."
[0079] When making a final determination of whether a person has fallen based on an OR condition, the processor 11 of the image processing device 10 determines that the person has fallen if at least one of the multiple determination results by the AI model 70 is "has fallen." When making a final determination of whether a person has fallen based on an OR condition, the processor 11 determines that the person has fallen because the AI model 70 has determined that three of the five captured images are "has fallen."
[0080] When the processor 11 of the image processing device 10 makes a final determination of a fall based on an AND condition, it is possible to reduce the number of undetected falls. Also, when the processor 11 of the image processing device 10 makes a final determination of a fall based on an OR condition, it is possible to reduce the number of false detections of a fall.
[0081] In the above specific example, multiple fall determination results based on images captured by the lower camera C1 are used to make a final fall determination based on an AND condition or an OR condition. However, this is not limited to this example. Multiple fall determination results based on images captured by the lower camera C1, the middle camera C2, and the upper camera C3 may also be used to make a final fall determination based on an AND condition or an OR condition. For example, if there is only one person riding on the escalator ESC1, the detection of that person will not be hindered by others. In this case, the processor 11 may input all of the images captured by each camera into the AI model 70 and make a final fall determination based on the multiple output determination results and an AND condition or an OR condition. Also, for example, if there are multiple people riding on the escalator ESC1, even if one of the people falls, the detection of the fallen person may be hindered by others. In this case, the processor 11 may input the images captured by the middle camera C2 into the AI model 70 and make a final fall determination based on the multiple output determination results and an AND condition or an OR condition. This is because the intermediate camera C2 captures an image of the escalator ESC1 from above, and therefore has a high possibility of capturing an image of the fallen person without being obstructed by others.
[0082] Summary of the Disclosure The above description of the first embodiment discloses at least the following techniques. Note that the components corresponding to the first embodiment are shown in parentheses, but the present invention is not limited to these.
[0083] <Technology 1> A person detection system (e.g., person detection system 1) includes a plurality of cameras (e.g., lower camera C1, middle camera C2, and upper camera C3) that capture images of an escalator (e.g., escalator ESC1), and an image processing device (e.g., image processing device 10) that acquires images (e.g., captured image I1, captured image I2, and captured image I3) captured by each of the plurality of cameras, detects a person (e.g., person P2) from the captured images, determines whether the person has fallen, and, if it is determined that the person has fallen, notifies the user that a person's fall has been detected.
[0084] As a result, the image processing device included in the person detection system can detect a person riding an escalator based on images of the escalator captured by multiple cameras and determine whether the person has fallen.The image processing device can then detect the person's fall based on the determination result.As a result, the image processing device can notify the determination result to, for example, station staff, etc.
[0085] <Technology 2> In the person detection system 1 described in Technology 1, the multiple cameras include at least a first camera (e.g., lower camera C1 or upper camera C3) that is arranged to be able to capture an image of the entrance to the escalator from behind in the direction of travel when the escalator is going up or down, and a second camera (e.g., middle camera C2) that is arranged to be able to capture an image of the escalator from above.
[0086] As a result, the first camera captures an image of the escalator entrance, allowing the image processing device to more accurately detect falls at the entrance. Also, the second camera captures an image of the escalator from above, so the image processing device can prevent a person's fall from going undetected, even if, for example, a crowd or other obstacle prevents the first camera from capturing an image of the person who has fallen.
[0087] <Technology 3> In the person detection system 1 described in Technique 2, the second camera is one or more cameras that capture images from the lower entrance (e.g., lower entrance E1) to the upper entrance (e.g., upper entrance E2) of the escalator.
[0088] This allows the second camera to capture images from the lower entrance to the upper entrance of the escalator, so the image processing device can prevent a person's fall from going undetected.
[0089] <Technology 4> In the person detection system described in any one of Techniques 1 to 3, if there is a location in the captured image where a warning sign (e.g., line F1) on an escalator step cannot be detected, the image processing device determines that a person has fallen at that location.
[0090] This allows the image processing device to detect a person's fall even if it is unable to detect a person who has fallen, for example, if the person has fallen in such a way that a warning sign on an escalator step is hidden from the camera's field of view.
[0091] <Technology 5> In the person detection system described in any one of Techniques 1 to 4, the image processing device changes the notification content based on the number of cameras among multiple cameras that captured an image showing the person who has been detected to have fallen.
[0092] This allows the image processing device to change the content of the notification to station staff and the like depending on the number of cameras that captured the image showing the person who has been detected to have fallen.
[0093] <Technology 6> In the human detection system described in any one of Techniques 1 to 5, when the image processing device detects a person falling, it controls the speed of the escalator and changes the notification content and speed control content based on the direction of travel.
[0094] This allows the image processing device to control the speed of the escalator when it detects a person falling. For example, when it detects a person falling, it can slow down the speed of the escalator. Furthermore, the image processing device can change the content of the notification to station staff, etc., depending on the direction of travel of the escalator. Furthermore, the image processing device can change the content of the control of the escalator speed depending on the direction of travel of the escalator.
[0095] <Technology 7> In the person detection system described in any one of Techniques 1 to 6, if the image processing device determines that a person has fallen at the entrance or exit of the escalator and has not returned to normal within a predetermined time (e.g., a predetermined time T), the image processing device stops the escalator.
[0096] This allows the image processing device to stop the escalator if a person who has fallen at the entrance or exit of the escalator does not recover within a predetermined time, thereby preventing the person who has fallen from getting caught in the escalator.
[0097] <Technology 8> In the person detection system described in any one of Techniques 1 to 7, if the image processing device determines that a person has fallen at a location other than the entrance or exit of the escalator and has not recovered within a predetermined time, the image processing device reduces the speed of the escalator.
[0098] This allows the image processing device to, for example, slow down the speed of the escalator if a person who has fallen in the middle of the escalator does not return to their normal position within a predetermined time. This allows the image processing device to assist the person in returning to their normal position. Furthermore, the image processing device can prevent further falls by stopping the escalator.
[0099] <Technology 9> In the person detection system described in any one of Techniques 1 to 8, when a person falls and then a person following that person also falls, the image processing device determines that a fall factor exists and stops the escalator.
[0100] This allows the image processing device to determine that the cause of the fall exists on the escalator when multiple passengers fall consecutively, allowing station staff, for example, to take action to address the cause of the fall.
[0101] <Technology 10> In the person detection system described in any one of Techniques 1 to 9, the image processing device inputs the captured image to an AI model (e.g., AI model 70) that calculates a score corresponding to the falling posture of a person appearing in the captured image, determines whether the person has fallen based on the score, and outputs the score and the determination result, and detects that the person has fallen based on the determination result output from the AI model.
[0102] This allows the image processing device to use the AI model to detect falls by passengers on escalators. By specializing the AI model in advance for detecting falls on escalators, the image processing device can detect falls by passengers on escalators with greater accuracy.
[0103] <Technology 11> In the person detection system described in Technology 10, when the AI model determines that a person in a captured image has fallen, it determines the continuity of the fall, that is, whether the person has recovered from the fall within a predetermined time, based on the captured image taken within a predetermined time after the person's fall; The image processing device detects the person's fall based on the results of the fall continuity determination output from the AI model.
[0104] This allows the image processing device to use the AI model to detect a fall by a passenger on an escalator based on whether the passenger recovers within a specified time after falling.
[0105] <Technology 12> In the person detection system described in Technology 10 or 11, the AI model detects a person and determines whether the person has fallen for each of a plurality of captured images, and the image processing device detects whether the person has fallen using an AND condition of whether the plurality of determination results output from the AI model are all "having fallen," or an OR condition of whether at least one of the plurality of determination results is "having fallen."
[0106] This allows the image processing device to use the AI model to detect a fall of a passenger on an escalator based on multiple fall determination results based on multiple captured images and an AND condition or an OR condition. This allows the image processing device to detect a fall of a passenger on an escalator when all of the multiple fall determination results based on multiple captured images are "falling." Alternatively, the image processing device can detect a fall of a passenger on an escalator when at least one of the multiple fall determination results based on multiple captured images is "falling."
[0107] <Technology 13> In the person detection system described in any one of Techniques 10 to 12, the AI model detects people and determines whether the people have fallen in each of the multiple images captured by all of the multiple cameras or a specific camera, and the image processing device detects whether the person has fallen using an AND condition of whether the multiple determination results output from the AI model are all "having fallen," or an OR condition of whether at least one of the multiple determination results is "having fallen."
[0108] This allows the image processing device to detect a fall of a passenger on an escalator based on multiple determination results based on multiple images captured by a specific camera among the multiple cameras and an AND condition or an OR condition. Alternatively, the image processing device can detect a fall of a passenger on an escalator based on multiple determination results based on multiple images captured by all of the multiple cameras and an AND condition or an OR condition.
[0109] Although the present embodiment has been described above with reference to the drawings, it goes without saying that the present disclosure is not limited to such examples. It is clear that a person skilled in the art can conceive of various modifications, alterations, substitutions, additions, deletions, and equivalents within the scope of the claims, and it is understood that these also naturally fall within the technical scope of the present disclosure. Furthermore, the components of the above-described present embodiment may be combined in any manner as long as they do not deviate from the spirit of the invention. [Industrial Applicability]
[0110] The present disclosure is useful as a person detection system. [Explanation of symbols]
[0111] 1. Person detection system 10 Image processing device 11 processors 12 Memory 13 Interface Device 20 Recorder 30 PoE Hub 40 Broadcasting Equipment 50, 60 notification terminals 70 AI models C1 Lower Camera C2 intermediate camera C3 upper camera ESC1 Escalator E1 Lower Entrance E2 Upper Entrance F1 Line G1 Luggage I1, I2, I3 images P1, P2 people R1, R2, R3 imaging range T predetermined time
Claims
1. A plurality of cameras for capturing images of the escalator; an image processing device that acquires images captured by each of the plurality of cameras, detects a person from the captured images, determines whether the person has fallen, and, if it is determined that the person has fallen, notifies the person that a fall has been detected; People detection system.
2. The plurality of cameras include at least a first camera that is provided so as to be able to capture an image of the entrance of the escalator from behind in the direction of travel when the escalator is ascending or descending, and a second camera that is provided so as to be able to capture an image of the escalator from above. The person detection system of claim 1 .
3. The second camera is one or more cameras that capture images from a lower entrance to an upper entrance of the escalator. The human detection system of claim 2 .
4. When a warning sign for a step of the escalator cannot be detected at a location in the captured image, the image processing device determines that the person has fallen at that location. The person detection system of claim 1 .
5. the image processing device changes the notification content based on the number of cameras among the plurality of cameras that captured an image showing the person for whom a fall has been detected. The person detection system of claim 1 .
6. When the image processing device detects that the person has fallen, the image processing device controls the speed of the escalator; changing the notification content and the speed control content based on the traveling direction; The person detection system of claim 1 .
7. When the image processing device determines that the person has not recovered within a predetermined time after falling at the entrance or exit of the escalator, the image processing device stops the escalator. The person detection system of claim 1 .
8. When the image processing device determines that the person has not recovered within a predetermined time after falling at a location other than the entrance or exit of the escalator, the image processing device reduces the speed of the escalator. The person detection system of claim 1 .
9. When a person following the person falls after the person falls, the image processing device determines that a cause of a fall exists and stops the escalator. The person detection system of claim 1 .
10. The image processing device includes: inputting the captured image into an AI model that calculates a score corresponding to the falling posture of the person appearing in the captured image, determines whether the person has fallen based on the score, and outputs the score and the determination result; Detecting a fall of the person based on the determination result output from the AI model. The person detection system of claim 1 .
11. When the AI model determines that the person in the captured image has fallen, it determines the continuity of the fall, i.e., whether the person has recovered from the fall within a predetermined time, based on captured images captured within the predetermined time after the person has fallen; The image processing device detects the fall of the person based on the determination result of the continuity of the fall output from the AI model. The person detection system of claim 10.
12. The AI model detects a person and determines whether the person has fallen for each of a plurality of captured images, The image processing device detects a fall of the person based on an AND condition of whether or not all of the plurality of determination results output from the AI model are "falling," or an OR condition of whether or not at least one of the plurality of determination results is "falling." The person detection system of claim 11 .
13. The AI model detects a person and determines whether the person has fallen in each of a plurality of images captured by all of the plurality of cameras or a specific camera, The image processing device detects a fall of the person based on an AND condition of whether or not all of the plurality of determination results output from the AI model are "falling," or an OR condition of whether or not at least one of the plurality of determination results is "falling." The person detection system of claim 12.
Citation Information
Patent Citations
JP1973084154A