Generating device
The generation device addresses the challenge of generating teacher data for detecting pedestrians in specific postures by tracking and estimating shape changes in moving images from surveillance cameras, resulting in efficient and accurate teacher data generation for machine learning models.
Patent Information
- Application Number
- JP2024059469
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-04-02
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-04-02
AI Technical Summary
The challenge is to easily generate teacher data for training machine learning models to detect pedestrians in specific postures, particularly those with difficulty moving, due to low occurrence frequency of such events and the effort required for human visual judgment.
A generation device that inputs moving images from surveillance cameras, trackingly detects image portions of pedestrians, estimates difficulty in moving based on shape changes, and stores determined image portions as teacher data.
Enables the easy generation of a large number of teacher data, facilitating the training of machine learning models to detect pedestrians in specific postures, such as difficulty in moving, with improved accuracy and efficiency.
Smart Images

Figure 0007684465000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a generation device and the like.
Background Art
[0002] At railroad crossings, within station premises, on roads, etc., surveillance cameras are often installed for the purpose of ensuring the safety of pedestrians and keeping watch over them. When a human tries to view and judge the captured images from these surveillance cameras, there are problems such as requiring a lot of manpower and a lot of time because there are many locations to be monitored. In recent years, the progress of image recognition technology has been remarkable. As an image recognition technology using AI (Artificial Intelligence), there is a technology called annotation that detects an object shown in an image, recognizes its type, and presents it in association with the range and type (name) where the object is shown (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Using such image recognition technology, a technology for quickly and mechanically detecting any abnormal situation that has occurred to a pedestrian from the captured images of surveillance cameras can be considered. To achieve this, it is necessary to train a machine learning model (such as an AI model) using an image of a situation to be detected, such as a pedestrian being in a specific appearance, as teacher data. However, the situations to be detected are often situations with a very low occurrence frequency that is different from normal, and it has been difficult to prepare a large number of necessary teacher data. That is, when extracting an image of the corresponding situation from actual captured images and using it as teacher data, it is difficult to prepare a large number of actual captured images as the extraction source because the occurrence frequency of the situation is low, and it also requires a great deal of effort for a human to visually judge and extract the corresponding situation.
[0005] The problem to be solved by the present invention is to enable easy generation of teacher data for training a machine learning model that detects pedestrians in a specific posture.
Means for Solving the Problem
[0006] A first invention for solving the above problem is a generation device (for example, the generation device 1 in FIG. 7) for teacher data for training a machine learning model that detects a person with difficulty in moving from a captured image of a monitoring camera that monitors a monitoring target location where a pedestrian passes, a moving image input means (for example, the moving image input unit 202 in FIG. 7) that inputs a moving image of a place where people come and go, a tracking detection means (for example, the tracking detection unit 204 in FIG. 7) that trackingly detects an image portion of a person shown in the moving image, an estimation means (for example, the estimation unit 206 in FIG. 7) that estimates whether or not the person shown in the image portion has reached a state of difficulty in moving based on a change in the shape of the image portion detected by the tracking detection means, a teacher data image storage control means (for example, the teacher data image storage control unit 208 in FIG. 7) that performs control to store the image portion positively determined by the estimation means in a predetermined storage unit as the teacher data, and a generation device comprising the above.
[0007] According to the first invention, teacher data for training a machine learning model that detects a person in a specific posture can be easily generated. That is, an image portion of a person shown in a moving image is trackingly detected, and an image portion showing a person estimated to have reached a specific posture such as difficulty in moving based on a change in the shape of the image portion can be generated and stored as teacher data. When a pedestrian reaches a state of difficulty in moving due to a fall or the like, the body posture often changes, so it can be estimated that the person has reached a state of difficulty in moving from the change in the shape of the image portion showing the person.
[0008] In addition, since the image portion in which a person is depicted is used as teacher data, the passing moving image may be an image taken at a location different from the location to be monitored. Therefore, it is easy to acquire and input a large number of passing moving images, and as a result, it becomes possible to easily generate a large number of teacher data.
[0009] A second invention is the invention described above, wherein the estimation means performs the estimation using, as the shape change, that the aspect ratio of the image portion has changed to a predetermined value or more. It is a generation device.
[0010] According to the second invention, it is possible to estimate that a person has reached a state of difficulty in moving by using, as a shape change, a change in the aspect ratio of the image portion in which the detected person is depicted. The image portion in which a person is depicted is vertically long while the person is walking, but often changes to a horizontally long shape when the person reaches a state of difficulty in moving due to a fall or the like. Therefore, it is possible to estimate that the person has reached a state of difficulty in moving due to a fall or the like from the change in the aspect ratio of the image portion.
[0011] A third invention is the invention described above, wherein the pedestrians include pedestrians and bicycle riders, the tracking detection means discriminates between a person and a bicycle shown in the passing moving image and performs tracking detection, the estimation means further estimates whether or not a bicycle rider shown in the image portion has reached a state of difficulty in moving based on a shape change of the image portion detected by the tracking detection means. It is a generation device.
[0012] According to the third invention, it is possible to detect that, in addition to pedestrians, a bicycle rider such as a fallen bicycle has reached a state of difficulty in moving.
[0013] A fourth invention is a machine learning model generation device that generates the machine learning model by causing the machine learning model to learn using the teacher data stored in the storage unit by the generation device of the above invention.
[0014] According to the fourth invention, a machine learning model for detecting a mobility-impaired pedestrian in a state of being mobility-impaired can be easily generated from a captured image of a surveillance camera that monitors a surveillance target location where pedestrians pass.
Brief Description of the Drawings
[0015]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Modes for Carrying Out the Invention
[0016] Hereinafter, preferred embodiments of the present invention will be described with reference to the drawings. Note that the applicable forms of the present invention are not limited to the following embodiments. Also, in the description of the drawings, the same reference numerals are assigned to the same elements.
[0017] This embodiment generates teacher data for training a machine learning model that detects mobility-impaired pedestrians in a state of difficulty moving from a captured image of a surveillance camera that monitors a surveillance target location where pedestrians pass. The surveillance target locations are places where an unspecified number of pedestrians pass, such as railroad crossings, within train stations, and roads. The surveillance camera is installed for the purpose of monitoring pedestrians in such surveillance target locations. A mobility-impaired pedestrian in a state of difficulty moving, which is the detection target by the machine learning model, means a pedestrian who has fallen due to tripping or a cyclist, or a pedestrian who is hunched due to poor physical condition, etc., and is considered to require some assistance. Therefore, the teacher data for training the machine learning model is an image of such a mobility-impaired pedestrian.
[0018] Figure 1 is a flowchart showing the flow of the process of generating teacher data performed by the teacher data generation device 1 of this embodiment. As shown in Figure 1, the generation device 1 first acquires and inputs a moving image of people coming and going captured at a location (step S1). The location where the moving image of people coming and going is captured may be the same location as the surveillance target location or a different location. Also, it does not have to be a real-time moving image and may be a recorded video. Also, since the occurrence frequency of mobility-impaired pedestrians who become teacher data is low, it is desirable to acquire and input a large amount of moving images of people coming and going.
[0019] Next, the image portions of the people shown in the acquired and input moving image of people coming and going are detected in a tracking manner (step S3). This detection can utilize known image recognition techniques. For example, by using a technique called annotation, objects shown in the moving image of people coming and going can be detected, and the type of the object (e.g., "person" or "bicycle") can also be recognized, and it can be obtained as a recognition result in association with the region (image portion) in the moving image of people coming and going where the object is shown. Also, the tracking detection of the image portions of people is performed for the period from the time when the person enters the screen to the time when the person exits the screen.
[0020] FIG. 2 shows an example of the detection result of a person shown in the passing video image, and shows an image (frame image) of one frame of the passing video image. The example shown in FIG. 2 is a passing video image taken at a place where vehicle passage is prohibited, and four people are detected. That is, three pedestrians M1 to M3 and one cyclist M4. And a rectangular frame showing the image portion where the detected person appears is displayed overlaid on the image. These individuals are distinguished, recognized as being of the type "person", assigned an identification number (object ID), and detected in a tracking manner.
[0021] Returning to FIG. 1, subsequently, it is estimated whether or not the person being detected in a tracking manner has reached a state of difficulty in moving (step S5). This estimation is performed, for example, based on the change in the shape of the image portion where the detected person appears.
[0022] Generally, when a walking person (pedestrian) is detected in a tracking manner, the shape of the image portion where the person appears is a vertically long rectangular shape, and the position of the image portion changes over time. However, if the person falls down on the way, the shape of the image portion changes from a vertically long rectangular shape to a horizontally long rectangular shape. For this reason, it is estimated that the person has reached a state of difficulty in moving because the shape of the image portion where the detected person appears has changed from a vertically long rectangular shape to a horizontally long rectangular shape, that is, the aspect ratio of the image portion has changed by a predetermined amount or more. For example, if the aspect ratio R of the image portion is defined as R = horizontal length W / vertical length H, when the image portion changes from a vertically long rectangular shape to a horizontally long rectangular shape, the aspect ratio R increases. In this case, it is estimated that the detected person has reached a state of difficulty in moving because the aspect ratio R (= horizontal length W / vertical length H) of the image portion has changed by a predetermined value (for example, doubled) or more.
[0023] Figure 3 is an example of a frame image in which the person shown in the moving video has reached a state of difficulty in moving, and is a frame image several frames after the image shown in Figure 2. In Figure 3, the pedestrian M1 detected in the frame image of Figure 2 has reached a state of difficulty in moving due to falling. As shown in Figure 4, the shape of the image portion G1 of the walking pedestrian M1 in the frame image shown in Figure 2 was a vertically long rectangular shape with a vertical length H1 and a horizontal length W1. Then, as shown in Figure 5, the shape of the image portion G1 of the fallen pedestrian M1 in the frame image shown in Figure 3 has changed to a horizontally long rectangular shape with a vertical length H2 and a horizontal length W2. That is, the aspect ratio R of the image portion has changed significantly from the aspect ratio R1 = W1 / H1 to the aspect ratio R2 = W2 / H2.
[0024] In addition, when the aspect ratio of the image portion of the person being detected in a tracking manner changes by a predetermined amount or more, and the shape of the image portion after the change continues for a predetermined time (predetermined number of frames), it may be estimated that the person has reached a state of difficulty in moving. Since it is a place where many people are coming and going, it is possible to avoid the possibility of false detection due to the influence of other people and improve the accuracy of the estimation. Also, a condition such as the position of the image portion not changing after the change in shape may be added.
[0025] Returning to Figure 1, if it is estimated that the person shown in the image portion being detected in a tracking manner has reached a state of difficulty in moving (step S7: YES), the image portion is extracted to generate teacher data (step S9).
[0026] FIG. 6 is a diagram for explaining the generation of teacher data. The upper part of FIG. 6 shows the images (frame images) of each consecutive frame constituting the moving image. Also, the image portion of the person detected in a tracking manner is shown as a dashed rectangle in each frame image. In the example of FIG. 6, the shape of the image portion of the person detected in a tracking manner was a vertically long rectangular shape in the frame images up to the (N - 1)-th frame, but changes to a horizontally long rectangular shape in the frame images from the next N-th frame onward. That is, in the frame image of the N-th frame, the shape of the image portion in which the detected person is shown has changed, and it is estimated that the person has reached a state of difficulty in moving. Starting from the point in time when it is estimated that the state of difficulty in moving has been reached, each frame image corresponding to a predetermined period (predetermined number of frames), that is, the frame images in which the person shown is in a state of difficulty in moving, is extracted. Then, for each of the extracted frame images, the image portion in which the person is shown is cut out and associated with the object ID and type (name) of the person to be used as teacher data.
[0027] Note that there may be a case where, as in the case where a person who has fallen stands up immediately and starts walking again, the person has reached a state of difficulty in moving but the state is immediately resolved. For this reason, if, after the point in time when it is estimated that the state of difficulty in moving has been reached, the shape of the image portion in which the detected person is shown changes (returns) to the shape before reaching the state, the frame images up to immediately before that may be extracted to generate teacher data.
[0028] FIG. 7 is a block diagram showing an example of the functional configuration of the teacher data generation device 1. As shown in FIG. 7, the generation device 1 includes an operation unit 102, a display unit 104, a communication unit 106, a processing unit 200, and a storage unit 300, and is configured as a kind of computer system.
[0029] The operation unit 102 is implemented by an input device such as a button switch or a touch panel, and outputs an operation signal corresponding to the performed operation to the processing unit 200. The display unit 104 is implemented by a display device such as an LCD (Liquid Crystal Display) or a touch panel, and performs various displays according to the display signal from the processing unit 200. The communication unit 106 is implemented by a wireless or wired communication device, and communicates with an external device via a given communication network.
[0030] The processing unit 200 is implemented by an arithmetic device such as a CPU (Central Processing Unit), and based on programs, data, etc. stored in the storage unit 300, gives instructions to and transfers data to each part constituting the generation device 1, and performs overall control of the generation device 1.
[0031] Further, the processing unit 200 executes the teacher data generation program 302 stored in the storage unit 300 to perform processing for generating teacher data for training a machine learning model that detects a mobility-impaired pedestrian in a mobility-impaired state from a captured image of a surveillance camera that monitors a surveillance target location where pedestrians pass (see FIG. 1). The processing unit 200 has, as functional processing blocks, a passing moving image input unit 202, a tracking detection unit 204, an estimation unit 206, and a teacher data image storage control unit 208. Each of these functional units of the processing unit 200 can be realized either software-wise by the processing unit 200 executing a program or by a dedicated arithmetic circuit. In the present embodiment, it will be described as being realized software-wise in the former case.
[0032] The passing moving image input unit 202 inputs a passing moving image obtained by photographing a location where people pass. Specifically, for example, it can be input from outside the device via the communication unit 106. Alternatively, it may be input by reading out moving image data stored in an external storage medium. The input passing moving image is stored as passing moving image data 310.
[0033] The tracking detection unit 204 trackingly detects the image portions of the people shown in the passing video image. Specifically, by using known image recognition technology, the objects shown in the passing video image are detected, and the types of the objects (for example, "person" or "bicycle") are also recognized, and the regions (image portions) in the passing video image where the objects are shown can be obtained as recognition results in association therewith. The trackingly detecting of the image portions of the people is performed for the period from the time when the person enters the screen to the time when the person exits the screen. Further, a plurality of each person shown in the passing video image is separately detected, and for each detected person, the image (frame image) in which the person is shown and the image portion in which the person is shown in each of these images are associated with an identification number (object ID) and generated and stored as tracking detection data 320 (see FIG. 2).
[0034] The estimation unit 206 estimates whether or not the person shown in the image portion has reached a state of difficulty in moving based on the change in the shape of the image portion detected by the tracking detection unit 204. For example, the estimation is performed using, as the shape change, the fact that the aspect ratio of the image portion has changed by a predetermined value or more.
[0035] Specifically, based on the tracking detection data 320 which is the result of the detection by the tracking detection unit 204, for each detected person, it is estimated whether or not the person has reached a state of difficulty in moving based on the change in the shape of the image portion in which the person is shown. For example, if the aspect ratio R of the image portion is defined as R = horizontal length W / vertical length H, then when the image portion changes from a vertically long rectangular shape to a horizontally long rectangular shape, the aspect ratio R increases. In this case, it is estimated that the person being detected has reached a state of difficulty in moving due to the aspect ratio R (= horizontal length W / vertical length H) of the image portion having changed by a predetermined value (for example, doubled) or more (see FIGS. 2 to 5).
[0036] The image memory control unit 208 for teacher data performs control to store the image portion positively determined by the estimation unit 206 in the storage unit 300 as teacher data. Specifically, from the frame images in which a person estimated to have reached a difficult-to-move posture by the estimation unit 206 is captured, frame images for a predetermined time (predetermined number of frames) after the estimated time are extracted. Then, from each of the extracted frame images, the image portion in which the person is captured is cut out, and is stored and accumulated in the storage unit 300 as teacher data 330 in association with the identification number (object ID) of the person and the type (“person”) (see FIG. 6).
[0037] The teacher data 330 stored in the storage unit 300 is input to the machine learning model generation device 3. The machine learning model generation device 3 generates a machine learning model 5 that detects a person with difficulty moving in a difficult-to-move posture from a captured image of a surveillance camera that monitors a surveillance target location where passersby pass by learning using this teacher data 330.
[0038] The storage unit 300 is realized by a storage device such as a hard disk, ROM (Read Only Memory), or RAM (Random Access Memory), stores programs, data, etc. for the processing unit 200 to integrally control the generation device 1, is used as a work area for the processing unit 200, and temporarily stores calculation results and the like executed by the processing unit 200 according to various programs. In this embodiment, a teacher data generation program 302, moving image data 310, tracking detection data 320, and teacher data 330 are stored.
[0039] [Operational Effects] According to the present embodiment, teacher data for training a machine learning model for detecting a person in a specific state can be easily generated. That is, the image portion of the person shown in the passing video is tracked and detected, and based on the shape change of the image portion, the image portion showing a person estimated to have reached a specific state such as difficulty in moving can be generated and stored as teacher data. When a passerby reaches a state of difficulty in moving due to a fall or the like, the posture often changes. Therefore, it is possible to estimate that the person has reached a state of difficulty in moving from the change in the shape of the image portion showing the person. Further, since the image portion showing the person is used as teacher data, the passing video may be an image taken at a location different from the monitoring target location. For this reason, it is easy to acquire and input a large number of passing videos, and as a result, it is possible to generate a large number of teacher data.
[0040] [Modification example] Note that the applicable embodiments of the present invention are not limited to the above-described embodiments, and it goes without saying that they can be appropriately changed without departing from the spirit of the present invention.
[0041] (A) Bicycle rider For example, for a bicycle rider, the person and the bicycle shown in the passing video may be identified and tracked and detected. In this case, it is estimated whether or not the person (bicycle rider) has reached a state of difficulty in moving based on the shape changes of the image portion showing the person (bicycle rider) and the image portion showing the bicycle.
[0042] FIG. 8 is a diagram showing an example of a frame image. In FIG. 8, for the bicycle rider M4, the person (bicycle rider M4) and the bicycle on which the person appears are distinguished from each other, and each is detected in a tracking manner. FIG. 9 is a frame image several frames after the image shown in FIG. 8. In FIG. 9, the bicycle rider M4 has fallen off the bicycle and is in a state of being difficult to move. In this case, for each of the image portion G4 in which the person (bicycle rider M4) appears and the image portion G5 in which the bicycle appears, it is estimated that the object ("person" or "bicycle") appearing has reached a state of being difficult to move due to a change in its shape (for example, a change in the aspect ratio by a predetermined value or more). Then, similarly, the image portion in which the estimated object appears is extracted and associated with the object ID and type ("person" or "bicycle") to be used as teacher data.
[0043] Since a bicycle moves almost integrally with its rider (person), if a bicycle rider reaches a state of being difficult to move due to falling or the like, the bicycle may also fall, and thus the shape of the image portion thereof may change. Therefore, using the image portion in which the fallen bicycle appears as teacher data is useful for detecting the fallen bicycle rider as a person with difficulty in moving in a state of being difficult to move.
[0044] (B) Non-detection during tracking Also, even when a person appearing in a moving image is no longer detected (becomes non-detected) during the process of being detected in a tracking manner, it may be estimated that the person has reached a state of being difficult to move.
[0045] FIG. 10 is a diagram showing an example of a frame image several frames after the image shown in FIG. 2. In FIG. 10, the pedestrian M1 has crouched down, making it difficult to move. However, this shows a case where tracking has been interrupted because the same object cannot be recognized due to a large change in the external shape, etc. In such a case, as shown in FIG. 11, the image portion immediately before the detection is interrupted is used as the teacher data. That is, in the example of FIG. 11, the shape of the image portion in which the pedestrian M1 being detected trackwise appears was a vertically long rectangular shape in the frame images up to the (N - 1)-th frame, but in the next N-th frame, it changed due to crouching down, and in the subsequent (N + 1)-th frame and later, detection was not performed (non-detection) because it was not recognized as the same object. In such a case, at the (N + 1)-th frame when detection stops, it is estimated that the person (pedestrian M1) has reached a state where it is difficult to move, and immediately before that, that is, the frame image of the N-th frame which was the last detected is extracted. Then, similarly, the image portion in which the person in the extracted frame image appears is cut out and associated with the object ID and type (name) of the person to be used as the teacher data.
[0046] (C) Generation of Machine Learning Model The generation device 1 may also have the function of the machine learning model generation device 3.
Explanation of Reference Numerals
[0047] 1... Generation device of teacher data 200... Processing unit 202... Incoming moving image input unit 204... Tracking detection unit 206... Estimation unit 208... Image storage control unit for teacher data 300... Storage unit 302... Teacher data generation program 310... Incoming moving image data 320... Tracking detection data 330... Teacher data 3... Machine learning model generation device 5... Machine learning model
Claims
1. A device for generating teacher data for training a machine learning model that detects a pedestrian who appears to have difficulty moving from an image captured by a surveillance camera that monitors a surveillance target area where pedestrians pass by, comprising: a traffic video input means for inputting a traffic video image obtained by photographing a place where people are passing by; a tracking detection means for tracking and detecting an image portion of a person captured in the traffic motion image; an estimation means for estimating that a person appearing in the image portion has difficulty moving when the shape of the image portion detected by the tracking detection means has changed and the position and shape of the image portion after the change continue for a predetermined period of time; A training data image storage control means for controlling storage of the image portion estimated by the estimation means to have reached the difficult-to-move state as training data in a predetermined storage unit; A generating device comprising:
2. A device for generating teacher data for training a machine learning model that detects a pedestrian who appears to have difficulty moving from an image captured by a surveillance camera that monitors a surveillance target area where pedestrians pass by, comprising: The passers-by include pedestrians and cyclists; a traffic video input means for inputting a traffic video image obtained by photographing a place where people are passing by; a tracking detection means for identifying a bicyclist and a bicycle captured in the traffic video image and for detecting the respective image portions by tracking; an estimation means for estimating whether the cyclist has become unable to move based on a change in shape of both the image portion of the cyclist and the image portion of the bicycle related to the cyclist among the image portions detected by the tracking detection means; a teacher data image storage control means for controlling storage of the image portion determined to be positive by the estimation means as the teacher data in a predetermined storage unit; A generating device comprising:
3. The estimation means determines that the change in shape is a change in the aspect ratio of the image portion that is equal to or greater than a predetermined value. A generating device according to claim 1 or 2.
Citation Information
Patent Citations
Object detection device and object detection method
JP2019168950A
Person detection device, method, and program
JP2021033343A
Sensor system, image sensor, and sensing method
JP2021039592A
System for implementing full automation of annotation in edge ai system
JP2023129172A
A system of detecting abnormal action
KR102628689B1