Learning device and learning method

JP7909200B2Active Publication Date: 2026-08-21PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022090819
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-06-03
Publication Date
2026-08-21
Estimated Expiration
2042-06-03

AI Technical Summary

Benefits of technology

【0011】 本発明によれば、処理条件設定画面から判定基準画像の撮影画像上の位置を設定し、学習対象画像が、監視エリアの撮影画像に判定基準画像が重畳合成されたもので生成される。このため、ユーザの主観に大きく左右されることなく、注目事象の発生状況をユーザが容易にかつ精度よく判断して、適切にラベル付けされた学習対象画像を効率よく収集して、精度の高い学習モデルを生成することができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007909200000001
    Figure 0007909200000001
  • Figure 0007909200000002
    Figure 0007909200000002
  • Figure 0007909200000003
    Figure 0007909200000003
Patent Text Reader

Abstract

To provide a learning device capable of generating a highly accurate learning model while allowing a user to easily and accurately determine the occurrence status of notable events without being greatly influenced by the user's subjectivity, by efficiently collecting properly labeled training target images, and a learning method.SOLUTION: In a monitoring system, image management server 3 generates a learning target image in which a determination reference image that is used as a reference when determining an entry of an object into a predetermined area is superimposed and synthesized with a picked-up image. A monitoring terminal 4 presents a learning target image to a user, and based on user operation, acquires information on the occurrence status of a notable event that appeared in the learning target image as label information. An image analysis server 2 generates a learning model by machine learning using the learning target image and the label information as learning information.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a learning device and a learning method for generating a learning model for detecting, as an attention event, a state in which an object has entered a predetermined area set on a captured image based on the captured image of a monitoring area.

Background Art

[0002] Systems for detecting abnormal events occurring in a monitoring area based on a captured image of the monitoring area have been widely spread. Further, in recent years, systems for detecting abnormal events using a learning model generated by machine learning such as deep learning have also been used.

[0003] In a system for detecting an abnormal event using a learning model, before operation, machine learning is performed using a learning target image (teacher image) based on the captured image of the monitoring area, thereby creating a learning model. Also, during operation, by inputting a detection target image based on the captured image of the monitoring area into the learning model, a detection result of an abnormal event can be obtained. Further, during operation, if additional learning is performed using the detection target image used for detecting the abnormal event as a learning target image, the accuracy of the learning model can be increased at any time.

[0004] As a technique for performing additional learning using such a detection target image as a learning target image, conventionally, a detection target image is presented to a user as a candidate for a learning target image in additional learning, and the user visually determines the suitability of the detection target image as a learning target image for additional learning, and the user selects the learning target image for additional learning (see Patent Document 1).

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

[0006] In conventional technology, the detection target images, which serve as candidates for training images in additional learning, are presented to the user with a detection result (correct / incorrect judgment result) indicating whether or not a predetermined event (appearance of an object) was detected. The user visually inspects the detection target images and, based on the correctness of the detection result, determines whether or not they are suitable as training images, thereby selecting training images for additional learning.

[0007] However, as with conventional technologies, simply visually inspecting the detection target images, which are captured images of the monitoring area, presented a problem: the judgment of whether or not an image is suitable for training was heavily influenced by the user's subjectivity, leading to inconsistent results. Furthermore, when detecting an abnormal event as the entry of an object (e.g., a person) into a designated area, it is difficult to determine the circumstances of the abnormal event. Simply visually inspecting the detection target images, which are captured images of the monitoring area, does not allow for easy and accurate determination of the abnormal event's status. As a result, labeling the training target images with label information regarding the abnormal event's status becomes inaccurate, making it impossible to create a highly accurate training model.

[0008] Therefore, the main objective of the present invention is to provide a learning device and learning method that enable users to easily and accurately determine the occurrence of a subject of interest without being heavily influenced by the user's subjectivity, efficiently collect appropriately labeled learning target images, and generate a highly accurate learning model. [Means for solving the problem]

[0009] The learning device of the present invention is a learning device that uses a processor to generate a learning model for detecting, based on captured images of a monitoring area, the state in which an object enters a predetermined area set on the captured images as a target event, wherein the processor displays a processing condition setting screen on which the captured images are arranged, and sets the position of a judgment criterion image that serves as a reference when determining whether an object has entered the predetermined area based on user operation,By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated, and the training target image is superimposed with the judgment criterion image. The system is configured to present the image to the user, acquire information regarding the occurrence of the notable event that appeared in the training image as label information based on the user's operation, and generate the learning model by machine learning using the training image and the label information as training information.

[0010] Furthermore, the learning method of the present invention is a learning method in which a processor performs a process to generate a learning model for detecting as a focus event the state in which an object enters a predetermined area set on the captured image of a monitoring area, and displays a processing condition setting screen on which the captured image is arranged, and sets the position of a judgment criterion image that serves as a reference when determining whether an object has entered the predetermined area based on user operation. By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated, and the training target image is superimposed with the judgment criterion image. The system is configured to present the image to the user, acquire information regarding the occurrence of the notable event that appeared in the training image as label information based on the user's operation, and generate the learning model by machine learning using the training image and the label information as training information. [Effects of the Invention]

[0011] According to the present invention, From the processing conditions setting screen, set the position of the judgment criterion image on the captured image. The training image is a composite image created by superimposing a judgment criterion image onto a photograph taken in the monitoring area. It is generated by [this method]. Therefore, without being heavily influenced by the user's subjective opinion, users can easily and accurately judge the occurrence of events of interest, efficiently collect appropriately labeled training images, and generate highly accurate learning models. [Brief explanation of the drawing]

[0012] [Figure 1] Overall configuration diagram of the monitoring system according to this embodiment [Figure 2] An explanatory diagram showing the judgment lines and judgment areas set on the camera's captured image. [Figure 3] An explanatory diagram showing the images processed by this system before operation. [Figure 4] An explanatory diagram showing an overview of the processes performed by this system before it is put into operation. [Figure 5] Explanatory drawing showing an image processed by this system during operation [Figure 6] Block diagram showing an overview of the processing performed by this system during operation [Figure 7] Block diagram showing a schematic configuration of an image analysis server, an image management server, and a monitoring terminal [Figure 8] Explanatory drawing showing a monitoring screen displayed on the monitoring terminal [Figure 9] Explanatory drawing showing a processing target setting screen displayed on the monitoring terminal [Figure 10] Explanatory drawing showing a camera confirmation screen displayed on the monitoring terminal [Figure 11] Explanatory drawing showing a processing condition setting screen displayed on the monitoring terminal [Figure 12] Explanatory drawing showing a processing condition setting screen displayed on the monitoring terminal [Figure 13] Explanatory drawing showing a notification screen displayed on the monitoring terminal [Figure 14] Explanatory drawing showing a notification screen displayed on the monitoring terminal [Figure 15] Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 16] [[ID=3~4]]Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 17] Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 18] Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 19] Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 20] Explanatory drawing showing a status confirmation screen displayed on the monitoring terminal [Figure 21] Explanatory drawing showing a learning model creation screen displayed on the monitoring terminal [Figure 22] Explanatory drawing showing a rollback setting screen displayed on the monitoring terminal

Mode for Carrying Out the Invention

[0013] The first invention made to solve the above problem is a learning device that uses a processor to generate a learning model for detecting, based on captured images of a monitoring area, the state in which an object enters a predetermined area set on the captured images as a noteworthy event, wherein the processor displays a processing condition setting screen on which the captured images are arranged, and sets the position of a judgment criterion image that serves as a reference when determining whether an object has entered the predetermined area based on user operation, By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated, and the training target image is superimposed with the judgment criterion image. The system is configured to present the image to the user, acquire information regarding the occurrence of the notable event that appeared in the training image as label information based on the user's operation, and generate the learning model by machine learning using the training image and the label information as training information.

[0014] According to this, From the processing conditions setting screen, set the position of the judgment criterion image on the captured image. The training image is a composite image created by superimposing a judgment criterion image onto a photograph taken in the monitoring area. It is generated by [this method]. Therefore, without being heavily influenced by the user's subjective opinion, users can easily and accurately judge the occurrence of events of interest, efficiently collect appropriately labeled training images, and generate highly accurate learning models.

[0015] Furthermore, the second invention is configured such that the processor generates a detection target image by superimposing the judgment criterion image onto the real-time captured image, performs processing to detect the attention event from the detection target image using the learning model, and when the attention event is detected, sets the detection target image as the learning target image for additional learning, presents the learning target image to the user, acquires the label information based on the user's operation, and generates the updated learning model by machine learning using the learning target image and the label information as learning information for additional learning.

[0016] According to this approach, the images used in the process of detecting the event of interest become the same images used for further training, allowing for more efficient additional training.

[0017] Furthermore, the third invention is configured such that the processor sets the detection target image specified by the user as the target of additional learning as the learning target image in the additional learning, based on the user's operation.

[0018] According to this, the images to be detected specified by the user are set as targets for additional training, allowing for more appropriate additional training.

[0019] Furthermore, the fourth invention is configured such that the processor generates a learning model for detecting, as the event of interest, the state in which an object has entered a predetermined area corresponding to a dangerous area in which an object is moving.

[0020] According to this, the system can detect a target event when an object approaches a hazardous area where objects are moving. A hazardous area where objects are moving includes, for example, a conveyor belt for transporting goods or a railway track where trains run.

[0021] Furthermore, the fifth invention is configured such that the processor superimposes and combines an image representing a linear judgment line as the judgment criterion image onto the captured image.

[0022] According to this method, by setting a linear judgment line as the criterion for detecting the event of interest, the event of interest can be appropriately detected. Therefore, it becomes possible to distinguish and detect events of interest that approach the object in a way that intersects the judgment line from other events of interest. Note that there may be one judgment line or multiple judgment lines.

[0023] Furthermore, the sixth invention is configured such that the processor superimposes and combines an image representing a polygonal determination area as the determination criterion image onto the captured image.

[0024] According to this, by setting a polygonal detection area as the criterion for detecting events of interest, events of interest can be detected appropriately.

[0025] Furthermore, the seventh invention is a learning method in which a processor performs a process to generate a learning model for detecting, as a notable event, when an object enters a predetermined area set on a captured image of a monitoring area, the method comprising: displaying a processing condition setting screen on which the captured image is placed, and setting the position of a judgment criterion image that serves as a reference when determining whether an object has entered the predetermined area based on user operation; By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated, and the training target image is superimposed with the judgment criterion image. The system is configured to present the image to the user, acquire information regarding the occurrence of the notable event that appeared in the training image as label information based on the user's operation, and generate the learning model by machine learning using the training image and the label information as training information.

[0026] According to this, similar to the first invention, users can easily and accurately judge the occurrence of events of interest without being heavily influenced by the user's subjectivity, efficiently collect appropriately labeled training images, and generate a highly accurate learning model.

[0027] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0028] Figure 1 is an overall diagram of the monitoring system according to this embodiment.

[0029] This system monitors abnormal events (events of interest) occurring in a monitoring area based on images captured within that area. The system comprises a camera 1, an image analysis server 2 (learning device), an image management server 3 (learning device), and a monitoring terminal 4 (terminal device). Camera 1, image analysis server 2, image management server 3, and monitoring terminal 4 are connected via a network such as the internet or a private network.

[0030] Camera 1 photographs the surveillance area.

[0031] Image analysis server 2 performs image analysis processing using a learning model generated by machine learning such as deep learning, detects abnormal events occurring in the monitoring area, and notifies the user (monitoring officer) of the abnormal event using monitoring terminal 4. In addition, image analysis server 2 generates a learning model using training images (teacher images) accumulated by image management server 3.

[0032] Image management server 3 stores and manages training images (teacher images) necessary for generating the learning model used by image analysis server 2. Training images are generated from images captured by camera 1.

[0033] Monitoring terminal 4 displays a monitoring screen. The monitoring screen displays real-time images captured by camera 1. This allows the user (monitoring officer) to check the current status of the monitoring area. Monitoring terminal 4 also displays a notification screen. This allows the user to quickly recognize if an abnormal event has occurred in the monitoring area. Furthermore, monitoring terminal 4 allows the user (administrator) to perform operations related to setting conditions for processing performed by image analysis server 2 and image management server 3.

[0034] While the monitoring terminal 4, image management server 3, and image analysis server 2 can be configured to operate on-premises, i.e., within the facility premises, the image management server 3 and image analysis server 2 may also be operated in the cloud.

[0035] Furthermore, in this embodiment, various processes are performed in the image management server 3 and the image analysis server 2, but these processes may also be performed on a single server. Alternatively, the processes performed in the image management server 3 and the image analysis server 2 may be shared among multiple servers in a different combination than that of this embodiment.

[0036] Next, we will explain the criteria for detecting abnormal events in the monitoring area. Figure 2 is an explanatory diagram showing the judgment line and judgment area set on the image captured by camera 1.

[0037] In this embodiment, the monitoring area is a workplace where a conveyor belt is installed. In a workplace, accidents such as hands getting caught or clothing getting caught may occur when a person approaches the conveyor belt. Therefore, in this embodiment, a condition in which there is a high risk of an accident occurring due to a person approaching the conveyor belt is detected as an abnormal event.

[0038] In this embodiment, as shown in Figure 2(A), a straight line is set on the image captured by camera 1 as a judgment criterion. In this example, two judgment lines, a first and a second, are set. The first judgment line is set at a position close to the conveyor (hazardous material), and the second judgment line is set at a position away from the conveyor. The judgment line is a boundary line that separates a predetermined area (hazardous area) where human bodies are prohibited from entering from the rest of the area (safe area).

[0039] Here, there are three states: the state in which the person's (object's) body intersects (contacts) both the first and second determination lines (first intersection state); the state in which the person's body intersects (contacts) only the second determination line and not the first determination line (second intersection state); and the state in which the person's body does not intersect (contact) either the first or second determination line (non-intersection state).

[0040] In this embodiment, the first and second crossing states are detected as abnormal events and a notification is issued. The first judgment line corresponds to the "danger" notification level, and the second judgment line corresponds to the "caution" notification level. That is, a state in which a person's body crosses (touches) both the first and second judgment lines (first crossing state) is considered high-risk, and a "danger" notification is issued. On the other hand, a state in which a person's body crosses (touches) only the second judgment line and does not cross (touch) the first judgment line (second crossing state) is considered low-risk, and a "caution" notification is issued.

[0041] The judgment line may be set to just one line, or it may be set to three or more lines.

[0042] Furthermore, the first and second judgment lines are drawn in different colors. For example, the first judgment line is drawn in red, and the second judgment line is drawn in yellow. This allows the user to easily visually determine the state of the intersection (contact) of a person's body with respect to the first and second judgment lines.

[0043] Furthermore, as shown in Figure 2(B), the judgment area may be set on the image captured by camera 1 as the judgment criterion. The boundary line of the judgment area (judgment line) is a polygon. The judgment area is defined such that, for example, the area inside the polygonal boundary line is a designated area (danger area) where entry by a person's body is prohibited, and the area outside is an area (safe area) where the presence of a person is permitted.

[0044] Furthermore, multiple judgment areas may be set. For example, another judgment area may be set outside the judgment area shown in Figure 2(B) so as to surround it. In this case, the inner judgment area may correspond to the "danger" notification level, and the outer judgment area may correspond to the "caution" notification level.

[0045] In this embodiment, the monitoring area is a workplace where a conveyor (transport device) is installed, but the monitoring area is not limited to this. For example, the platform and tracks at a railway station may be the monitoring area. In this case, conditions that pose a high risk of accidents such as users falling from the platform onto the tracks or coming into contact with a train are detected as abnormal events and reported.

[0046] Furthermore, in this embodiment, abnormal events (events of interest) occurring in the monitoring area are detected based on images captured in the monitoring area, but the detected events of interest are not limited to abnormal events.

[0047] Next, we will explain the overview of the processing performed by this system before operation. Figure 3 is an explanatory diagram showing the images processed by this system before operation. Figure 4 is an explanatory diagram showing the overview of the processing performed by this system before operation.

[0048] In this embodiment, two judgment lines, a first and a second, are set on the image captured by camera 1, and a state in which a person's body intersects (touches) the first and second judgment lines (first intersection state, second intersection state) is detected as an abnormal event. The process of detecting abnormal events is performed by image analysis processing using a learning model generated by machine learning such as deep learning.

[0049] In this system, a process (main learning) is performed to generate a learning model before operation. In this embodiment, the learning model is generated by learning using training target images (teacher images) generated from images captured by camera 1.

[0050] At this time, as shown in Figure 3, a training image is generated by superimposing a judgment criterion image (an image representing the judgment line), which serves as the basis for determining the presence or absence of abnormal events, onto the image captured by camera 1 that photographs the monitoring area. The generated training image is displayed on the annotation screen of the monitoring terminal 4.

[0051] On the annotation screen, the user (administrator) visually inspects the displayed training image, determines the state of the monitoring area (first crossing state, second crossing state, non-crossing state), and inputs label information related to the state of that monitoring area (annotation work). At this time, since the judgment criterion image is superimposed on the training image, the user can easily and accurately determine the occurrence of abnormal events by visually inspecting the training image, thereby streamlining the user's annotation work.

[0052] As shown in Figure 4, in this embodiment, the image management server 3 performs a process to generate a training image by superimposing a judgment criterion image onto the image captured by the camera 1 (image synthesis process). In addition, the image management server 3 performs labeling by adding label information related to the state of the monitoring area (first crossing state, second crossing state, non-crossing state) to the training image (annotation process). Furthermore, the image management server 3 registers the training image and label information in the training image database.

[0053] Next, the image analysis server 2 retrieves the training target images and label information stored in the image management server 3 from the image management server 3, and generates a training model using machine learning with the training target images and label information (training model generation process).

[0054] Next, we will explain the overview of the processes performed by this system during operation. Figure 5 is an explanatory diagram showing the images processed by this system during operation. Figure 6 is an explanatory diagram showing the overview of the processes performed by this system during operation.

[0055] During operation, as shown in Figure 5, a detection target image is generated by superimposing a judgment criterion image (an image representing the judgment line), which serves as the basis for determining the presence or absence of abnormal events, onto the real-time image captured by camera 1.

[0056] As shown in Figure 6, in this embodiment, the image analysis server 2 performs a process to generate a detection target image by superimposing a judgment criterion image onto the image captured by the camera 1 (image synthesis process). In addition, the image analysis server 2 performs image analysis processing using a learning model on the detection target image, and based on the analysis results, an anomaly detection process is performed to determine whether or not there is an abnormal event, i.e., a state in which a person's body crosses (comes into) the first and second judgment lines (first crossing state, second crossing state).

[0057] When an abnormal event is detected, the monitoring terminal 4 notifies the user of the occurrence of the abnormal event. At this time, a notification screen containing notification information corresponding to the severity of the abnormal event is displayed, followed by a status confirmation screen. The status confirmation screen displays the detected image that became the target of the notification due to the detected abnormal event (see Figure 5).

[0058] Here, the detected image, which has been identified as an abnormal event and is subject to notification, becomes the training image for additional training. The user (administrator) visually inspects the displayed training image on the status confirmation screen, determines the state of the monitoring area (first crossing state, second crossing state, non-crossing state), and inputs label information related to the state of that monitoring area (annotation work). At this time, since the judgment criterion image is superimposed on the training image, the user can easily and accurately determine the state of the monitoring area by visually inspecting the training image, thus streamlining the user's annotation work even during additional training.

[0059] Image analysis server 2 performs labeling (annotation processing), adding label information related to the state of the monitoring area (first crossing state, second crossing state, non-crossing state) to the training images. The training images and label information are registered in the training image database on image management server 3 as training information for further training.

[0060] Next, the image analysis server 2 retrieves the training target images and label information stored in the image management server 3, and uses that training target images and label information to generate a new training model (training model generation process). Next, the image analysis server 2 applies the new training model to the image analysis process (training model management process).

[0061] Furthermore, the learning model may be updated each time an abnormal event is detected and notification is issued. In this embodiment, when notification is issued, the detected image that became the target of notification due to the detection of an abnormal event, i.e., the learning target image in the additional learning, is labeled, so the learning model may be updated using the newly added learning target image at this timing.

[0062] Next, we will describe the schematic configuration of the image analysis server 2, the image management server 3, and the monitoring terminal 4. Figure 7 is a block diagram showing the schematic configuration of the image analysis server 2, the image management server 3, and the monitoring terminal 4.

[0063] The image management server 3 comprises a communication unit 31, a storage unit 32, and a processor 33.

[0064] The communication unit 31 communicates with the camera 1, the image analysis server 2, and the monitoring terminal 4.

[0065] The memory unit 32 stores programs executed by the processor 33, etc. The memory unit 22 stores registration information for the learning target image database. The learning target image database registers learning target images (training images) for each state of the monitoring area (first crossing state, second crossing state, non-crossing state).

[0066] The processor 33 performs various processes by executing programs stored in the memory unit 32. In this embodiment, the processor 33 performs image synthesis processing and annotation processing, among other things.

[0067] In the image synthesis process, the processor 33 superimposes judgment criterion images (judgment line images, judgment area images) onto the images captured by camera 1 to generate training target images (teacher images) to be used for pre-operation training. At this time, the image synthesis process may be performed using real-time images captured by camera 1, or it may be performed using images captured by camera 1 that have been stored in the device.

[0068] In the annotation process, the processor 33, in response to user input operations performed on the monitoring terminal 4, labels the training images generated by the image synthesis process by adding label information regarding the occurrence of abnormal events that appeared in those training images.

[0069] The image analysis server 2 comprises a communication unit 21, a storage unit 22, and a processor 23.

[0070] The communication unit 21 communicates with the camera 1, the image management server 3, and the monitoring terminal 4.

[0071] The memory unit 22 stores programs executed by the processor 33, etc. The memory unit 22 also stores information about previously updated learning models.

[0072] The processor 23 performs various processes by executing programs stored in the memory unit 22. In this embodiment, the processor 23 performs learning model generation processing, learning model management processing, image synthesis processing, image analysis processing, anomaly detection processing, and annotation processing.

[0073] In the learning model generation process, the processor 23 obtains training images (teacher images) from the image management server 3 and uses these training images to generate a learning model using machine learning such as deep learning.

[0074] In the learning model management process, the processor 23 manages the learning model used in the image analysis process. In this embodiment, the processor 23 performs a rollback process to revert the learning model used in the image analysis process back to a previous learning model specified by the user. The rollback process is executed when the user determines that a rollback is necessary because the currently running learning model has many problems, specifically because of frequent false positives, and instructs the system to perform a rollback.

[0075] In the image synthesis process, the processor 23 acquires real-time images captured by the camera 1, and generates a detection target image by superimposing a judgment criterion image (judgment line image, judgment area image) onto the captured images.

[0076] In the image analysis process, the processor 23 uses the learning model generated in the learning model generation process to perform image analysis on the detection target image acquired in the image synthesis process, and obtains the probability that an abnormal event occurred, namely, a state in which a person's body crosses (contacts) the first and second judgment lines (first crossing state, second crossing state). Specifically, the detection target image is input to the learning model, and the probability of the abnormal event output from the learning model as an analysis result is obtained.

[0077] In the anomaly detection process, the processor 23 detects an abnormal event, that is, a state in which a person's body crosses (comes into) the first and second judgment lines (first crossing state, second crossing state). At this time, the presence or absence of an abnormal event (first crossing state, second crossing state) is determined from the accuracy of the abnormal event obtained by the image analysis process, based on the sensitivity specified by the user.

[0078] In the annotation process, the processor 23, in response to user input operations performed on the monitoring terminal 4, labels candidate images for further learning, i.e., detected images in which abnormal events have been detected and which are subject to notification, by adding label information regarding the occurrence of the abnormal events that appeared in those images.

[0079] The monitoring terminal 4 comprises a display 41, an input device 42, a communication unit 43, a storage unit 44, and a processor 45.

[0080] The display 41 shows the screen. The input device 42 is a keyboard or mouse, and detects user input.

[0081] The communication unit 43 communicates with the image management server 3 and the image analysis server 2.

[0082] The memory unit 44 stores programs and other data executed by the processor 45.

[0083] The processor 45 performs various processes by executing programs stored in the memory unit 44. In this embodiment, the processor 45 performs tasks such as display input control.

[0084] In the display input control process, the processor 45 displays various screens, such as the monitoring screen (see Figure 8), on the display 41, and acquires input operation information in response to the user's operation of the input device 42. The user's input operation information is sent to the image management server 3 and the image analysis server 2.

[0085] Next, we will explain the monitoring screen 101 displayed on the monitoring terminal 4. Figure 8 is an explanatory diagram showing the monitoring screen 101.

[0086] The monitoring screen 101 is provided with a monitoring area selection unit 102. The monitoring area selection unit 102 allows the user to select the target monitoring area. In this example, the user can select the workspace as the monitoring area.

[0087] Furthermore, the monitoring screen 101 is equipped with a camera image display unit 103. The camera image display unit 103 displays images (camera images) captured by multiple cameras 1 that photograph the monitoring area (workplace) selected by the monitoring area selection unit 102, arranged side by side. The camera image display unit 103 displays real-time captured images (live images). This allows the user (monitoring officer) to visually check the images captured by all cameras 1 that photograph the target monitoring area (workplace) and confirm the current status of the monitoring area.

[0088] When a user performs a predetermined operation (right-click) on one of the display fields of multiple monitoring areas (workspaces) displayed in the monitoring area selection unit 102, a menu 104 is displayed. If the user selects a processing target setting in the menu 104, the system transitions to the processing target setting screen 111 (see Figure 9).

[0089] Next, we will explain the processing target setting screen 111 displayed on the monitoring terminal 4. Figure 9 is an explanatory diagram showing the processing target setting screen 111.

[0090] The processing target setting screen 111 is provided with a camera image display unit 112. The camera image display unit 112 displays images (camera images) taken by multiple cameras 1 that photograph the monitoring area (workplace) selected on the monitoring screen 101, side by side. The camera image display unit 112 allows setting whether or not each camera 1 should be included in the anomaly detection processing performed by the image analysis server 2.

[0091] Specifically, when a user performs a predetermined operation (right-click) on the display field for the name of each camera 1 in the camera image display unit 112, a menu 113 is displayed, and the user can select either to be processed or not to be processed from the menu 113. If "to be processed" is selected, the corresponding camera 1 is set as the target of anomaly detection processing, and if "not to be processed" is selected, the corresponding camera 1 is excluded from anomaly detection processing. In this example, camera #1 is set as the target of anomaly detection processing, and cameras #2 to #8 are set as not to be processed.

[0092] Furthermore, when a user selects a target camera 1 by operating one of the display fields for each camera 1 on the camera image display unit 112, the system transitions to a camera confirmation screen 121 (see Figure 10) related to the selected camera 1.

[0093] Furthermore, the processing target setting screen 111 is provided with a "back" button 114. When the user operates the "back" button 114, they return to the monitoring screen 101 (see Figure 8).

[0094] Next, we will explain the camera confirmation screen 121 displayed on the monitoring terminal 4. Figure 10 is an explanatory diagram showing the camera confirmation screen 121.

[0095] The camera confirmation screen 121 is equipped with a camera attribute information display unit 122. The camera attribute information display unit 122 displays the attribute information of camera 1, specifically the installation date and time, name, and IP address of camera 1. This allows the user to check the attribute information of camera 1.

[0096] Furthermore, the camera confirmation screen 121 is equipped with a camera image display unit 123. The camera image display unit 123 displays the image captured by camera 1. This allows the user to visually check the image captured by camera 1 and confirm the shooting status of the target work area.

[0097] Furthermore, the camera confirmation screen 121 is provided with a "Get Camera Information" button 124. When the user operates the "Get Camera Information" button 124, a process is performed to obtain the attribute information of camera 1 from the image analysis server 2, and the attribute information of camera 1 is displayed on the camera attribute information display unit 122. At the same time, a process is performed to obtain the captured image from camera 1, and the captured image from camera 1 is displayed on the camera image display unit 123.

[0098] Furthermore, the camera confirmation screen 121 has a "Camera Confirmation" tab 125 and a "Processing Condition Settings" tab 126. When the "Camera Confirmation" tab 125 is selected on the camera confirmation screen 121, the user will be redirected to the processing condition settings screen 131 (see Figures 11 and 12).

[0099] Next, we will explain the processing condition setting screen 131 displayed on the monitoring terminal 4. Figures 11 and 12 are explanatory diagrams showing the processing condition setting screen 131.

[0100] As shown in Figures 11 and 12, the processing condition setting screen 131 is provided with a camera image display unit 132. The camera image display unit 132 displays the image captured by camera 1.

[0101] Furthermore, the processing condition setting screen 131 is provided with a judgment criterion setting section 133. The judgment criterion setting section 133 is provided with a judgment criterion selection field 141, a coordinate display field 142, an "Add" button 143, a "Delete" button 144, and a "Set" button 145.

[0102] Here, the example shown in Figure 11 is one in which a judgment line is set as the judgment criterion.

[0103] In this case, when the user operates the judgment criterion selection field 141, a pull-down menu is displayed, and the user selects a first judgment line (red judgment line) and a second judgment line (yellow judgment line) from the pull-down menu. At this time, the system transitions to judgment line input mode, and the user can use the input device 42 to specify the positions of the two endpoints of the straight line representing the judgment line on the captured image displayed on the camera image display unit 132. The first judgment line corresponds to the "danger" notification level, and the second judgment line corresponds to the "caution" notification level.

[0104] The coordinate display area 142 shows the coordinates of the two endpoints of the line representing the judgment line. Alternatively, you can specify the coordinates of three or more endpoints and use a polyline connecting the specified points as the judgment line.

[0105] When a user presses the "Add" button 143, they can add a new judgment line. When a user presses the "Set" button 144, the judgment line specified by the user is set. When a user presses the "Delete" button 145, they can delete a previously set judgment line.

[0106] On the other hand, the example shown in Figure 12 is one in which a judgment area is set as the judgment criterion.

[0107] In this case, when the user operates the judgment criterion selection field 141, a pull-down menu is displayed, and the user selects a judgment area from the pull-down menu. At this time, the system transitions to the judgment area input mode, and the user can use the input device 42 to specify the positions of multiple endpoints of the polygon representing the judgment area on the captured image displayed on the camera image display unit 132.

[0108] The coordinate display area 142 shows the coordinates of multiple endpoints of the polygon representing the determination area.

[0109] When a user clicks the "Add" button 143, they can add a new judgment area. When a user clicks the "Set" button 144, the judgment area specified by the user is set. When a user clicks the "Delete" button 145, they can delete a previously set judgment area.

[0110] Furthermore, the processing condition setting screen 131 is provided with a mask area setting section 134. In the mask area setting section 134, the user can specify mask areas to be excluded from the anomaly detection processing. The mask area setting section 134 is provided with a mask area selection field 146, a coordinate display field 147, an "Add" button 148, a "Delete" button 148, and a "Set" button 150.

[0111] When the user operates the mask area selection field 146, a pull-down menu is displayed, allowing the user to select a mask area. At this time, the system transitions to mask area input mode, and the user can use the input device 42 to specify the positions of multiple endpoints of the polygon representing the mask area on the captured image displayed on the camera image display unit 132.

[0112] The coordinate display area 147 shows the coordinates of multiple endpoints of the polygon representing the mask area.

[0113] When a user clicks the "Add" button 148, they can add a new mask area. When a user clicks the "Set" button 149, the mask area specified by the user is set. When a user clicks the "Delete" button 150, they can delete a set mask area.

[0114] Furthermore, the processing condition setting screen 131 is provided with a sensitivity setting unit 135. In the sensitivity setting unit 135, the user can input a numerical value (0 to 100) for the sensitivity, which serves as the reference value when determining the presence or absence of an abnormal event in the process of detecting an abnormal event, i.e., when a person's body crosses (comes into) the first and second judgment lines (anomaly detection process). Note that a higher numerical value for sensitivity increases the sensitivity of detecting abnormal events.

[0115] Furthermore, the processing condition setting screen 131 is equipped with a "Register" button 136. When the user operates the "Register" button 136, the information entered on this screen is registered as setting information.

[0116] Next, we will explain the notification screen 161 displayed on the monitoring terminal 4. Figures 13 and 14 are explanatory diagrams showing the notification screen 161.

[0117] When an abnormal event is detected, namely a state in which a person's body crosses (comes into) the first and second judgment lines (first crossing state, second crossing state), the notification screen 161 shown in Figures 13 and 14 is displayed as a pop-up on the monitoring screen 101 (see Figure 8).

[0118] On the notification screen 161, either the word "Danger" or "Caution" is displayed to indicate the notification level, depending on the detected abnormal event (first crossing state, second crossing state). The notification screen 161 shown in Figure 13 is for the case where the notification level is Danger, that is, when a state is detected where the person's body crosses both the first and second judgment lines. In this case, the word "Danger" is displayed on the notification screen 161. The notification screen 161 shown in Figure 14 is for the case where the notification level is Caution, that is, when a state is detected where the person's body crosses only the second judgment line. In this case, the word "Caution" is displayed on the notification screen 161.

[0119] The notification screen 161 is equipped with a "Confirm" button 162. When the user operates the "Confirm" button 162, the system transitions to the status confirmation screen 171 (see Figures 15 to 20).

[0120] Next, we will explain the status confirmation screen 171 displayed on the monitoring terminal 4. Figures 15 to 20 are explanatory diagrams showing the status confirmation screen 171.

[0121] The status confirmation screen 171 is equipped with a camera image display unit 172. The camera image display unit 172 displays the detected image in which an abnormal event has been detected and which is subject to notification. Judgment lines are superimposed on the detected image. Therefore, by visually inspecting the detected image, the user can easily and accurately determine the state of the intersection (contact) of a person's body with the first and second judgment lines.

[0122] Furthermore, the status confirmation screen 171 is provided with an annotation operation unit 173. The annotation operation unit 173 allows the user to perform input operations to label candidate learning target images in additional learning, that is, detected target images that have been detected as abnormal events and are subject to notification. The annotation operation unit 173 is provided with a button 181 for inputting that it is a correct detection, and three buttons 182, 183, and 184 for inputting the actual status in the case of a false detection.

[0123] The user visually inspects the image to be detected and, upon confirming that it is a correct detection, operates button 181. The user also visually inspects the image to be detected and, upon confirming that it is a false detection and that the person's body intersects (touches) both the first detection line (red detection line) and the second detection line (yellow detection line) (first intersection state), operates button 182. The user also visually inspects the image to be detected and, upon confirming that it is a false detection and that the person's body intersects (touches) only the second detection line (yellow detection line) and does not intersect (touches) the first detection line (second intersection state), operates button 183. The user also visually inspects the image to be detected and, upon confirming that it is a false detection and that the person's body does not intersect (touches) either the first or second detection line (non-intersection state), operates button 184.

[0124] Here, the examples shown in Figures 15 to 17 represent the status confirmation screen 171, which is accessed after transitioning from the notification screen 161 (see Figure 13) where the notification level is "Dangerous". In this case, the button 182 corresponding to positive detection is displayed in a grayed-out (unselectable) state.

[0125] In the example shown in Figure 15, the detected image displayed on the camera image display unit 172 shows that the person's body is in contact with both the first judgment line (red judgment line) and the second judgment line (yellow judgment line) (first intersection state), and the notification level is "dangerous". Therefore, in this case, the user visually confirms that it is a correct detection by looking at the detected image and operates the button 181 to input that it is a correct detection.

[0126] Furthermore, in the example shown in Figure 16, the detected image displayed on the camera image display unit 172 shows the person's body crossing only the second judgment line (yellow judgment line) (second crossing state), and the actual notification level is "caution." Therefore, in this case, the user visually inspects the detected image to confirm that it is a false detection and that it is in the second crossing state, and then operates the button 183 to input that it is in the second crossing state.

[0127] Furthermore, in the example shown in Figure 17, the detected image displayed on the camera image display unit 172 shows that the person's body does not intersect (contact) either the first or second determination line (non-intersection state), and therefore, no notification would normally be issued. In this case, the user visually inspects the detected image to confirm that it is a false detection and that the state is non-intersection, and then operates the button 184 to input that the state is non-intersection.

[0128] On the other hand, the example shown in Figures 18 to 20 is the status confirmation screen 171 when transitioning from the notification screen 161 (see Figure 14) where the notification level is "caution". Here, the button 183 corresponding to positive detection is displayed in a grayed-out state (unselectable).

[0129] In the example shown in Figure 18, the detected image displayed on the camera image display unit 172 shows that the person's body is only crossing the second judgment line (yellow judgment line) (second crossing state), and the notification level is "caution". Therefore, in this case, the user visually confirms that it is a correct detection by looking at the detected image and operates the button 181 to input that it is a correct detection.

[0130] Furthermore, in the example shown in Figure 19, the detected image displayed on the camera image display unit 172 shows a state where the person's body intersects (touches) both the first judgment line (red judgment line) and the second judgment line (yellow judgment line) (first intersection state), and the actual notification level is "danger." Therefore, in this case, the user visually checks the detected image to confirm that it is a false detection and that it is the first intersection state, and then operates the button 182 to input that it is the first intersection state.

[0131] Furthermore, in the example shown in Figure 20, the detected image displayed on the camera image display unit 172 shows that the person's body does not intersect (contact) either the first or second determination line (non-intersection state), and therefore, no notification would normally be issued. In this case, the user visually inspects the detected image to confirm that it is a false detection and that the state is non-intersection, and then operates the button 184 to input that the state is non-intersection.

[0132] Furthermore, the status confirmation screen 171 is equipped with a learning model information display unit 174. The learning model information display unit 174 displays the identification information of the learning model currently in operation, that is, the learning model currently used for image analysis processing, specifically the creation date and time of the learning model.

[0133] Furthermore, the status confirmation screen 171 is equipped with a button 175 for instructing the user to roll back the trained model. If the user determines that the currently running trained model has many problems, specifically that there are many false positives, and therefore a rollback to a previous trained model is necessary, they operate button 175. This causes the rollback settings screen 211 (see Figure 22) to pop up on the status confirmation screen 171.

[0134] Furthermore, the status confirmation screen 171 has a "Notification" tab 176 and a "Learning" tab 177. In the status confirmation screen 171, the "Notification" tab 176 is selected, and when the user operates the "Learning" tab 177, the system transitions to the learning model creation screen 201 (see Figure 21).

[0135] Next, we will explain the learning model creation screen 201 displayed on the monitoring terminal 4. Figure 21 is an explanatory diagram showing the learning model creation screen 201.

[0136] The learning model creation screen 201 is provided with a learning target image list display unit 202. The learning target image list display unit 202 displays a list of information regarding candidate learning target images to be used for additional learning. Candidate learning target images are detection target images that have been detected as abnormal events and are subject to notification. Specifically, for each candidate learning target image, the learning target image list display unit 202 displays the time of capture, a thumbnail image, the accuracy of the abnormal event detection result (whether it was a correct or false detection), and label information to be attached to the learning target image.

[0137] When a user performs a predetermined operation (right-click) on the display field of any of the images to be trained in the list of images to be trained display unit 202, menu 203 is displayed. In menu 203, the user can specify whether or not to include candidate images to be trained in the additional training. If "Exclude from additional training" is selected, the candidate images to be trained are excluded from the training process.

[0138] Furthermore, the learning model creation screen 201 is provided with a creation memo input field 204. In the creation memo input field 204, the user can enter text that represents the content (creation memo) of the learning model to be created.

[0139] Furthermore, the learning model creation screen 201 is equipped with a "Create Learning Model" button 205. When the user operates the "Create Learning Model" button 205, a process is performed to create a learning model using the image selected as the target for additional learning from among the candidate learning target images displayed in the learning target image list display unit 202.

[0140] Next, we will explain the rollback settings screen 211 displayed on the monitoring terminal 4. Figure 22 is an explanatory diagram showing the rollback settings screen 211.

[0141] On the status confirmation screen 171 (see Figures 15 to 20), when the user operates the button 175 to instruct the rollback of the learning model, the rollback settings screen 211 shown in Figure 22 is displayed as a pop-up on the status confirmation screen 171.

[0142] The rollback settings screen 211 includes a model selection section 212. The model selection section 212 displays a list of the creation date and content for each learning model. The user can select a learning model by manipulating the column for any of the learning models.

[0143] Furthermore, the rollback settings screen 211 is equipped with a "Perform Switch" button 213 and a "Back" button 214. When the user operates the "Perform Switch" button 213, a rollback process is performed to switch the learning model used for image analysis processing to the learning model selected in the model selection unit 212. When the user operates the "Back" button 214, the user returns to the status confirmation screen 171 (see Figures 15 to 20).

[0144] As described above, embodiments have been explained as examples of the technology disclosed in this application. However, the technology in this disclosure is not limited to these embodiments and can be applied to embodiments that have been modified, replaced, added, or omitted. Furthermore, it is possible to create new embodiments by combining the components described in the above embodiments. [Industrial applicability]

[0145] The learning device and learning method according to the present invention have the effect of enabling users to easily and accurately judge the occurrence of a target event without being greatly influenced by the user's subjectivity, efficiently collect appropriately labeled learning target images, and generate a highly accurate learning model. They are useful as a learning device and learning method for generating a learning model for detecting a target event when an object enters a predetermined area set on a captured image, based on captured images of a monitoring area. [Explanation of Symbols]

[0146] 1 Camera 2. Image analysis server (learning device) 3. Image management server (learning device) 4. Monitoring terminal (terminal device) 23 processors 33 processors

Claims

1. A learning device that uses a processor to generate a learning model for detecting, based on images of a monitoring area, the state in which an object enters a predetermined area set on the captured image as a target event, The aforementioned processor, The processing condition setting screen on which the captured image is placed is displayed, and the position of the judgment criterion image, which serves as the basis for determining whether an object has entered the predetermined area based on user operation, is set. By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated. The user is presented with the learning target image in which the judgment criterion image is superimposed and combined. Based on user operations, information regarding the occurrence status of the notable event that appeared in the training image is obtained as label information. A learning device characterized by generating the learning model by machine learning using the aforementioned learning target images and the aforementioned label information as learning information.

2. The aforementioned processor, A detection target image is generated by superimposing the judgment criterion image onto the real-time captured image. Using the aforementioned learning model, a process is performed to detect the target event from the image to be detected. When the aforementioned event of interest is detected, the detected image is set as the training target image in the additional training, and the training target image is presented to the user. Based on user actions, the label information is obtained, The learning device according to claim 1, characterized in that it generates the updated learning model by machine learning using the aforementioned learning target image and the aforementioned label information as learning information for additional learning.

3. The aforementioned processor, The learning device according to claim 2, characterized in that, based on user operations, the detection target image designated by the user as the target of additional learning is set as the learning target image in the additional learning.

4. The aforementioned processor, The learning device according to claim 1, characterized in that it generates a learning model for detecting the state in which an object enters a predetermined area corresponding to a dangerous area in which an object is moving, as the event of interest.

5. The aforementioned processor, The learning device according to claim 1, characterized in that an image representing a linear judgment line is superimposed and combined onto the captured image as the judgment criterion image.

6. The aforementioned processor, The learning device according to claim 1, characterized in that an image representing a polygonal judgment area is superimposed and combined onto the captured image as the judgment criterion image.

7. A learning method in which a processor performs a process to generate a learning model for detecting, based on images of a monitoring area, the state in which an object enters a predetermined area set on the captured images as a target event, The processing condition setting screen on which the captured image is placed is displayed, and the position of the judgment criterion image, which serves as the basis for determining whether an object has entered the predetermined area based on user operation, is set. By performing a superimposition process of the judgment criterion image on the captured image, a training target image is generated. The user is presented with the learning target image in which the judgment criterion image is superimposed and combined. Based on user operations, information regarding the occurrence status of the notable event that appeared in the training image is obtained as label information. A learning method characterized by generating the learning model by machine learning using the aforementioned learning target images and the aforementioned label information as learning information.

Citation Information

Patent Citations

  • Monitoring system and monitoring method

    JP2018205900A

  • Machine-learning device, machine-learning method, and recording medium having machine-learning program stored therein

    WO2021106028A1

  • Accident sign detection system and accident sign detection method

    WO2021205982A1