Control device, control method, and storage medium
Patent Information
- Application Number
- CN202280021394.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-24
- Filing Date
- 2022-01-20
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2042-01-20
AI Technical Summary
[0022]本发明的程序用于使计算机执行包括如下步骤的处理:能够切换第1监控模式及第2监控模式的步骤,所述第1监控模式通过使监控摄像机进行摄像获取第1摄像图像,且根据所提供的指示来改变摄像范围,所述第2监控模式通过使监控摄像机进行摄像获取第2摄像图像,且使用基于机器学习的学习完成模型,检测映入第2摄像图像的物体,并根据检测结果来改变摄像范围的步骤;及将在第1监控模式下获取的第1摄像图像作为针对机器学习的教师图像输出的步骤。
Smart Images

Figure CN117063480B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a control device, a control method, and a storage medium. Background Technology
[0002] Japanese Patent Application Publication No. 2004-056473 discloses the following: A monitoring and control device is equipped with a neural network (NW) that outputs recognition information corresponding to images captured by a camera based on learning results, a control mechanism that implements control based on the recognition information, a short-term storage mechanism that temporarily stores image data, and a storage mechanism that records the image data. The NW learns the relationship between the image and the urgency of the event represented by the image, identifies the urgency corresponding to the image from the camera, and the control mechanism controls the frame rate of the image data recorded to the storage mechanism through the recognition information of the NW.
[0003] Japanese Patent Publication No. 2006-523043 describes a method for detecting moving objects and controlling a monitoring system, comprising a processing module adapted to receive image information from at least one image forming sensor. The monitoring system performs motion detection analysis on the captured images and controls the camera in a specific manner when a moving object is detected.
[0004] The method and system configuration for a video surveillance system disclosed in Japanese Patent Application Publication No. 2009-516480 are as follows: The system includes multiple cameras, each with its own field of view. Each camera performs at least one of the following: zooming to change its field of view, tilting to rotate the camera around a horizontal tilt axis, and panning to rotate the camera around a vertical pan axis. The system also includes a processor configured to: receive signals representing images within the field of view of at least one camera; identify a target object using the received signals; determine a direction toward the target object from the camera that identified the target object; and transmit the determined direction to the other cameras among the multiple cameras. Summary of the Invention
[0005] One embodiment of the present invention provides a control device, control method, and storage medium capable of effectively collecting teacher images for machine learning.
[0006] means for solving technical problems
[0007] The control device of the present invention includes a processor for controlling a surveillance camera. In this control device, the processor performs the following processing: switching between a first monitoring mode and a second monitoring mode; the first monitoring mode involves the surveillance camera capturing a first image and changing the camera range according to provided instructions; the second monitoring mode involves the surveillance camera capturing a second image and using a machine learning-based learning completion model to detect objects reflected in the second image and changing the camera range according to the detection results; and outputting the first image acquired in the first monitoring mode as a teacher image for machine learning.
[0008] The processor preferably outputs the first camera image as the teacher image based on the manual operation performed on the surveillance camera.
[0009] Manual operation is the switching operation from the second monitoring mode to the first monitoring mode. The processor prefers to output the first camera image acquired in the first monitoring mode after switching from the second monitoring mode to the first monitoring mode as the teacher image.
[0010] The surveillance camera can change the field of view by changing at least one of panning, tilting and zooming. The switching operation is preferably the operation of at least one of panning, tilting and zooming when changing to the second monitoring mode.
[0011] After the processor switches from the second monitoring mode to the first monitoring mode, it outputs the first camera image as the teacher image according to the provided output instructions.
[0012] The processor preferably performs the following processing: assigning a judgment result that does not conform to the object detection to the second camera image acquired in the second monitoring mode before the switch and outputting it as a teacher image; and assigning a judgment result that conforms to the object detection to the first camera image acquired in the first monitoring mode after the switch and outputting it as a teacher image.
[0013] The processor preferably switches to the second monitoring mode after switching from the second monitoring mode to the first monitoring mode, and if no operation is performed in the first monitoring mode within a specified time.
[0014] The processor prefers that, after switching from the first monitoring mode to the second monitoring mode, and if a manual operation is performed again after a specified time since the last manual operation, the second camera image will not be output as the teacher's image.
[0015] The processor preferably detects objects reflected in the teacher's image and appends the position information of the detected objects within the teacher's image to the teacher's image.
[0016] The processor preferably sets the detection benchmark for object detection when detecting objects reflected in the teacher's image to be lower than the detection benchmark for object detection when detecting objects reflected in the second camera image.
[0017] The processor preferably performs location information change processing based on the provided instructions.
[0018] The processor preferably determines the position of the object mapped into the teacher image according to the provided instructions, and appends the determined position information of the object in the teacher image to the teacher image.
[0019] In addition to the teacher image, the processor preferably outputs an enlarged image generated by enlarging the teacher image.
[0020] The enlargement process is preferably at least one of the following: inversion, shrinking, noise addition, and deep learning-based style transformation.
[0021] The control method of the present invention includes: the step of being able to switch between a first monitoring mode and a second monitoring mode, wherein the first monitoring mode acquires a first video image by having a monitoring camera capture images and changes the camera range according to a provided instruction; the second monitoring mode acquires a second video image by having a monitoring camera capture images and uses a machine learning-based learning completion model to detect objects reflected in the second video image and changes the camera range according to the detection result; and the step of using the first video image acquired in the first monitoring mode as a teacher image output for machine learning.
[0022] The program of the present invention is used to enable a computer to perform a process including the following steps: being able to switch between a first monitoring mode and a second monitoring mode, wherein the first monitoring mode acquires a first video image by having a monitoring camera capture images and changes the camera range according to provided instructions; the second monitoring mode acquires a second video image by having a monitoring camera capture images and uses a machine learning-based learning completion model to detect objects reflected in the second video image and changes the camera range according to the detection results; and outputting the first video image acquired in the first monitoring mode as a teacher image for machine learning. Attached Figure Description
[0023] Figure 1 This is a schematic structural diagram illustrating an example of the overall structure of the monitoring system according to the first embodiment.
[0024] Figure 2 This is a block diagram illustrating an example of the hardware structure of a surveillance camera and its management device.
[0025] Figure 3This is a block diagram illustrating an example of the functions of the CPU included in a management device.
[0026] Figure 4 This is a conceptual diagram representing an example of manual PTZ in manual monitoring mode.
[0027] Figure 5 This is a concept diagram representing an example of object detection and processing.
[0028] Figure 6 This is a conceptual diagram representing an example of automatic PTZ in automatic monitoring mode.
[0029] Figure 7 This is a conceptual diagram representing an example of a falsely detected object in automatic monitoring mode.
[0030] Figure 8 This is a conceptual diagram representing an example of manual PTZ during automatic monitoring mode.
[0031] Figure 9 This is a concept map representing an example of learning processing.
[0032] Figure 10 This is a flowchart illustrating an example of the monitoring process involved in the first embodiment.
[0033] Figure 11 This is a flowchart illustrating an example of the monitoring process involved in the second embodiment.
[0034] Figure 12 This is a flowchart illustrating an example of the monitoring process involved in the third embodiment.
[0035] Figure 13 This is a flowchart illustrating an example of the monitoring process involved in the fourth embodiment.
[0036] Figure 14 This is a concept diagram representing the first variant of teacher image output processing.
[0037] Figure 15 This is a concept diagram representing the second variation of teacher image output processing.
[0038] Figure 16 This is a concept diagram representing the third variation of teacher image output processing.
[0039] Figure 17 This is a concept diagram representing the fourth variation of teacher image output processing.
[0040] Figure 18 This is a concept diagram illustrating a variation of object detection.
[0041] Figure 19This is a block diagram illustrating an example of how a program stored on a storage medium is installed in a computer. Detailed Implementation
[0042] Hereinafter, an example of the control device, control method and procedure involved in the present invention will be described with reference to the accompanying drawings.
[0043] First, let me explain the terms used in the following explanation.
[0044] CPU stands for Central Processing Unit. NVM stands for Non-volatile memory. RAM stands for Random Access Memory. IC stands for Integrated Circuit. ASIC stands for Application Specific Integrated Circuit. PLD stands for Programmable Logic Device. FPGA stands for Field-Programmable Gate Array. SoC stands for System-on-a-chip. SSD stands for Solid State Drive. USB stands for Universal Serial Bus. HDD stands for Hard Disk Drive. EEPROM stands for Electrically Erasable and Programmable Read Only Memory. EL stands for "Electro-Luminescence". I / F stands for "Interface". CMOS stands for "Complementary Metal Oxide Semiconductor". CCD stands for "Charge Coupled Device".
[0045] SWIR stands for "Short Wave Infrared". LAN stands for "Local Area Network".
[0046] [First Implementation]
[0047] As an example, such as Figure 1 As shown, the monitoring system 10 includes a monitoring camera 12 and a management device 16. The monitoring system 10 is, for example, a system for monitoring a construction site. The monitoring camera 12 is, for example, installed at a height such as on the roof of a building near the construction site. The management device 16 is used, for example, by a site supervisor or other user who oversees the workers at the construction site. The user uses the management device 16 to monitor whether any hazards occur at the construction site during work. The monitoring system 10 is a system designed to reduce the monitoring burden on users.
[0048] The surveillance camera 12 includes a camera device 18 and a rotating device 20. The camera device 18 captures images of a subject, for example, by receiving light in the visible wavelength band reflected by the subject. Alternatively, the camera device 18 can capture images of a subject by receiving near-infrared light, which is in the short-wave infrared wavelength band reflected by the subject. The short-wave infrared wavelength band refers, for example, to the wavelength band of approximately 900 nm to 2500 nm. Light in the short-wave infrared wavelength band is generally also referred to as SWIR light.
[0049] The camera device 18 is mounted on the rotating device 20. The rotating device 20 rotates the camera device 18. For example, the rotating device 20 changes the camera direction of the camera device 18 to a panning direction and a tilting direction. The panning direction, for example, refers to the horizontal direction. The tilting direction, for example, refers to the vertical direction.
[0050] The rotating device 20 includes a base 22, a panning rotating component 24, and a pitch rotating component 26. The panning rotating component 24 is cylindrical and mounted on the upper surface of the base 22. The pitch rotating component 26 is arm-shaped and mounted on the outer peripheral surface of the panning rotating component 24. A camera device 18 is mounted on the pitch rotating component 26. By rotating around a pitch axis TA parallel to the horizontal direction, the pitch rotating component 26 changes the camera direction of the camera device 18 to the pitch direction.
[0051] The base 22 supports the panning rotating component 24 from below. The panning rotating component 24 changes the imaging direction of the imaging device 18 to the panning direction by rotating about a panning axis PA that is parallel to the vertical direction.
[0052] A driving source (e.g., is built into the substrate 22) is provided. Figure 2(The panning motor 24A and the pitch motor 26A are shown). The drive source of the base 22 is mechanically connected to the panning motor 24A and the pitch motor 26A. For example, the drive source of the base 22 is connected to the panning rotating member 24 and the pitch rotating member 26 via a power transmission mechanism (not shown). The panning rotating member 24 rotates about the panning axis PA by receiving power from the drive source of the base 22, and the pitch rotating member 26 rotates about the pitch axis TA by receiving power from the drive source of the base 22.
[0053] like Figure 1 As shown, the monitoring system 10 generates video images by capturing images within a camera range 31 set within the monitoring area 30 using the camera device 18. The monitoring system 10 performs panning and tilting movements, capturing the entire monitoring area 30 by changing the camera range 31. The construction site, which is the monitoring area 30, contains various subjects such as heavy equipment and workers. Heavy equipment includes loaders, bulldozers, cranes, dump trucks, etc.
[0054] The imaging device 18 is, for example, a digital video camera with an image sensor (not shown). The image sensor receives subject light representing the subject, performs photoelectric conversion on the received subject light, and outputs an electrical signal with a signal level corresponding to the amount of light received as image data. The image data output by the image sensor corresponds to the captured image. The image sensor is a CMOS image sensor or a CCD image sensor, etc. The imaging device 18 can capture color images or monochrome images. Furthermore, the captured image can be a still image or a moving image.
[0055] Furthermore, the camera device 18 has a zoom function. The zoom function refers to the ability to reduce or increase the field of view 31 (i.e., zoom in or zoom out). The zoom function of the camera device 18 can be an optical zoom function performed by moving the zoom lens or an electronic zoom function performed by image processing of the image data. Alternatively, the zoom function of the camera device 18 can be a combination of optical zoom and electronic zoom.
[0056] The management device 16 includes a management device main body 13, a receiving device 14, and a display 15. The management device main body 13 has a built-in computer 40 (reference). Figure 2 The monitoring system 10 is controlled by a receiver 14 and a display 15 connected to the main body 13 of the management device.
[0057] The receiving device 14 receives various instructions from the user using the monitoring system 10. Examples of receiving devices 14 include a keyboard, mouse, and / or touch panel. The management unit 13 manages the various instructions received using the receiving device 14. The display 15 displays various information (e.g., images and text) under the control of the management unit 13. Examples of display devices 15 include a liquid crystal display (LCD) or an EL display.
[0058] The surveillance camera 12 is communicatively connected to a communication network NT (such as the Internet or LAN) via the management device 16 and operates under the control of the management device main body 13. The connection between the surveillance camera 12 and the management device 16 can be either a wired connection or a wireless connection.
[0059] The management device 16 acquires video images output from the camera device 18 of the surveillance camera 12 and uses a machine learning-based learning completion model to detect specific objects (e.g., heavy equipment) reflected in the video images. When a specific object is detected, the management device 16 causes the surveillance camera 12 to pan, tilt, and zoom to track the detected object. Hereinafter, the action of changing the camera range 31 by panning, tilting, and zooming will be referred to as "PTZ". In addition, the action of changing the camera range 31 according to the detection result of the object reflected in the video images will be referred to as "automatic PTZ".
[0060] Furthermore, the management device 16 is capable of changing the camera range 31 based on the user's operation of the receiving device 14. Hereinafter, the action of changing the camera range 31 according to the instructions provided to the receiving device 14 will be referred to as "manual PTZ". In manual PTZ, the user can set the camera range 31 to any position and size within the monitoring area 30 by operating the receiving device 14.
[0061] Furthermore, the monitoring mode of monitoring the monitoring area 30 via manual PTZ will be referred to as "manual monitoring mode," and the monitoring mode of monitoring the monitoring area 30 via automatic PTZ will be referred to as "automatic monitoring mode." Users can switch between manual and automatic monitoring modes of the monitoring system 10. The manual monitoring mode is an example of the "first monitoring mode" according to the technology of this invention. The automatic monitoring mode is an example of the "second monitoring mode" according to the technology of this invention.
[0062] As an example, such as Figure 2 As shown, the rotating device 20 of the surveillance camera 12 includes a controller 34. The controller 34 controls the panning motor 24A, the tilt motor 26A, and the camera device 18 under the control of the management device 16.
[0063] The main body 13 of the management device 16 includes a computer 40. The computer 40 includes a CPU 42, an NV M 44, RAM 46, and a communication I / F 48. The management device 16 is an example of a "control device" according to the technology of this invention. The computer 40 is an example of a "computer" according to the technology of this invention. The CPU 42 is an example of a "processor" according to the technology of this invention.
[0064] CPU42, NVM44, RAM46, and communication I / F48 are connected to bus 49. Figure 2 In the example shown, for ease of illustration, bus 49 is represented as a single bus, but multiple buses can also be represented. Bus 49 can be a serial bus or a parallel bus that includes a data bus, an address bus, and a control bus.
[0065] NVM44 stores various types of data. Examples of NVM44 include various non-volatile storage devices such as EEPROM, SSD, and / or HDD. RAM46 temporarily stores various information and is used as working memory. Examples of RAM46 include DRAM or SRAM.
[0066] The NVM44 stores the program PG. The CPU42 reads the required program from the NVM44 and executes the read program PG on the RAM46. By executing the process according to the program PG, the CPU42 controls the entire monitoring system 10, including the management device 16.
[0067] The communication I / F48 is an interface implemented using hardware resources such as FPGA. The communication I / F48 can be communicatively connected to the controller 34 of the surveillance camera 12 via the communication network NT, and various information is transmitted and received between the CPU 42 and the controller 34.
[0068] The bus 49 is also connected to a receiver 14 and a display 15. The CPU 42 performs actions according to the instructions received through the receiver 14 and displays various information on the display 15.
[0069] Furthermore, the NVM44 stores a learned model LM used for the aforementioned object detection. The learned model LM is a machine learning model for object detection generated by performing machine learning on multiple teacher images mapped to a specific object. Additionally, the NVM44 stores teacher images TD. Teacher images TD are teacher images used for further learning of the learned model LM. Teacher images TD are images that meet specified conditions from the video images acquired by the surveillance camera 12.
[0070] As an example, such as Figure 3As shown, the CPU 42 executes actions according to the program PG, thereby realizing multiple functional units. The program PG enables the CPU 42 to function as a camera control unit 50, a mode switching control unit 51, an image acquisition unit 52, a display control unit 53, an object detection unit 54, a teacher image output unit 55, and a machine learning unit 56.
[0071] The camera control unit 50 controls the controller 34 of the surveillance camera 12 to cause the camera device 18 to perform camera shooting and zooming, and to cause the rotating device 20 to perform panning and tilting. That is, the camera control unit 50 causes the surveillance camera 12 to perform camera shooting and changes the camera range 31.
[0072] The mode switching control unit 51 controls the switching between automatic and manual monitoring modes of the monitoring system 10 based on instructions received from the receiving device 14. When the mode switching control unit 51 is in manual monitoring mode, it causes the camera control unit 50 to manually adjust the camera range 31 according to instructions provided to the receiving device 14. When the mode switching control unit 51 is in automatic monitoring mode, it performs automatic PTZ adjustment of the camera range 31 based on object detection results from the object detection unit 54.
[0073] The image acquisition unit 52 captures a video image P output from the monitoring camera 12 by causing the monitoring camera 12 to capture an image via the camera control unit 50. The image acquisition unit 52 then supplies the video image P acquired from the monitoring camera 12 to the display control unit 53. The display control unit 53 displays the video image P supplied by the image acquisition unit 52 on the display 15.
[0074] In manual monitoring mode, the image acquisition unit 52 supplies the camera image P acquired from the monitoring camera 12 as the first camera image P1 to the teacher image output unit 55. On the other hand, in automatic monitoring mode, the image acquisition unit 52 supplies the camera image P acquired from the monitoring camera 12 as the second camera image P2 to the object detection unit 54.
[0075] The object detection unit 54 uses the learned model LM stored in the NVM 44 to detect specific objects (e.g., heavy equipment) projected into the second camera image P2. The object detection unit 54 supplies the detection results to the display control unit 53 and the camera control unit 50. Based on the detection results supplied from the object detection unit 54, the display control unit 53 displays the detected objects recognizablely on the display 15. Based on the detection results supplied from the object detection unit 54, the camera control unit 50 adjusts the camera range 31 so that the detected object is located in the center of the camera range 31 and expands the scope of the detected object.
[0076] The teacher image output unit 55 stores the first camera image P1 as a teacher image TD in the NVM 44 based on the manual operation performed by the user on the monitoring camera 12 using the receiving device 14. In this embodiment, the teacher image output unit 55 stores the first camera image P1 acquired in the manual monitoring mode as a teacher image TD in the NVM 44 when the user switches from automatic monitoring mode to manual monitoring mode using the receiving device 14. The switching operation refers to changing at least one of panning, tilting, and zooming.
[0077] The machine learning unit 56 uses the teacher images TD stored in the NVM44 to perform supplementary learning on the learned model LM, thereby updating the learned model LM. For example, when a predetermined number of teacher images TD have been accumulated in the NVM44, the machine learning unit 56 uses the accumulated teacher image TDs to perform supplementary learning on the learned model LM. After updating the learned model LM, the object detection unit 54 uses the updated learned model LM to perform object detection.
[0078] The learning-complete model (LM) is constructed using neural networks. For example, the learning-complete model (LM) is constructed using a multi-layered neural network, namely a deep neural network (DNN), which is the object of deep learning. As a DNN, for example, a convolutional neural network (CNN) is used that treats images as objects.
[0079] Figure 4 This represents an example of manual PTZ in manual monitoring mode. Figure 4 In the example shown, two heavy pieces of equipment H1 and H2 are reflected as objects in the camera image P displayed on the monitor 15. The user can change the camera range 31 to the area of interest by operating a keyboard or mouse, which serves as the receiving device 14. Figure 4 The diagram shows that the heavy equipment H2, in which there are people nearby, is the object of monitoring, and the operation of changing the camera range 31 is performed so that the area of concern including the heavy equipment H2 is consistent with the camera range 31.
[0080] Figure 5 This illustrates an example of object detection processing performed by the object detection unit 54, which uses a learned complete model (LM). In this embodiment, the learned complete model (LM) is constructed using a CNN. The object detection unit 54 inputs the second camera image P2 as an input image to the learned complete model (LM). The learned complete model (LM) generates a feature map FM representing the feature quantities of the second camera image P2 through convolutional layers.
[0081] The object detection unit 54 slides windows W of various sizes across the feature map FM to determine whether object candidates exist within the window W. If the object detection unit 54 determines that object candidates exist within the window W, it cuts out an image R from the feature map FM that includes the object candidates within the window W, and inputs the cut image R into the classifier. The classifier outputs the labels and scores of the object candidates contained in the image R. The label indicates the type of object. The score indicates the probability that the object candidate belongs to the type represented by the label. Figure 5 In the example shown, heavy equipment H1 is extracted as an object candidate, and the classifier determines that the label for heavy equipment H1 is "forklift". Furthermore, the score representing the probability that heavy equipment H1 is "forklift" is "0.90".
[0082] The object detection unit 54 outputs the location information, label, and score of the image R containing objects with a score of a predetermined value or higher as the detection result. Additionally, in Figure 5 In the example shown, one object is detected from the second camera image P2, but sometimes more than two objects are detected.
[0083] As an example, such as Figure 6 As shown, the display control unit 53 displays a rectangular frame F within the camera image P, surrounding the object detected by the object detection unit 54. Furthermore, the display control unit 53 displays a label L near the frame F indicating the type of object within the frame F. Additionally, the display control unit 53 can also display a score. Moreover, when two or more objects are detected, the display control unit 53 displays multiple frames F within the camera image P.
[0084] Figure 6 This represents an example of automatic PTZ in automatic monitoring mode. Figure 6 In the example shown, two heavy equipment pieces H1 and H2 are reflected as objects in the camera image P displayed on the monitor 15. The object detection unit 54 detects heavy equipment H1, but heavy equipment H2 is not detected as an object. In this case, the camera control unit 50 controls the change of the camera range 31 so that the area including heavy equipment H1 is consistent with the camera range 31. As a result, automatic PTZ is performed to track heavy equipment H1.
[0085] Furthermore, if there are two or more objects detected by the object detection unit 54 in the camera image P, for example, the camera control unit 50 controls the change of the camera range 31 so that the area including the object with the highest score is consistent with the camera range 31.
[0086] Figure 7 This illustrates an example where the object detection unit 54 falsely detects an object in automatic monitoring mode. Figure 7In the example shown, a car that is not heavy equipment was mistakenly detected as a forklift, and automatic PTZ was performed to track the car. However, when such a misdetection occurs, it is possible that the monitored object exists in another area. For example, in Figure 7 In the example shown, the user focuses on heavy equipment H2 where people are nearby as the monitoring target. In this case, the user operates the receiver 14 to perform manual PTZ so that the area of interest including the heavy equipment H2 is aligned with the camera range 31. Thus, after an object is detected in automatic monitoring mode, if the object is not the user's desired monitoring target, the user sometimes operates the receiver 14 to perform manual PTZ.
[0087] Figure 8 This indicates an example of performing manual PTZ in automatic monitoring mode. Figure 8 In the example shown, based on the automatic PTZ caused by the car falsely detected by the object detection unit 54, the user performs manual PTZ to focus on the desired area of interest (see reference). Figure 7 The camera range is consistent with 31. In automatic monitoring mode, when the user operates the receiving device 14 to perform manual PTZ, the above-mentioned mode switching control unit 51 switches the monitoring mode from automatic monitoring mode to manual monitoring mode.
[0088] The teacher image output unit 55 outputs the first camera image P1 as the teacher image TD when the monitoring mode is switched from automatic monitoring mode to manual monitoring mode. For example, when the user performs manual PTZ, thus switching the monitoring mode to manual monitoring mode, the teacher image output unit 55 outputs the first camera image P1 at the moment when the manual PTZ stops as the teacher image TD.
[0089] Figure 9 This is an example of a learning process that uses a teacher image TD to perform additional learning on a learned model LM. In the learning process, the teacher image TD is input into the learned model LM. A positive label L1 representing the type of object contained in the first camera image P1 is assigned to the teacher image TD. The objects mapped into the first camera image P1 are objects that were not detected by the object detection unit 54, so the positive label L1 is assigned, for example, by the user's judgment of the type of object.
[0090] The learned model LM outputs a detection result RT based on the input teacher image TD. The detection result RT consists of the aforementioned label L and score. Based on this detection result RT and the positive resolution label L1, a loss function is applied. Furthermore, the various coefficients (weight coefficients, bias, etc.) of the learned model LM are updated according to the results of the loss function, and the learned model LM is updated accordingly.
[0091] Alternatively, label L can be a simple label indicating whether the detected object is a correct match (e.g., whether it is heavy equipment). In this case, for example, label L can be represented by two values, "1" or "0", with the correct match label L1 set to "1" and the incorrect match label L0 set to "0". Furthermore, the correct match label L1 is an example of a "judgment result that matches the detection of an object" according to the technology of this invention. The incorrect match label L0 is an example of a "judgment result that does not match the detection of an object" according to the technology of this invention.
[0092] In this way, the object that the user manually sets as the monitoring target is highly likely to be the correct detection, so the first camera image P1, which is assigned the correct detection label L1, is output as the teacher image TD. Thus, by using the teacher image TD to supplement the learned model LM, the accuracy of object detection is improved. Furthermore, it is possible to detect new types of objects. For example, if a bulldozer, as a type of heavy equipment, cannot be detected in automatic monitoring mode, the user can supplement the learning in manual monitoring mode with the teacher image TD including the bulldozer as the monitoring target, thereby enabling the bulldozer to be detected as heavy machinery again.
[0093] Next, refer to Figure 10 The function of monitoring system 10 will be explained.
[0094] Figure 10 This is a flowchart illustrating an example of the monitoring process executed via CPU42. Additionally, Figure 10 The monitoring process shown is an example of the "control method" involved in the technology of this invention. Furthermore, for ease of explanation, the description will be based on the premise that the video recording performed by the camera device 18 is executed at a predetermined frame rate.
[0095] Figure 10 In the monitoring process shown, firstly, in step S10, the mode switching control unit 51 causes the camera control unit 50 to start operating in automatic monitoring mode. When the automatic monitoring mode is started, the monitoring camera 12 controls the camera range 31 (refer to the monitoring area 30) set within the monitoring area 30. Figure 1 The camera is used as the object for recording. After step S10, the monitoring process moves to step S11.
[0096] In step S11, the image acquisition unit 52 acquires the camera image P output from the monitoring camera 12 and supplies it to the object detection unit 54 as the second camera image P2. At this time, the camera image P is displayed on the display 15 via the display control unit 53. After step S11, the monitoring process proceeds to step S12.
[0097] In step S12, the object detection unit 54 uses the learned model LM to perform object detection processing to detect specific objects (e.g., heavy equipment) projected into the second camera image P2 (see reference). Figure 5 After step S12, the monitoring process moves to step S13.
[0098] In step S13, the camera control unit 50 determines whether the object detection unit 54 has detected an object. If no object is detected in step S13, the determination is negative, and the monitoring process proceeds to step S14. If an object is detected in step S13, the determination is positive, and the monitoring process proceeds to step S15.
[0099] In step S14, the camera control unit 50 changes the camera range 31 to the panning or tilting direction by causing the monitoring camera 12 to pan or tilt (see reference). Figure 1 After step S14, the monitoring process returns to step S11. In step S11, the image acquisition unit 52 performs camera image acquisition processing again.
[0100] In step S15, the camera control unit 50 performs automatic PTZ (referencing) by changing the camera range 31 based on the detection result of the object detected by the object detection unit 54. Figure 6 and Figure 7 After step S15, the monitoring process moves to step S16.
[0101] In step S16, the mode switching control unit 51 determines whether the monitoring mode has switched from automatic monitoring mode to manual monitoring mode by manually operating the receiving device 14 via user operation. If the mode has not switched to manual monitoring mode in step S16, the determination is negative, and the monitoring process returns to step S15. If the mode has switched to manual monitoring mode in step S16 (see reference...), the determination is negative. Figure 8 If the condition is affirmative, the monitoring process proceeds to step S17. For example, in automatic monitoring mode, when the user operates the receiving device 14 to perform manual PTZ, the condition is affirmative.
[0102] In step S17, the camera control unit 50 performs manual PTZ (referencing) to change the camera range 31 according to the instruction provided by the user to the receiving device 14. Figure 4 After step S17, the monitoring process moves to step S18.
[0103] In step S18, the teacher image output unit 55 outputs the first camera image P1 acquired in manual monitoring mode as the teacher image TD (reference). Figure 8 After step S18, the monitoring process moves to step S19.
[0104] In step S19, the mode switching control unit 51 determines whether the monitoring mode has switched from manual monitoring mode to automatic monitoring mode by user operation of the receiving device 14. If the mode has not switched to automatic monitoring mode in step S19, the determination is negative, and the monitoring process proceeds to step S20. If the mode has switched to automatic monitoring mode in step S19, the monitoring process returns to step S10.
[0105] In step S20, the mode switching control unit 51 determines whether the conditions for ending the monitoring process (hereinafter referred to as "ending conditions") are met. One example of an ending condition is when the receiving device 14 receives an instruction to end the monitoring process. In step S20, if the ending condition is not met, the determination is negative, and the monitoring process returns to step S17. In step S20, if the ending condition is met, the determination is positive, and the monitoring process ends.
[0106] As explained above, the management device 16, acting as a control device, can switch between manual monitoring mode and automatic monitoring mode. In manual monitoring mode, the monitoring camera 12 captures a first image P1 and adjusts the camera range 31 according to provided instructions. In automatic monitoring mode, the monitoring camera 12 captures a second image P2 and uses a machine learning-based learning completion model (LM) to detect objects reflected in the second image P2, adjusting the camera range 31 based on the detection results. Furthermore, the management device 16 outputs the first image P1 acquired in manual monitoring mode as a teacher image TD for machine learning. Thus, according to the technology of the present invention, the user can effectively collect teacher images TD for machine learning without special operation.
[0107] Furthermore, the management device 16 outputs the first camera image P1 as the teacher image TD based on the manual operation performed on the monitoring camera 12. This manual operation is a switch from automatic monitoring mode to manual monitoring mode, and the management device 16 outputs the first camera image acquired during the switch from automatic to manual monitoring mode as the teacher image TD. Moreover, the monitoring camera 12 can change the camera range 31 by changing at least one of panning, tilting, and zooming; the switching operation is an operation of at least one of panning, tilting, and zooming when changing from automatic monitoring mode. Thus, according to the technology of the present invention, the teacher image TD can be effectively collected according to the user's intention.
[0108] [Second Implementation]
[0109] In the first embodiment, an example is shown where the first camera image is output as a teacher image TD based on the case where the monitoring mode is switched from automatic monitoring mode to manual monitoring mode. In the second embodiment, the first camera image is output as a teacher image TD based on the output instruction provided by the user.
[0110] Figure 11 This indicates the function of the monitoring system 10 according to the second embodiment. For example... Figure 11 As shown, in this embodiment, step S30 is added between step S17 and step S18. The other steps are the same as in the first embodiment.
[0111] In this embodiment, after manual PTZ is started in step S17, the monitoring process is transferred to step S30.
[0112] In step S30, it is determined whether the user has given an output instruction by operating the receiving device 14. For example, the user operates the mouse, which acts as the receiving device 14, and clicks a dedicated button displayed on the display 15 to give an output instruction. If an output instruction was given in step S30, the determination is affirmative, and the monitoring process proceeds to step S18. If no output instruction was given in step S30, the determination is negative, and the monitoring process proceeds to step S19.
[0113] In step S18, similar to the first embodiment, the teacher image output unit 55 performs teacher image output processing, which outputs the first camera image P1 acquired in manual monitoring mode as the teacher image TD.
[0114] Thus, in this embodiment, after the management device 16 switches from automatic monitoring mode to manual monitoring mode, it uses the first camera image P1 as the teacher image TD according to the provided output instruction, thereby enabling the effective collection of the teacher image TD according to the user's expectations.
[0115] [Third Implementation]
[0116] In the first embodiment, an example is shown where the first camera image is output as the teacher image TD when switching from automatic monitoring mode to manual monitoring mode. In the third embodiment, in addition to the first camera image, the second camera image P2 acquired in the automatic monitoring mode before the switch is also output as the teacher image TD.
[0117] Figure 12 This indicates the function of the monitoring system 10 according to the third embodiment. For example... Figure 12 As shown, in this embodiment, step S40 is added between step S16 and step S17. The other steps are the same as in the first embodiment.
[0118] In this embodiment, in step S16, when switching to manual monitoring mode, the determination is affirmative, and the monitoring process is transferred to step S40.
[0119] In step S40, the teacher image output unit 55 will display the second camera image P2 (reference image) acquired in the automatic monitoring mode before the switch. Figure 8 The image P2 is output as the teacher's image TD. When the user switches from automatic monitoring mode to manual monitoring mode, for the second camera image P2 acquired in the previous automatic monitoring mode, as shown... Figure 8 As shown, the object detection performed by the object detection unit 54 is considered a false detection, therefore the teacher image output unit 55 outputs the second camera image P2 as a teacher image TD with a non-correct detection label L0. After step S40, the monitoring process proceeds to step S17.
[0120] In addition, in this embodiment, in step S18, the teacher image output unit 55 outputs the first camera image P1 acquired in manual monitoring mode as a teacher image TD with the correct solution label L1.
[0121] Thus, in this embodiment, the management device 16 outputs the second camera image P2, acquired in the automatic monitoring mode before the switch, as the teacher image TD with a non-positive resolution label L0, and the first camera image P1, acquired in the manual monitoring mode after the switch, as the teacher image TD with a positive resolution label L1. Therefore, in this embodiment, the positive resolution label L1 or the non-positive resolution label L0 can be automatically assigned to the teacher image TD, reducing the user's workload. Furthermore, in addition to the first camera image P1, the second camera image P2 is used to supplement the learning of the completed model LM, thereby improving the detection accuracy of object detection.
[0122] [Fourth Implementation]
[0123] Next, the fourth embodiment will be described. The fourth embodiment is a variation of the third embodiment. In the third embodiment, when the monitoring mode is switched from automatic monitoring mode to manual monitoring mode, the second camera image P2 acquired in the previous automatic monitoring mode is output as the teacher image TD. In the fourth embodiment, after the monitoring mode is switched from automatic monitoring mode to manual monitoring mode, if certain conditions are met, the second camera image P2 acquired in the previous automatic monitoring mode is output as the teacher image TD.
[0124] Figure 13This describes the function of the monitoring system 10 according to the fourth embodiment. In this embodiment, in step S19, the mode switching control unit 51 determines whether, after switching from automatic monitoring mode to manual monitoring mode, no operation has been performed within a predetermined time in manual monitoring mode (i.e., whether the state of no operation has lasted for a predetermined time). In step S19, if no operation has been performed within the predetermined time, it is determined to be positive, and the monitoring process returns to step S10. In step S19, if an operation has been performed before the predetermined time has elapsed, it is determined to be negative, and the monitoring process proceeds to step S20. That is, in this embodiment, after switching from automatic monitoring mode to manual monitoring mode, if no operation has been performed within the predetermined time in manual monitoring mode, the process returns to automatic monitoring mode.
[0125] Furthermore, in this embodiment, step S50 is added between step S16 and step S40. The other steps are the same as in the third embodiment.
[0126] In this embodiment, in step S16, when switching to manual monitoring mode, the determination is affirmative, and the monitoring process is transferred to step S50.
[0127] In step S50, the mode switching control unit 51 determines whether a predetermined time has elapsed since the last switching operation. Specifically, if the mode switching control unit determines "yes" in step S16, it starts timing from the moment the monitoring mode is switched to manual monitoring mode. If the mode switching control unit determines "yes" in step S19, it switches the monitoring mode to automatic monitoring mode. The time elapsed until the mode switching control unit determines "yes" again in step S16 is then used to determine whether the predetermined time has elapsed.
[0128] In step S50, if the current switching operation has not elapsed within the specified time since the previous switching operation, it is determined to be negative, and the monitoring process proceeds to step S40. In step S50, if the current switching operation has elapsed within the specified time since the previous switching operation, it is determined to be positive, and the monitoring process proceeds to step S17.
[0129] Thus, in this embodiment, if a predetermined time has elapsed since the last switching operation, the teacher image output unit 55 will not output the second camera image P2 acquired in the automatic monitoring mode before the switch as the teacher image TD. This corresponds to a situation where, for example, after a user switches the monitoring mode to manual monitoring mode, they remain inactive due to leaving the location of the management device 16, thus switching back to automatic monitoring mode, and then switching back to manual monitoring mode upon returning to the location of the management device 16. This is because, in such a situation, there is a high probability that the user will not observe the second camera image P2 acquired in automatic monitoring mode before switching to manual monitoring mode, and it cannot be considered that the object detection is a false detection, thus preventing the switching of the monitoring mode to manual monitoring mode. In other words, this is because it is assumed that the user switched from manual monitoring mode to automatic monitoring mode due to a continuous state of inactivity, and the switching operation was performed to easily return to manual monitoring mode.
[0130] Thus, in this embodiment, it is possible to prevent the second camera image P2 from being output as the teacher image TD under circumstances that are not desired by the user.
[0131] then, Figures 14-17 Various variations of teacher image output processing performed by the teacher image output unit 55 are shown.
[0132] [First Variation]
[0133] Figure 14 This represents the first variation of the teacher's image output processing. For example... Figure 14 As shown, in the first variant, the teacher image output unit 55 detects objects reflected in the teacher image TD and appends the position information of the detected objects in the teacher image TD to the teacher image TD.
[0134] For example, the teacher image output unit 55 uses the learned completion model LM to detect objects from the first camera image P1, which is the output object of the teacher image TD, and appends the position information of the detected objects to the first camera image P1. Furthermore, the teacher image output unit 55 outputs the first camera image P1 with the appended position information as the teacher image TD.
[0135] In addition, when the teacher image output unit 55 takes the second camera image P2 as the output object, the same position information additional processing can be performed on the second camera image P2.
[0136] Furthermore, the detection benchmark for object detection when the teacher image output unit 55 uses the learned model LM to detect objects is preferably lower than the detection benchmark for object detection when the object detection unit 54 uses the learned model LM to detect objects. For example, the detection benchmark is the lower limit of the score at which an object candidate is determined to be a specific object. For example, when the object detection unit 54 uses the learned model LM to detect objects, if the score is 0.9 or higher, the object is determined to be heavy equipment; when the teacher image output unit 55 uses the learned model LM to detect objects, if the score is 0.7 or higher, the object is determined to be heavy equipment.
[0137] In this way, by making the detection benchmark of object detection when the object is mapped into the teacher image TD lower than the detection benchmark of object detection when the object is mapped into the second camera image P2 in automatic monitoring mode, the detection accuracy of the learning-completed model LM is improved, and objects that could not be detected until now are detected.
[0138] [Second Variation]
[0139] Figure 15 This represents the second variation of the teacher's image output processing. For example... Figure 15 As shown, in the second variation, in addition to the location information additional processing shown in the first variation, the user can also change the location information.
[0140] In this embodiment, the teacher image output unit 55 uses a learning completion model (LM) to detect objects from the first camera image P1, which is the output object of the teacher image TD, and displays the position information of the detected objects together with the first camera image P1 on the display 15 via the display control unit 53. The user can change and determine the position information displayed on the display 15. For example, the user can use the receiving device 14 to change and determine the position, shape, and size of the rectangle representing the object's position information. Figure 15 In the example shown, when the user detects a person who is not a heavy machine as an object by learning the model LM, the user changes the location information so that the location information represents the area of the heavy machine H2.
[0141] The teacher image output unit 55 changes the position information according to the instruction provided to the receiving device 14, and outputs the first camera image P1 with the changed position information as the teacher image TD.
[0142] In addition, when the teacher image output unit 55 uses the second camera image P2 as the output object, the same position information change processing can be performed on the second camera image P2.
[0143] Thus, according to this variation, the user can change the location information to an appropriate location, thereby improving the accuracy of the additional learning of the learned model LM.
[0144] [3rd Variation]
[0145] In the third variation, the teacher image output unit 55 does not perform object detection using the learning completion model LM, but determines the position of the object mapped into the teacher image TD according to the instructions provided by the user, and appends the determined position information of the object in the teacher image TD to the teacher image TD.
[0146] As an example, such as Figure 16 As shown, the teacher image output unit 55 displays the first camera image P1, which is the output object of the teacher image TD, on the display 15 via the display control unit 53. Based on instructions provided to the receiving device 14, the teacher image output unit 55 determines the position of the heavy equipment H2 mapped into the first camera image P1, which is the output object of the teacher image TD, and outputs the first camera image P1 with the position information of the heavy equipment H2 appended as the teacher image TD. For example, by using the receiving device 14 to change the position, shape, and size of the rectangle representing the position information of the object, the user can determine the position of the object.
[0147] Similarly, when the second camera image P2 is the output object, the teacher image output unit 55 can also add position information according to the instructions provided by the user.
[0148] According to this variation, the user can determine the position of the object mapped into the teacher image TD, thus improving the accuracy of machine learning in the learning completion model LM.
[0149] [4th Variation]
[0150] In the fourth variation, the teacher image output unit 55 performs data augmentation by enlarging the teacher image TD to further improve the accuracy of the machine learning of the learned model LM. As an example, such as... Figure 17 As shown, in addition to the first camera image P1, which is the output object of the teacher image TD, the teacher image output unit 55 also outputs an enlarged image P1E of the first camera image P1 that has been inverted as the teacher image TD. Therefore, by increasing the number of teacher images TD, the accuracy of machine learning in the learning completion model LM is improved.
[0151] Furthermore, the enlargement process used to generate the enlarged image P1E is not limited to inversion. The enlargement process can be at least one of inversion, reduction, noise addition, and deep learning-based style transformation.
[0152] Similarly, when the second camera image P2 is used as the output object, the teacher image output unit 55 can increase the number of teacher images TD by performing enhancement processing.
[0153] Furthermore, the various processing described in the first to fourth variations can be performed after the teacher image TD output from the teacher image output unit 55 is stored in a storage device such as the NVM44.
[0154] In the above-described embodiments and variations, heavy equipment as an object is detected from the camera image, but the detected object is not limited to heavy equipment. For example, such as Figure 18 As shown, roadblocks B placed around heavy equipment H1 to ensure safety can be detected. Alternatively, after heavy equipment H1 is detected, it can be checked periodically to see if roadblocks B are placed around it. The same techniques described above can be applied to roadblock detection as to the case of heavy equipment.
[0155] The technology of this invention is particularly useful in situations where it is difficult to obtain images of teachers, such as heavy equipment or roadblocks at construction sites.
[0156] Furthermore, in the above embodiments, in NVM44 (reference) Figure 2 The invention stores a monitoring and processing program PG, but the technology of the present invention is not limited thereto. For example, such as Figure 19 As shown, the program PG can also be stored in any portable storage medium 100 that is a non-temporary storage medium such as an SSD or USB memory. In this case, the program PG stored in the storage medium 100 is installed in the computer 40, and the CPU 42 performs the aforementioned monitoring process based on the program PG.
[0157] Furthermore, the program PG can be pre-stored on the storage device of another computer or server device connected to computer 40 via a communication network (not shown), and the program PG can be downloaded to computer 40 and installed upon request from management device 16. In this case, monitoring processing is performed by computer 40 based on the installed program PG.
[0158] As the hardware resource for performing the aforementioned monitoring processing, various processors, as shown below, can be used. For example, as described above, a general-purpose processor, i.e., a CPU, can function as the hardware resource for performing monitoring processing by executing software, i.e., a program PG. Furthermore, as processors, special-purpose circuits, such as FPGAs, PLDs, or ASICs, can be used, which have circuit structures specifically designed for performing particular processes. Memory is built into or connected to any processor, and any processor performs monitoring processing using memory.
[0159] The hardware resources for performing monitoring processing can consist of one of these various processors, or a combination of two or more processors of the same or different types (e.g., a combination of multiple FPGAs or a combination of a CPU and an FPGA). Furthermore, the hardware resources for performing monitoring processing can also be a single processor.
[0160] As examples of processors, firstly, there is the method where, as exemplified by user-end computers and servers, a processor is constructed using a combination of one or more CPUs and software, and this processor functions as a hardware resource for performing monitoring processing. Secondly, there is the method where, as exemplified by SOC, a processor is used where a single IC chip implements the functions of the entire system, including multiple hardware resources for performing monitoring processing. In this way, monitoring processing is implemented using one or more of the aforementioned processors as hardware resources.
[0161] Furthermore, the hardware architecture of these various processors, more specifically, enables the use of circuits composed of circuit elements such as semiconductor components.
[0162] Furthermore, the monitoring process described above is merely one example. Therefore, without departing from the intended purpose, unnecessary steps can certainly be removed, new steps can be added, or the processing order can be changed.
[0163] The descriptions and illustrations above are detailed explanations of the parts involved in the technology of this invention, and are merely one example of the technology of this invention. For example, the descriptions related to the above-described structure, function, effect, and effect are examples of the structure, function, effect, and effect of the parts involved in the technology of this invention. Therefore, without departing from the technical spirit of this invention, unnecessary parts may be deleted from the descriptions and illustrations above, or new elements may be added or replaced. Furthermore, to avoid complications and to facilitate understanding of the parts involved in the technology of this invention, explanations of technical common sense, etc., that do not require special explanation in aspects enabling the implementation of this invention have been omitted from the descriptions and illustrations above.
[0164] In this specification, "A and / or B" has the same meaning as "at least one of A and B". That is, "A and / or B" can mean only A, only B, or a combination of A and B. Furthermore, in this specification, the same approach applies to situations where three or more items are connected by "and / or".
[0165] All documents, patent applications and technical standards described in this specification, and the specific and separately described documents, patent applications and technical standards incorporated herein by reference, are incorporated herein by reference to the same extent.
Claims
1. A control device comprising a processor for controlling a surveillance camera, wherein, The processor performs the following processing: Switch between a first monitoring mode and a second monitoring mode. The first monitoring mode involves capturing a first image using a monitoring camera and changing the camera range according to a provided instruction. The second monitoring mode involves capturing a second image using the monitoring camera and using a machine learning-based learning completion model to detect objects reflected in the second image and changing the camera range according to the detection results. The first camera image acquired in the first monitoring mode will be used as the teacher image output for the machine learning. If, after the processor switches from the first monitoring mode to the second monitoring mode, and a certain period of time has elapsed since the last manual operation that switched from the second monitoring mode to the first monitoring mode, the second camera image will not be output as the teacher image.
2. The control device according to claim 1, wherein, The processor outputs the first camera image as the teacher image based on the manual operation performed on the surveillance camera.
3. The control device according to claim 2, wherein, The processor will output the first camera image acquired in the first monitoring mode after switching from the second monitoring mode to the first monitoring mode as the teacher image.
4. The control device according to claim 3, wherein, The surveillance camera can change the camera range by changing at least one of panning, tilting, and zooming. The switching operation from the first monitoring mode to the second monitoring mode is an operation of at least one of panning, tilting, and zooming when changing the second monitoring mode.
5. The control device according to claim 3, wherein, After the processor switches from the second monitoring mode to the first monitoring mode, it outputs the first camera image as the teacher image according to the provided output instruction.
6. The control device according to claim 3, wherein, The processor performs the following processing: The second camera image acquired in the second monitoring mode before the switch is assigned a judgment result that does not conform to the detection of the object and is output as the teacher image; and The first camera image acquired under the switched first monitoring mode is assigned a judgment result that matches the detection of the object and is output as the teacher image.
7. The control device according to claim 6, wherein, After the processor switches from the second monitoring mode to the first monitoring mode, if it does not perform any operation in the first monitoring mode for a certain period of time, it switches back to the second monitoring mode.
8. The control device according to claim 1, wherein, The processor detects objects reflected in the teacher image and appends the position information of the detected objects within the teacher image to the teacher image.
9. The control device according to claim 8, wherein, The processor sets the detection benchmark for object detection when detecting objects reflected in the teacher's image to be lower than the detection benchmark for object detection when detecting objects reflected in the second camera image.
10. The control device according to claim 8, wherein, The processor performs location information change processing based on the provided instructions to modify the location information.
11. The control device according to claim 1, wherein, The processor determines the position of the object mapped into the teacher image according to the provided instructions, and appends the determined position information of the object in the teacher image to the teacher image.
12. The control device according to any one of claims 1 to 11, wherein, In addition to the teacher image, the processor will also output an enlarged image generated by enlarging the teacher image as the teacher image.
13. The control device according to claim 12, wherein, The scaling process is at least one of the following: inversion, scaling, noise addition, and deep learning-based style transformation.
14. A control method, comprising: The steps for switching between a first monitoring mode and a second monitoring mode are as follows: the first monitoring mode involves capturing a first image using a monitoring camera and changing the camera range according to a provided instruction; the second monitoring mode involves capturing a second image using the monitoring camera and using a machine learning-based learning completion model to detect objects reflected in the second image and changing the camera range according to the detection results. The step of using the first camera image acquired in the first monitoring mode as the teacher image output for the machine learning; and After switching from the first monitoring mode to the second monitoring mode, if the manual operation is performed after a certain period of time since the last manual operation of switching from the second monitoring mode to the first monitoring mode, the second camera image will not be output as the teacher image.
15. A computer-readable storage medium storing a program for causing a computer to perform a process comprising the following steps: The steps include being able to switch between a first monitoring mode and a second monitoring mode. The first monitoring mode involves capturing a first image using a monitoring camera and changing the camera range according to a provided instruction. The second monitoring mode involves capturing a second image using the monitoring camera and using a machine learning-based learning completion model to detect objects reflected in the second image and changing the camera range according to the detection results. The step of using the first camera image acquired in the first monitoring mode as the teacher image output for the machine learning; and After switching from the first monitoring mode to the second monitoring mode, if the manual operation is performed after a certain period of time since the last manual operation of switching from the second monitoring mode to the first monitoring mode, the second camera image will not be output as the teacher image.
Citation Information
Patent Citations
Monitoring controller
JP2004056473A
Method and system for monitoring
JP2006523043A
Method and system for operating a video surveillance system
JP2009516480A
Image processing apparatus, image processing method, program
JP2018117280A
Imaging apparatus and control method of the same
JP2019106694A