surveillance system

The monitoring system efficiently collects basic images for machine learning by using real-time guidance and image capture, enhancing object recognition accuracy and adapting to new environments.

JP7739200B2Active Publication Date: 2025-09-16HITACHI CONSTRUCTION MACHINERY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022028659
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-02-25
Publication Date
2025-09-16
Estimated Expiration
2042-02-25

AI Technical Summary

Technical Problem

Existing surveillance systems face challenges in efficiently collecting basic learning images for improving object recognition accuracy due to memory capacity limitations.

Method used

A monitoring system that includes a camera, processing devices, and a portable terminal for real-time image capture and guidance, allowing efficient collection of basic images for machine learning by displaying guidance images and extracting relevant images based on predefined instructions.

Benefits of technology

Enhances object recognition accuracy by efficiently collecting and updating recognition models with effective basic images, adapting to new environments and improving the efficiency of data generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007739200000001
    Figure 0007739200000001
  • Figure 0007739200000002
    Figure 0007739200000002
  • Figure 0007739200000003
    Figure 0007739200000003
Patent Text Reader

Abstract

To efficiently collect basic images for machine learning of a target object recognition model and improve the accuracy of object recognition.SOLUTION: Provided is a monitoring system comprising: a first memory that stores a recognition model; a first processing device that recognition processes a target object from captured images by a camera, on the basis of the recognition model; a second memory in which training data for generating the recognition model is stored; a second processing device that generates data for updating the recognition model from the basic images for machine learning, on the basis of the training data; and a portable terminal which is capable of displaying the captured images in real time. The first processing device causes the portable terminal to display a guidance image in which guidance data is added to the real-time captured images, extracts the basic images for machine learning of the recognition model from the captured images having been captured by the camera, by operation of the portable terminal conforming to the guidance image, and outputs the same to the second processing device.SELECTED DRAWING: Figure 4
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a monitoring system that recognizes a target object, for example, a monitoring system that monitors the surroundings of a construction machine. [Background technology]

[0002] Patent Document 1 describes an excavator support device that includes an environmental information acquisition unit that acquires environmental information about the shovel's surroundings and a determination unit that determines objects around the shovel based on the acquired environmental information. In this excavator support device, objects around the shovel are determined using a trained model obtained by machine learning. Then, training data is generated from the environmental information acquired by the environmental information acquisition unit, and additional training is performed by an external device using this training data, thereby updating the trained model. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] International Publication No. 2020 / 091002 Summary of the Invention [Problem to be solved by the invention]

[0004] In order to improve the accuracy of object recognition using machine learning, it is effective to collect basic learning images of objects to be judged in the actual environment and update the recognition model (learning model) required for object recognition processing. However, due to various limitations such as memory capacity, the basic images required to improve object recognition accuracy must be collected efficiently.

[0005] An object of the present invention is to provide a surveillance system that can efficiently collect basic images for machine learning of a target object recognition model, thereby improving the accuracy of object recognition. [Means for solving the problem]

[0006] In order to achieve the above object, the present invention provides a camera, a first memory for storing a recognition model which is data for recognizing a target object, a first processing device for processing a captured image input from the camera to recognize the target object based on the recognition model, and a processing unit for generating the recognition model. Data based on basic images for machine learning A second memory that stores training data; ,before In a monitoring system comprising a second processing device that generates update data for the recognition model stored in the first memory based on the training data, and a portable terminal that can receive images captured by the camera and display them in real time, the first processing device transmits the real-time images captured by the camera to the portable terminal in response to an execution command signal from the portable terminal, displays a guidance image on the portable terminal in which default guidance data has been added to the captured image, and extracts basic images for machine learning of the recognition model from the images captured by the camera by operating the portable terminal in accordance with the guidance image in response to a recording start signal from the portable terminal, and outputs the extracted basic images to the second processing device. [Effects of the Invention]

[0007] According to the present invention, basic images for machine learning of a target object recognition model can be efficiently collected, thereby improving the object recognition accuracy. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is a diagram showing an application example of a monitoring system according to an embodiment of the present invention. [Figure 2] 1 is a functional block diagram of a monitoring system according to an embodiment of the present invention; [Figure 3] FIG. 10 is a diagram illustrating an example of an image generated by a monitoring system according to an embodiment of the present invention and displayed on a monitor of a construction machine. [Figure 4] FIG. 10 is a diagram illustrating an example of a first guidance image displayed on a monitor of a mobile terminal in the monitoring system according to an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram illustrating an example of a second guidance image displayed on the monitor of the mobile terminal in the monitoring system according to one embodiment of the present invention. [Figure 6] 1 is a flowchart showing a guidance process for collecting basic images by a monitoring system according to an embodiment of the present invention. [Figure 7] Diagram illustrating the difference in the distance of the target object relative to the camera [Figure 8] The first figure explains how the appearance of a target object in a captured image changes depending on the distance and angle from the camera. [Figure 9] The second figure explains how the appearance of a target object in a captured image changes depending on the distance and angle from the camera. [Figure 10] The third figure explains how the appearance of the target object in the captured image changes depending on the distance and angle from the camera. [Figure 11] The fourth figure explains how the appearance of the target object in the captured image changes depending on the distance and angle from the camera. [Figure 12] The first figure shows how the target object looks different depending on its orientation relative to the camera. [Figure 13] The second figure shows how the target object looks different depending on its orientation relative to the camera. [Figure 14] The third figure shows how the object appears depending on its orientation relative to the camera. [Figure 15] The fourth figure shows how the target object looks different depending on its orientation relative to the camera. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings.

[0010] <Construction machinery> 1 is a diagram showing an example of application of a monitoring system according to one embodiment of the present invention. While the diagram illustrates an example in which a monitoring system 5 according to this embodiment is applied to a hydraulic excavator 1, applications of the monitoring system 5 include various construction machines such as wheel loaders, cranes, and crawler carriers used for various types of work such as civil engineering work, construction work, and demolition work. While the hydraulic excavator 1 is a crawler type, the construction machine to which the monitoring system 5 is applied may be a wheel type or a stationary type.

[0011] The hydraulic excavator 1 includes a vehicle body 4 and a front work implement 10 attached to the vehicle body 4. The vehicle body 4 includes a running body 2 and a rotating body 3 that is rotatably mounted on the running body 2. The running body 2 travels by driving left and right crawlers with a hydraulic traveling motor. The rotating body 3 has a driver's cab 7 in which an operator sits, and is driven to rotate relative to the running body 2 by a hydraulic rotating motor provided on the vehicle body 4. The front work implement 10 is a work machine having an articulated work arm 11 to which an attachment 12 appropriate for the work is attached, and is connected to the front of the rotating body 3. In the example shown in FIG. 1, the attachment 12 is a bucket. The front work implement 10 is driven by multiple hydraulic cylinders 13-15.

[0012] The hydraulic excavator 1 shown in FIG. 1 is equipped with a camera 30 that captures images of the area around the revolving unit 3 in real time. The camera 30 shown in FIG. 1 is attached to the upper rear end of the revolving unit 3, at a position higher than the height of a person standing on the same ground as the hydraulic excavator 1, and is installed in a position that overlooks the area behind the revolving unit 3 diagonally with a viewing angle of approximately 180° left and right. Although not shown in the figure, the camera 30 may be installed on the left, right, or front of the revolving unit 3 in addition to the rear of the revolving unit 3, in which case it also captures images to the left, right, and front of the revolving unit 3. A wide-angle video camera equipped with an imaging element such as a CCD or CMOS and a wide-angle lens, particularly a camera with excellent durability and weather resistance, is preferably used as the camera 30.

[0013] Also provided inside the operator's cab 7 are an operation device for operating each actuator of the hydraulic excavator 1, and a monitor (display device) 190 for displaying images captured by the camera 30 and various information. When the controller 100 determines that a target object (e.g., a human) is present around the hydraulic excavator 1, information about the surrounding target object is displayed on the monitor 190 based on input data from the controller 100. The monitor 190 in this embodiment is a touch panel, and also serves as the input device 180 ( FIG. 2 ) for inputting data to the controller 100 by touch operation. However, the input device 180 may be provided separately from the monitor 190 and may be made up of various switches or the like.

[0014] <Monitoring system> The monitoring system 5 shown in FIG. 1 is a system that monitors a monitoring range using images (image data) captured by the camera 30, and is configured to include a controller 100, a server 53, and a mobile terminal 70. The monitoring range can be set to all or part of the field of view of the camera 30, and for example, it can be set to an area at a certain distance from the vehicle body 4. The controller 100 is an on-board computer mounted on the hydraulic excavator 1, and the server 53 is a computer installed in the management center 52. The mobile terminal 70 is a mobile PC carried by a worker, such as a tablet PC, smartphone, smartwatch, or notebook PC, and is, for example, loaned by the manager of the construction company that uses the hydraulic excavator 1 and carried by the worker.

[0015] The mobile terminal 70 is carried by a worker, and the other components of the monitoring system 5 are installed separately in the hydraulic excavator 1 and the management center 52. The management center 52 is located at a location away from the work site where the hydraulic excavator 1 is operating. The management center 52 is installed, for example, at a facility such as the head office, branch office, or factory of the manufacturer of the hydraulic excavator 1, a rental company for the hydraulic excavator 1, a data center specializing in server operation, or a facility of the owner of the hydraulic excavator 1. The status of the hydraulic excavator 1 is remotely managed by a server 53 installed in the management center 52. There may be cases where a plurality of hydraulic excavators 1 are subject to management by a single server 53.

[0016] The controller 100 and the server 53 are connected to (or include) communication devices 160, 150, respectively, and are configured to be able to perform two-way communication via the communication devices 160, 150 and the wide area network 50. The wide area network 50 is a mobile phone communication network (mobile communication network) provided by a mobile phone carrier or the like, the Internet, or the like. When the controller 100 of the hydraulic excavator 1 and a wireless base station 51 are connected via a mobile phone communication network as shown in FIG. 1, for example, data transmitted from the hydraulic excavator 1 to the wireless base station 51 is transmitted to the server 53 via the Internet. The communication devices 160, 150 are wireless communication devices capable of wireless communication with the wireless base station 51 connected to the wide area network 50, and have a communication interface including a communication antenna having a sensitivity band within a predetermined frequency band.

[0017] The mobile terminal 70 also has a built-in communication device 170 (FIG. 2) and is configured to be able to perform two-way communication with the controller 100 via the communication devices 170, 160 and a local network. For example, a wireless base station (access point) is installed at the work site, and data is exchanged between the controller 100 and the mobile terminal 70 via the wireless base station.

[0018] The communication device 160 connected to the controller 100 can be configured to directly or indirectly exchange data with the server 53 and the mobile terminal 70 using a known short-distance communication method. The communication method between the controller 100 and the server 53 and the communication method between the controller 100 and the mobile terminal 70 may be different.

[0019] <controller> Fig. 2 is a functional block diagram of the monitoring system. As shown in Fig. 2, a camera 30, a monitor 190, a communication device (first communication device) 160, and an input device 180 are connected to a controller 100. The controller 100 is configured to include a memory 110 (first memory) such as a ROM, RAM, HDD, SSD, etc., as well as a CPU 120 (first processing device), an input / output interface, and other peripheral circuits.

[0020] The ROM of the controller 100 (as well as the server 53 and the mobile terminal 70) is a non-volatile memory such as an EEPROM, and stores programs capable of executing various calculations. The RAM is a volatile memory, and serves as a work memory for directly inputting and outputting data to and from the CPU. The RAM temporarily stores necessary data while the CPU is executing a program. The non-volatile memory stores a recognition model 131 and a base image 132. The recognition model 131 and the base image 132 may be stored in a single non-volatile memory, or may be stored in different non-volatile memories.

[0021] The recognition model 131 is data for recognizing a target object (a human or other obstacle). The recognition model 131 is, for example, a YOLO (You Only Look Once) model, which is a type of convolutional neural network. The YOLO model is a recognition model created by deep learning based on machine learning using a neural network, and is stored in advance in the memory 110. The recognition model 131 is also generated for each target object based on additional machine learning in the server 53, for example, and is stored or updated in the memory 110.

[0022] The base image 132 is image data extracted from an image captured by the camera 30, and is data that serves as the basis for machine learning of the target object recognition model 131.

[0023] Signals from various sensors and devices are input to the input / output interface of the controller 100 (as well as the server 53 and the mobile terminal 70). The input / output interface converts the input signals into a format that can be calculated by the CPU 120 as needed. The CPU 120 is a processing device that loads a program stored in ROM into RAM and executes it, and performs predetermined calculation processing based on data input from the input / output interface, ROM, and RAM in accordance with the program. The input / output interface converts the calculation results of the CPU 120 into a data format for output and outputs them to the corresponding device.

[0024] The CPU 120 executes each process of object recognition 101, display image generation 102, recording control 103, communication control 104, and guidance image generation 106 in accordance with the program stored in the memory 110.

[0025] -Object recognition- Object recognition 101 is a process of acquiring a captured image (image data) input from camera 30 and recognizing the presence of a target object within a monitoring range from the acquired captured image based on recognition model 131. Specifically, in object recognition 101, CPU 120 determines whether the reliability calculated in the recognition process targeting the captured image is equal to or greater than a first reliability (a set value). If the reliability is equal to or greater than the first reliability, it is determined that a target object exists within the monitoring range. If the reliability is less than the first reliability, it is not determined that an object exists within the monitoring range.

[0026] The reliability calculated in the recognition process for a captured image is an evaluation value calculated in the object recognition 101 process and corresponds to the degree of match with the parameter set of the target object (e.g., a human) preset in the recognition model 131. If there is a perfect match with the recognition model 131, the reliability is 100%. The parameter set of the target object is the image features (edge, color, position, etc.). For example, in the case of YOLOv3, the parameter is the setting value obtained by machine learning for each of the 72 convolutional layers (feature filters). The first reliability is a preset threshold for determining the presence or absence of a target object within the monitoring range, and can be set to, for example, a value of approximately 90%. The object recognition 101 process generates data on the class, position coordinates, and size of the target object recognized as being present within the monitoring range. The class is data on the type (classification) of the target object. The position coordinates are data on the horizontal and vertical coordinates of the target object within the captured image. The size is the size of the target object on the captured image (data on the length in the horizontal and vertical directions of a rectangular frame surrounding the target object on the captured image).

[0027] -Display image generation- Display image generation 102 is a process for generating a display image by combining a frame surrounding a target object with an image captured in real time by camera 30, based on data (position coordinates and size) of the target object determined to exist within the monitoring range by object recognition 101. The data of the display image generated by the process of display image generation 102 is output from controller 100 to monitor 190 and displayed on the screen of monitor 190.

[0028] FIG. 3 shows an example of an image generated by the processing of display image generation 102 and displayed on the screen of monitor 190. In the display image shown in the figure, a frame 193 surrounding a target object 192 is superimposed on an image 191 captured in real time by camera 30. FIG. 3 shows an example of a display image in which a worker (human) standing near camera 30 (i.e., near the rear of hydraulic excavator 1) is captured as target object 192. In this display image, text data representing the class of the target object (for example, "human" in the example shown) can also be displayed alongside frame 193. By displaying target object 192 surrounded by frame 193 on the display image in this way, the operator of the hydraulic excavator 1 can be visually informed of the position, etc., of target objects present around the hydraulic excavator 1.

[0029] - Guidance image generation - Guidance image generation 106 is a process for generating a guidance image (FIGS. 4 and 5) in response to a command signal for executing a guidance mode from mobile terminal 70. The guidance image generated by the process of guidance image generation 106 is transmitted to mobile terminal 70 via communication device 160 after being processed by communication control 104.

[0030] A guidance image is an image obtained by synthesizing predetermined guidance data with a real-time image captured by the camera 30. Here, the predetermined guidance data is data related to guiding the procedure for acquiring a captured image that serves as the basis for machine learning. The guidance data includes, for example, a demarcation line L, a designated frame F, and a message field 198 ( FIG. 4 ). The demarcation line L is a line that divides the captured image 191 into multiple areas A. The designated frame F is a frame that displays a designated area A1, which is one of the areas A designated to display a target object 192 on the captured image 191. FIG. 4 illustrates a state in which a worker, who is the target object 192, is imaged within the frame F designated as the designated area A1. The message field 198 is a display field for text that provides guidance on the action required of the worker. As the guidance data to be displayed in the message field 198, for example, a message is generated that prompts the target object 192 to be displayed in a predetermined posture in the designated area A1.

[0031] -Recording control- The recording control 103 is a process that extracts basic images for machine learning of the target object recognition model 131 from images captured by the camera 30 in response to a recording start signal from the mobile terminal 70 and records the extracted images in the memory 110. The captured images extracted as basic images are captured by the worker operating the mobile terminal 70 according to the guidance images. As described above, the guidance images are images captured in real time by the camera 30, and the basic images are captured using a live view method using the camera 30 and the mobile terminal 70. The basic images 132 are extracted by extracting them from multiple still images that make up a video. For example, when the camera 30 captures a video consisting of 30 still images per second, many of the 30 captured images acquired per second are generally similar. Therefore, in the process of the recording control 103, the basic images 132 are extracted and recorded in the memory 110 by thinning out the images captured by the camera 30 (e.g., extracting them at 0.2-second intervals).

[0032] The basic image data recorded in the memory 110 includes, in addition to image data, additional data such as the identification number of the hydraulic excavator 1 on which the camera 30 is mounted, identification information for the camera 30, and the date and time the basic image was captured. The identification information for the camera 30 can be, for example, the coordinates of the camera 30 in the local coordinate system of the hydraulic excavator 1 (position data indicating which part of the hydraulic excavator 1 the camera 30 is installed on). By including this additional data, it is possible to determine when and where the basic image was captured, and to determine in what situations the update data for the recognition model 131 generated based on the basic image 132 is effective. In the object recognition 101 process, in order to display the image shown in FIG. 3, the class, position coordinates, and size are generated as data for the target object recognized to be present within the monitoring range. The class is data on the type (classification) of the target object. The position coordinates are data on the horizontal and vertical coordinates of the target object in the captured image. The size is data on the size of the target object in the captured image (data on the horizontal and vertical lengths of a rectangular frame surrounding the target object in the captured image).

[0033] -Communication Control- Communication control 104 is a general term for processing that controls input and output of data to and from controller 100 via communication device 160. For example, when portable terminal 70 is operated and a command signal to execute guidance mode is input to controller 100, CPU 120 generates a guidance image, executes processing of communication control 104, and starts transmitting the guidance image to portable terminal 70. When input device 180 is operated and a command signal to cancel guidance mode is input to controller 100, processing of communication control 104 is executed, and transmission of the guidance image to portable terminal 70 is terminated. Communication control 104 also includes processing that checks the connection status between communication device 160 and wireless base station 51 (FIG. 1), and determines whether communication with server 53 is possible via communication device 160.

[0034] As part of the processing of the communication control 104, if the CPU 120 is able to communicate with the server 53, it outputs (transmits and uploads) the data of the basic images 132 recorded in the memory 110 to the server 53 via the communication device 160. When the upload of the data of the basic images 132 is successfully completed, a reception completion notice is sent from the server 53 via the communication device 160. In the processing of the communication control 104, when this reception completion notice is received, data indicating that the data of the basic images 132 recorded in the memory 110 has been sent is added for the successfully uploaded basic images 132. The data of the basic images 132 that have already been sent is not resent to the server 53 but is periodically erased or overwritten.

[0035] Furthermore, in the processing of the communication control 104, the CPU 120 communicates with the server 53 to acquire information (version data) about the recognition model 131 stored in the memory 54 of the server 53. At this time, the controller 100 determines whether a recognition model 131 with a newer version than the recognition model 131 stored in the memory 110 by the CPU 120 is recorded in the memory 54 of the server 53. If a new recognition model 131 is available in the server 53, the controller 100 transmits a signal to the server 53 requesting transmission of update data for the recognition model 131. When the server 53 receives the request signal and transmits the update data for the recognition model 131, the controller 100 receives the data and updates (upgrades) the recognition model 131 recorded in the memory 110. In other words, the parameter group constituting the recognition model 131 stored in the memory 110 is updated.

[0036] <server> A monitor 55 ( FIG. 1 ), a communication device (second communication device) 150, and an input device 56 are connected to the server 53. The monitor 55 is a liquid crystal display device or the like, and the input device 56 is a keyboard, a mouse, or the like. When the server 53 receives operation data, data of basic images 132, and the like from the hydraulic excavator 1, the server 53 records the data in the memory 54 and displays the data stored in the memory 54 on the monitor 55 in response to a predetermined operation input from the input device 56. The manager can learn the status of the hydraulic excavator 1, etc., by viewing predetermined information on the monitor 55. Images (including the basic images 132) captured by the camera 30 of the hydraulic excavator 1 can also be displayed on the monitor 55 in a similar manner.

[0037] The server 53 is also a computer like the controller 100, and is configured to include a memory 54 (second memory) such as ROM, RAM, HDD, SSD, etc., as well as a CPU 57 (second processing device), an input / output interface, and other peripheral circuits.

[0038] The memory 54 stores a recognition model 131 and training data for generating the recognition model 131. For the training data, image data that has been visually confirmed by an administrator to include a target object (e.g., a human) can be selected from the basic images 132 sent from the controller 100 of the hydraulic excavator 1. This training data includes not only image data but also additional data such as the position, size, and class of the target object. For example, the horizontal and vertical coordinates of the target object in the image can be used as the position data of the target object. For example, the horizontal and vertical lengths of the frame 193 ( FIG. 3 ) surrounding the target object in the image can be used as the size data of the target object. As described above, the class of the target object is classification information (e.g., a human) that indicates the type of the target object. Such additional data can be generated by, for example, having the administrator input data using the input device 56 while viewing the image on the monitor 55, and then recording the data in the memory 54 in association with the basic image 132 selected as training data (training data generation operation).

[0039] The CPU 57 executes various processes in accordance with the programs stored in the memory 54. For example, when the CPU 57 receives data of the basic image 132 from the controller 100 via the communication device 150, it records the data of the basic image 132 received from the controller 100 in the memory 54. After finishing recording the basic image 132, the CPU 57 transmits a reception completion notice indicating that reception of the basic image 132 has been successfully completed to the controller 100 of the hydraulic excavator 1 via the communication device 150. The CPU 57 also displays the basic image 132 on the monitor 55 in response to a predetermined operation of the input device 56.

[0040] Furthermore, the CPU 57 executes processing to generate update data for the recognition model 131 stored in the memory 110 of the controller 100 from the basic image 132 transmitted from the controller 100. This processing to generate update data is executed for the basic image 132 by machine learning (for example, a neural network method using YOLOv3) based on the training data stored in the memory 54. The update data is an update program for updating the recognition model 131 recorded in the memory 110 of the hydraulic excavator 1 to a new version (applying parameters calculated by machine learning to the recognition model 131 stored in the memory 110). Note that this update data is not limited to the update program, and may be the new version of the recognition model 131 itself. The update data is recorded in the memory 54, and is transmitted to the controller 100 of the hydraulic excavator 1 in response to a request signal from the controller 100, as described above.

[0041] <Mobile device> The portable terminal 70 is a mobile computer that can receive and display in real time images 191 captured by the camera 30 mounted on the hydraulic excavator 1. The portable terminal 70 includes memories (not shown) such as ROM, RAM, HDD, and SSD, a CPU 71 (third processing device), a monitor 72, a communication device (third communication device) 170, and an input device 73. The monitor 72 is a liquid crystal display device or the like, and the input device 73 is typically a touch panel or touch pad, although a keyboard, mouse, etc. can also be used.

[0042] This mobile terminal 70 receives images captured by the camera 30 (including the above-mentioned guidance image) from the controller 100 via the communication device 170, and can display a live view on the monitor 72 through the operation of the CPU 71 in response to the operator's operation of the input device 73. This allows the operator to visually confirm the situation around the hydraulic excavator 1 in real time through the monitor 72 of the mobile terminal 70, and can also capture the above-mentioned basic image 132 by operating the mobile terminal 70 in accordance with the guidance image.

[0043] -Guidance image- 4 and 5 are diagrams showing examples of guidance images displayed on the monitor of a mobile terminal, with FIG. 4 being a screen before recording of a basic image starts, and FIG. 5 being a screen during recording of a basic image.

[0044] The guidance image 194a shown in Fig. 4 is generated by the execution of the process of guidance image generation 106 in the controller 100 in response to a guidance mode execution command signal input from the mobile terminal 70 by a predetermined operation. Then, under the control of the communication control 104 executed in parallel with the process of guidance image generation 106 in the controller 100, the guidance image 194a is transmitted from the controller 100 to the mobile terminal 70 via the communication device 160. On the mobile terminal 70 side, when a predetermined operation (for example, selection of the type or posture of the target object) is performed on the input device 73, the guidance image 194a received by the communication device 170 is displayed in real time on the monitor 72 (Fig. 4). Fig. 4 illustrates an example of the guidance image 194a capturing an image of a person in a crouching posture as the target object 192.

[0045] The guidance image 194a displayed on the mobile terminal 70 displays a camera ID 195, a recording start button 196, a stop button 197, and a message field 198 in addition to an image 191 (live view video) captured by the camera 30.

[0046] The photographed image 191 displayed on the guidance image 194a is divided into a plurality of areas A by a plurality of demarcation lines L, and a designated frame F representing a designated area A1 designated by the controller 100 from the plurality of areas A is displayed superimposed on the photographed image 191. The guidance image 194a shown in FIG. 4 illustrates an example in which the photographed image 191 is divided into 16 parts (4 parts vertically and 4 parts horizontally).

[0047] The camera ID 195 is identification information of the camera 30 that captured the captured image 191, and in the example of Fig. 4, it is a combination of the machine ID (SH01) of the hydraulic excavator 1 on which the camera 30 is mounted and the installation position (Rear) on the hydraulic excavator 1. The recording start button 196 is a button for instructing the controller 100 to send a recording start signal that instructs the controller 100 to start recording the basic image 132. The stop button 197 is a button for instructing the CPU 71 to stop displaying the guidance image 194a.

[0048] The message field 198 is a display area that displays text of the action required of the worker who is operating the mobile terminal 70 while looking at the guidance image 194a. In the example of the guidance image 194a in Fig. 4 before recording starts, the message field 198 displays guidance such as "Please stand within the designated area and start recording."

[0049] 4, to capture a basic image 132 in which a human being is captured as a target object, for example, the worker first stands in the monitoring range of the camera 30 and appears in the captured image 191 displayed on the mobile terminal 70. Then, in accordance with the message in the message field 198, the worker stands in the designated area A1 (for example, standing so that his feet are within the designated area A1) designated by the controller 100 (or designated by the worker himself by operating the input device 73). In other words, the worker himself becomes the human being as the target object 192 captured in the basic image 132. After standing in the designated area A1 in accordance with the message in the message field 198, the worker presses the recording start button 196 to remotely instruct the controller 100 of the hydraulic excavator 1 to start capturing (recording) the basic image 132.

[0050] In response to a recording start signal input from the mobile terminal 70 by operating the recording start button 196, the controller 100 starts the processing of the recording control 103 described above. At the same time, the controller 100 executes the processing of the guidance image generation 106 and the communication control 104, and generates the guidance image 194b shown in Fig. 5. The generated guidance image 194b is transmitted via the communication device 160 and displayed in real time on the monitor 72 of the mobile terminal 70.

[0051] Guidance image 194b in Fig. 5 differs from guidance image 194a in Fig. 4 in that it does not have a stop button 197, that start recording button 196 has been replaced with a stop recording button 199, and that "Recording in progress" is displayed on captured image 191. Also, the message displayed in message field 198 has been switched to guidance regarding posture, saying "Please crouch within the designated area for 3 seconds." While this shows an example of guidance regarding the worker's posture, guidance regarding the orientation of target object 192 relative to camera 30 may also be provided in addition to posture.

[0052] 5, a human being (the worker himself / herself in the example of the same figure) as the target object 192 crouches down in the designated area A1 and maintains this posture for three seconds, and a basic image 132 showing the human being crouching down in the designated area A1 is captured. After that, when the recording end button 199 is operated, a signal instructing the end of recording is transmitted from the mobile terminal 70 to the controller 100. Upon receiving the recording end signal, the controller 100 stops the processing of the recording control 103 and switches the guidance image to be transmitted to the mobile terminal 70 to guidance image 194a (FIG. 4).

[0053] To finish collecting the basic images 132, when the cancel button 197 on the guidance image 194a is pressed, the guidance image 194a closes on the monitor 72 of the mobile terminal 70, and a command signal to cancel the guidance mode is sent from the mobile terminal 70 to the controller 100. The controller 100 that has received the command signal to cancel the guidance mode cancels the guidance mode and stops the processing of the guidance image generation 106. The same applies when a command signal to cancel the guidance mode is input from the input device 180 of the cab 7.

[0054] <Operation> 6 is a flowchart showing the guidance processing procedure by the controller 100 when collecting basic images. When a command signal to execute the guidance mode is input, the controller 100 starts the processing shown in the figure, and the processing is repeated until a command signal to cancel the guidance mode is input.

[0055] In step S01, the controller 100 acquires data of the photographed image 191 from the camera 30, and the procedure proceeds to step S02.

[0056] When the procedure proceeds to step S02, the controller 100 initializes the setting of the designated area A1 in which the target object 192 on the photographed image 191 should be captured, and then the procedure proceeds to step S03.

[0057] When the procedure proceeds to step S03, the controller 100 sets a designated area A1 on the captured image 191 in accordance with a preset rule. The designated area A1 that is initially set when the guidance mode is activated is a preset initial setting area. If an opportunity to perform step S03 arises again while the guidance mode is ongoing, the next area in accordance with the rule or an area designated by the worker is displayed as the designated area A1.

[0058] In the following step S04, the controller 100 generates a guidance image 194a (FIG. 4) including the photographed image 191 with the designated frame F combined therewith, and the like, and then proceeds to step S05.

[0059] When the procedure proceeds to step S05, the controller 100 transmits the guidance image 194a to the mobile terminal 70 via the communication device 160, and then proceeds to step S06.

[0060] When the procedure proceeds to step S06, the controller 100 determines whether the recording start button 196 has been pressed and a recording start signal has been input from the mobile terminal 70. If a recording start signal has been input, the controller 100 proceeds to step S07, and if no recording start signal has been input, the controller 100 proceeds to step S11.

[0061] When a recording start signal is input from the mobile terminal 70 and the procedure proceeds to step S07, the controller 100 starts recording the captured image 191 in the memory 110 at predetermined time intervals as the basic image 132, generates a guidance image 194b (Figure 5), and proceeds to step S08.

[0062] In the following step S08, the controller 100 transmits the guidance image 194b via the communication device 160, and the procedure proceeds to step S09.

[0063] Then, in step S09, when the controller 100 receives a recording end signal transmitted from the mobile terminal 70 in response to the operation of the recording end button 199, the controller 100 proceeds to step S10 to interrupt the recording of the basic image 132, and returns the procedure to step S03.

[0064] Furthermore, if a recording start signal is not input in step S06 described above, the controller 100 proceeds to step S11, where it determines whether or not a signal specifying the designated area A1 has been input from the mobile terminal 70 by the worker operating the input device 73. If a signal specifying the designated area A1 has been input, the controller 100 returns to step S03, and if not, proceeds to step S12.

[0065] When the procedure proceeds to step S12, the controller 100 determines whether a command signal to cancel the guidance mode has been input from the mobile terminal 70 (or the input device 180 of the driver's cab 7). If a command signal to cancel the guidance mode has not been input, the controller 100 returns the procedure to step S06 and repeats the determinations of steps S06, S11, and S12 until an instruction to start recording, specify the designated area A1, or cancel the guidance mode is given. If a command signal to cancel the guidance mode has been input, the controller 100 cancels the guidance mode and ends the flow of FIG. 6.

[0066] -effect- (1) In this embodiment, the controller 100 recognizes target objects 192 around the hydraulic excavator 1 from captured images 191 input from the camera 30, based on the recognition model 131 stored in the memory 110. Furthermore, the server 53 generates update data for the recognition model 131 from basic images 132 for machine learning captured by the camera 30, based on the training data stored in the memory 54, and updates the recognition model 131.

[0067] Then, the controller 100 supports the collection of basic images 132 that serve as the basis for data for updating the recognition model 131 by displaying, on the mobile terminal 70, guidance images 194a and 194b including the image 191 captured by the camera 30 in a live view as described above. Specifically, the controller 100 transmits the guidance images 194a and 194b to the mobile terminal 70 in response to a command signal to execute the guidance mode, and extracts the basic image 132 from the image 191 captured by the camera 30 in response to an operation of the mobile terminal 70, and transmits the basic image 132 to the server 53. The guidance images 194a and 194b are obtained by combining guidance data with the image 191 captured in real time by the camera 30, and the user can confirm how the position and posture of the target object 192 are displayed in the live view video of the camera 30 in accordance with the guidance. This makes it possible to easily capture the basic image 132 in accordance with the guidance.

[0068] Therefore, according to this embodiment, it is possible to efficiently collect basic images 132 for machine learning of the recognition model 131 of a target object 192 such as a human being, and by enriching the teacher data, it is possible to improve the recognition accuracy of the target object 192. Since it is possible to efficiently collect basic images 132 that are effective at the operating site of the hydraulic excavator 1, it is possible to adapt the recognition model 131 to a new environment when the hydraulic excavator 1 is used at a new site, for example, when a new monitoring system 5 is introduced or when the operating site of the hydraulic excavator 1 is changed. Since it is possible to efficiently collect effective and necessary basic images 132 using default guidance images, it is also possible to improve the efficiency of the process of generating update data for the recognition model 131 by the server 53 and the work of creating teacher data by the administrator.

[0069] (2) According to the monitoring system 5 of this embodiment, for example, a photographer other than the worker who appears as the target object 192 in the basic image 132 can operate the mobile terminal 70, and the photographer can give instructions to the worker regarding the position and posture while looking at the mobile terminal 70, thereby capturing the basic image 132. When a person other than the worker who is the subject of the image gives instructions while looking at the live view image from the camera 30, it is also possible to configure the live view image from the camera 30 to be output to the monitor 190 in the cab 7. However, in these cases, two people are required: the worker who is the subject of the image and the photographer, and communication between the photographer giving instructions and the worker receiving the instructions can be time-consuming.

[0070] In contrast, as described above, the basic images 132 can be captured while checking the live view video of the camera 30 of the hydraulic excavator 1 on the mobile terminal 70. Therefore, by having the worker operating the mobile terminal 70 himself become the subject of the image, the number of personnel required to collect the basic images 132 can be reduced, and the basic images 132 can be collected efficiently without the hassle of communication between personnel.

[0071] (3) As described above, by displaying the designated area A1 in the designated frame F on the real-time captured image 191 and conveying instructions in the message field 198, the target object 192 can be easily displayed at the desired position in the basic image 132 showing the monitoring range of the camera 30.

[0072] (4) In particular, by displaying specific messages in the guidance images 194a and 194b encouraging the target object 192 to be captured in a predetermined posture in the designated area A1, the worker operating the mobile terminal 70 can take action to capture the basic image 132 without hesitation.

[0073] As mentioned above, the camera 30 is mounted on top of the revolving body 3 of the hydraulic excavator 1, and as shown in Figure 7, the appearance of the target object 192 in the captured image 191 changes depending on the distance and angle from the hydraulic excavator 1 (camera 30).

[0074] For example, Fig. 8 illustrates how a target object 192 (human) appears in the center of a photographed image 191 in the left-right direction and at a short distance from the hydraulic excavator 1. The target object 192 (human) shown in Fig. 8 is in an upright position, but because it is photographed from above by the wide-angle camera 30, a top view of the target object 192 appears large in the photographed image 191 in an area below the center of the screen.

[0075] 9 shows an example of how a target object 192 (human) located in the center of the photographed image 191 in the left-right direction and at a location slightly away from the hydraulic excavator 1 appears. As shown in FIG. 9, even for a target object 192 (human) with the same posture, the depression angle decreases as the human moves away from the hydraulic excavator 1. As a result, the photographed image 191 shows the object as seen from the side (front), and the size of the target object 192 in the photographed image 191 also becomes smaller.

[0076] 10 illustrates an example of how a target object 192 (human) appears on the right side of the photographed image 191 in the left-right direction and at a short distance from the hydraulic excavator 1. FIG. 11 illustrates an example of how a target object 192 (human) appears on the right side of the photographed image 191 in the left-right direction and at a location slightly distant from the hydraulic excavator 1. Because the camera 30 has a wide angle, even if the target object 192 (human) has the same posture, as shown in FIGS. 10 and 11, the closer it is to the left or right end of the photographed image 191, the more the target object 192 appears tilted to the left or right compared to FIGS. 8 and 9.

[0077] 12 to 15 are diagrams showing differences in how the target object 192 appears depending on the posture relative to the camera 30. The target object 192 shown in FIGS. 12 to 15 is assumed to appear around the position of the target object 192 illustrated in FIG. 9. Even when the target object 192 (human) is photographed facing forward at the same position, the appearance of the target object 192 in the photographed image 191 differs significantly depending on whether the target object 192 (human) is standing upright (FIG. 12) or leaning forward (FIG. 13). Furthermore, even when the target object 192 is in the same upright or leaning forward posture, the appearance of the target object 192 in the photographed image 191 changes as the orientation of the target object 192 relative to the camera 30 changes, as shown in FIGS. 14 and 15.

[0078] As described above, in this embodiment, the message field 198 is used to provide guidance to the worker regarding the position, posture, etc. of the target object 192. This makes it possible to comprehensively collect variations of the basic images 132 with different postures and orientations of the target object 192 for each area A of the captured image 191, as described with reference to FIGS. 8 to 15. By comprehensively collecting variations in this way and performing machine learning, it becomes possible to recognize the target object 192 with high accuracy even if the position, posture, orientation, etc. of the target object 192 relative to the camera 30 changes.

[0079] 12 to 15 explain the difference in appearance depending on the posture of the target object 192, but by collecting and learning basic images 132 of the target object 192 with different equipment (jacket, helmet, safety belt, etc.), it becomes possible to recognize the target object 192 with different equipment with high accuracy.

[0080] (5) A command to execute the guidance mode can be given from the mobile terminal 70, but in this embodiment, the command is executed from the input device 180 provided in the cab 7 of the hydraulic excavator 1. In this case, a situation can be naturally created in which the operator sitting in the cab is aware that the basic images 132 are being collected (that there are workers nearby). Note that the controller 100 may be provided with an interlock function that prohibits the operation of the various hydraulic actuators of the hydraulic excavator 1 while the guidance mode is being executed.

[0081] (6) In this embodiment, the camera 30, memory 110, and CPU 120 are mounted on the hydraulic excavator 1, and the memory 54 and CPU 57 are mounted on the server 53. In this way, by having the function of generating the recognition model 131 and its update data carried out by a computer external to the hydraulic excavator 1, it is possible to centrally generate and update the recognition models 131 of multiple construction machines on a single server 53.

[0082] However, for example, if the generation of the recognition model 131 or update data executed by the server 53 is intended for only a single hydraulic excavator 1, the functions of the server 53 may be provided in the controller 100 of the hydraulic excavator 1, or an on-board computer may be added to the hydraulic excavator 1. In this case, the memories 110 and 54 may be configured as a single memory or may be configured as multiple memories. Similarly, the CPUs 120 and 57 may be configured as a single CPU or may be configured as multiple CPUs.

[0083] <Modification> Although one embodiment of the present invention has been described above, the present invention is not limited to the above embodiment and can be modified in various ways without departing from the spirit of the present invention. Some modified examples are given below.

[0084] -Variation 1- In the above embodiment, an example has been described in which the camera 30 is installed on the hydraulic excavator 1. However, in order to collect basic images 132 that form the basis of update data for the recognition model 131, a configuration may be adopted in which the camera is set up to simulate an on-board environment of the hydraulic excavator 1. As an example of this, FIG. 1 also shows an example in which a camera 33 with the same angle of view as the camera 30 is installed on a support 34 so that the height above ground and the optical axis angle (depression angle) are set to simulate installation on the hydraulic excavator 1. In the example shown in the same figure, an image 191 captured by the camera 33 installed on the support 34 is input to a monitoring simulator 240. The monitoring simulator 240 is a computer having the same configuration and functions as the controller 100, and therefore a detailed description thereof will be omitted.

[0085] In this modification, similar to the controller 100 of the above embodiment, the monitoring simulator 240 acquires a captured image 191 by the camera 33, generates guidance images 194a and 194b from the captured image 191, and transmits them to the mobile terminal 70. In addition, the monitoring simulator 240 receives a recording start signal from the mobile terminal 70 and records the basic image 132.

[0086] In this way, even if the camera 33 simulating the vehicle-mounted state of the camera 30 of the hydraulic excavator 1 is used, the basic images 132 can be collected in the same manner as in the above embodiment.

[0087] -Variation 2- In the above embodiment, an example has been described in which the basic image 132 is transmitted from the controller 100 of the hydraulic excavator 1 to the server 53 installed in the management center 52, but the transmission destination of the basic image 132 is not necessarily limited to the server 53. As illustrated in FIG. 1 , the system configuration may be such that the controller 100 of the hydraulic excavator 1 transmits the basic image 132 to a data server 58 installed in a location separate from the hydraulic excavator 1 and the management center 52. The basic image 132 transmitted from the controller 100 is stored in the memory 59 of the data server 58. In this example, the server 53 receives the basic image 132 from the data server 58 via the communication device 150, and generates update data for the recognition model 131 based on the received basic image 132.

[0088] In this example, the basic image 132 is stored in the data server 58, so the capacity of the memory 54 of the server 53 can be reduced.

[0089] -Variation 3- In the above embodiment, an example has been described in which the controller 100 of the hydraulic excavator 1 generates the guidance images 194a, 194b and transmits them to the mobile terminal 70, but the guidance images 194a, 194b do not necessarily have to be generated by the controller 100. For example, a configuration may be adopted in which the controller 100 simply transmits a real-time captured image 191 from the camera 30, and the mobile terminal 70 overlays guidance data on the captured image 191 to generate and display the guidance images 194a, 194b. The captured image 191 received by the mobile terminal 70 may be recorded in the memory of the mobile terminal 70.

[0090] According to this example, the processing for generating guidance images 194a and 194b in the controller 100 is shared by the mobile terminal 70, thereby reducing the processing load on the controller 100 and effectively suppressing delays in transmitting images to the mobile terminal 70.

[0091] -Variation 4- In the above embodiment, an example has been described in which a command to execute the guidance mode is given by the input device 180 of the operator's cab 7 of the hydraulic excavator 1, but as mentioned above, a configuration in which a command to execute the guidance mode is given from, for example, the mobile terminal 70 may also be used. Furthermore, the guidance mode can be activated only when the front work unit 10, the traveling unit 2, and the rotating unit 3 of the hydraulic excavator 1 are not operating. Examples of a state in which the front work unit 10, the traveling unit 2, and the rotating unit 3 are not operating include a state in which the key is on and power is supplied to the controller 100, but the front work unit 10, the traveling unit 2, and the rotating unit 3 cannot be operated by, for example, a so-called gate lock valve, or when the engine is idling.

[0092] In this example, the guidance mode is executed in a state where the operation of the hydraulic excavator 1 is prohibited, and the basic images 132 can be collected.

[0093] -Variation 5- In the above embodiment, a human being is given as a specific example of the target object 192, but the present invention can also be applied to cases where the target object 192 is a work vehicle, a sign, or the like, in addition to an obstacle such as a rock. For example, when a sign is the target object 192, the sign is selected as the target object 192, a guidance image 194a is generated, and a message such as "Please place the sign within the designated area and start recording" is displayed in the message field 198 of Fig. 4. This makes it possible to capture a basic image 132 in which the sign is reflected in the designated area A1.

[0094] -Variation 6- In the above embodiment, the monitoring system 5 is applied to a crawler-type hydraulic excavator 1 as an example, but as mentioned above, the monitoring system 5 can be applied to various construction machines such as wheel-type hydraulic excavators, wheel loaders, dump trucks, and cranes.

[0095] -Variation 7- In the above embodiment, an example has been described in which the YOLO model is applied to the recognition model 131, but the recognition model 131 is not limited to the YOLO model. A CNN (Convolutional Neural Network) model for object recognition, such as an SSD (Single Shot Multibox Detector), can also be adopted as the recognition model 131. [Explanation of symbols]

[0096] 1...hydraulic excavator (construction machinery), 5...monitoring system, 30...camera, 33...camera, 34...support, 53...server, 54...memory (second memory), 57...CPU (second processing device), 70...mobile terminal, 110...memory (first memory), 120...CPU (first processing device), 131...recognition model, 132...basic image, 132...basic image, 191...captured image, 192...target object (worker), 194a, 194b...guidance image, 198...message field (guidance data), A...area, A1...designated area (guidance data), F...designated frame (guidance data), L...dividing line (guidance data)

Claims

1. A camera and a first memory that stores a recognition model, which is data for recognizing a target object; a first processing device that recognizes the target object from the captured image input from the camera based on the recognition model; a second memory storing training data based on basic images for machine learning, the training data being data for generating the recognition model; a second processing device that generates update data for the recognition model stored in the first memory based on the training data; a mobile terminal capable of receiving images captured by the camera and displaying them in real time; In a monitoring system comprising: The first processing device includes: In response to an execution command signal from the mobile terminal, a real-time image captured by the camera is transmitted to the mobile terminal, and a guidance image in which predetermined guidance data is added to the captured image is displayed on the mobile terminal; In response to a recording start signal from the mobile terminal, extracting a base image for machine learning of the recognition model from a captured image taken by the camera by operating the mobile terminal according to the guidance image; The extracted basic image is output to the second processing device. A monitoring system characterized by:

2. 2. The monitoring system according to claim 1, A monitoring system characterized in that the guidance data includes dividing lines that divide the image captured by the camera into multiple areas, a designation frame that displays a designated area on the captured image where the target object should be captured, and a message field that provides guidance on the actions required of the worker operating the mobile terminal.

3. 3. The monitoring system according to claim 2, A surveillance system characterized in that a message prompting the target object to be captured in a predetermined posture in the designated area is displayed in the message field.

4. 2. The monitoring system according to claim 1, the camera, the first memory, and the first processing device are mounted on a construction machine; The second memory and the second processing device are mounted on a server that exchanges data with the construction machine. A monitoring system characterized by:

5. 2. The monitoring system according to claim 1, The camera is mounted on a pole at a height above ground that simulates its installation on a construction machine. A monitoring system characterized by:

Citation Information

Patent Citations

  • Individual recognition device and passage control device

    JP2003141541A

  • Information processing apparatus, shop system, and program

    JP2015035094A

  • Server for learning, image collection assisting system for insufficient learning, and image estimation program for insufficient learning

    JP2019192082A

  • Construction site image acquisition system, construction site image acquisition device, and construction site image acquisition program

    JP2021174284A

  • Person recognition apparatus

    US20070031010A1