Behavior detection system and behavior detection method
The action detection system efficiently detects unsafe behaviors by combining multiple types of specific data through a multimodal model, addressing inefficiencies in conventional systems that rely on single-image data analysis.
Patent Information
- Application Number
- JP2025022814
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2026-08-26
AI Technical Summary
Conventional action detection systems rely solely on image data, which limits their ability to accurately detect certain behaviors and often require extensive training data, making them inefficient and costly.
An action detection system that utilizes multiple types of specific data, including first specific data for extracting person and construction equipment parts, second specific data for representing chronological posture, and third specific data for area representation, combined through a multimodal model to efficiently detect predetermined actions.
The system effectively detects unsafe behaviors by extracting a richer set of features from multiple data types, reducing the need for extensive training data and improving accuracy while being cost-effective.
Smart Images

Figure 2026136942000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an action detection system for detecting a predetermined action and an action detection method for detecting a predetermined action, and particularly to an action detection system and an action detection method for detecting an unsafe action.
Background Art
[0002] As a system for recognizing a predetermined action, for example, there is one shown in Patent Document 1 regarding a person's handwashing operation.
[0003] FIG. 13 is a diagram showing an example of a determination method by the handwashing recognition system shown in Patent Document 1.
[0004] Referring to FIG. 13, in the handwashing recognition system, first, in the hand region extraction unit, an image captured by a digital camera or the like is input, and a hand region is extracted based on the input image. Next, in the motion detection unit, the posture of the user's hand is detected from the hand region data, and the operation steps being performed by the user are detected.
[0005] Next, in the detergent foam region detection unit, a color component specified in advance within the hand region is detected. At this time, if the detected color component includes a color component corresponding to the foam of the specified detergent, it is determined that the user is using an appropriate detergent.
[0006] Next, in the determination unit, the amount of detergent foam adhering to the user's hand is detected, and if it is more than a predetermined threshold value, it is determined that handwashing is being performed with a sufficient amount of detergent foam. Note that when making the determination, the threshold value is set after learning using a dataset of images of the handwashing operation taken in advance, and the determination is made based on comparison with the threshold value.
[0007] In such a handwashing recognition system, by extracting the hand region and detecting the detergent foam region from the input image, it is possible to determine whether the user is performing handwashing with an appropriate detergent and a sufficient amount of foam.
Prior Art Documents
[0008] [Patent Document 1] Patent No. 7447998 [Overview of the project] [Problems that the invention aims to solve]
[0009] However, conventional systems like the one described above analyze data based on only one type of image data (the input image), which means that they may not be able to properly detect certain types of behavior, or they may require a massive amount of training data to properly detect behavior. Therefore, there has been a great need for a behavior detection system that can efficiently detect specific behaviors.
[0010] This invention was made to solve the above-mentioned problems and aims to provide an action detection system and an action detection method that can efficiently detect predetermined actions. [Means for solving the problem]
[0011] To achieve the above objective, the invention described in claim 1 is an action detection system for detecting a predetermined action at a work site where a person and construction equipment are present, comprising: an image acquisition unit that acquires images constituting video footage taken at the work site; a specific data creation unit that creates multiple types of specific data relating to specific information of a person from the images; and a detection unit that acquires multiple types of specific data and combines them to detect a predetermined action.
[0012] The invention described in claim 2 is the configuration of the invention described in claim 1, wherein multiple types of specific data are created by performing different processing on the image.
[0013] The invention described in claim 3 is the configuration of the invention described in claim 1 or claim 2, wherein the specific data includes first specific data which has been processed to cut out parts of a person and construction equipment, second specific data which represents the posture of a person in chronological order, and third specific data which represents the area of the person and the area of the construction equipment, and the detection unit detects a predetermined action by linking the first specific data, the second specific data and the third specific data.
[0014] The invention described in claim 4 is configured in the invention described in claim 1 or claim 2, wherein the detection unit extracts a score related to a predetermined action based on numerically represented features from specific data, and detects the predetermined action by comparing the extracted score with a set threshold.
[0015] The invention described in claim 5 is an action detection method for detecting a predetermined action at a work site where a person and construction equipment are present, comprising: an image acquisition step of acquiring images constituting video footage taken at the work site; a determination step of determining whether or not a person is included in the image; a specific data creation step of creating multiple types of specific data relating to specific information of a person from the image; and an action detection step of acquiring multiple types of specific data and linking them to detect a predetermined action, wherein the specific data creation step creates multiple types of specific data if it is determined that a person is included, and does not create multiple types of specific data if it is determined that a person is not included.
[0016] The invention described in claim 6 is the configuration of the invention described in claim 5, wherein the determination step includes setting a flag that distinguishes an image determined to contain a person as an image for creating specific data. [Effects of the Invention]
[0017] According to the behavior detection system of this invention, since multiple types of specific data are used in the detection unit, predetermined behaviors can be detected efficiently. [Brief explanation of the drawing]
[0018] [Figure 1]It is a diagram for explaining an action detection system according to an embodiment of the present invention. [Figure 2] It is a diagram showing an example of an image acquired by an image acquisition unit. [Figure 3] It is a diagram showing an example of an image in which portions of a person and construction equipment are cut out from the image shown in FIG. 2. [Figure 4] It is a diagram showing an example of posture data representing a person's posture in a time series. [Figure 5] It is a diagram showing an example of image data representing a person's area and a construction equipment area. [Figure 6] Diagram for explaining a multimodal model [Figure 7] It is a flowchart showing an action detection method according to an embodiment of the present invention. [Figure 8] It is a flowchart showing a flow of outputting detection results to a display device. [Figure 9] It is a diagram for explaining post-processing. [Figure 10] It is a diagram showing an example of a detection result list screen output to a display device. [Figure 11] It is a diagram showing an example of an unsafe action detection result details screen output to a display device. [Figure 12] It is a diagram showing an example of images before and after the image displayed in FIG. 11. [Figure 13] It is a diagram showing an example of a determination method by a handwashing recognition system shown in Patent Document 1.
Embodiments for Carrying Out the Invention
[0019] FIG. 1 is a diagram for explaining an action detection system according to an embodiment of the present invention.
[0020] Action Detection System 1 is a system for detecting predetermined actions in a work site where people and construction equipment are present. "People" are not particularly limited, but include, for example, workers performing various tasks at the work site (construction workers, concrete workers, pipework workers, civil engineering workers, building workers, scaffolding and demolition workers, etc.) and site supervisors managing the process. "Construction equipment" refers to materials and machinery used in civil engineering and construction work, and is not particularly limited, but includes, for example, trucks such as dump trucks, cranes, excavators, concrete trucks and other vehicles, stepladders, scaffolding, ladders, timber, steel frames, metal materials, temporary materials, etc. "Predetermined actions" refer to actions by people that alter their relationship with construction equipment, and are not particularly limited, but include, for example, actions that could endanger the safety of the person or those associated with that person (unsafe actions). "Work site" is not particularly limited, but includes, for example, a place where construction work is being carried out or its surroundings. The behavior detection system can be applied, for example, to detecting unsafe behaviors performed by individuals near construction equipment at a work site.
[0021] Referring to Figure 1, the behavior detection system 1 detects a predetermined behavior (unsafe behavior) using images that constitute video footage of the work site captured by the camera 11. The video footage of the work site is transmitted via a network (not shown). The network can be wired or wireless, and any type of communication network such as the Internet, LAN (Local Area Network), or VPN (Virtual Private Network) can be used.
[0022] The imaging device 11 is a device that captures an image of an object as a single or continuous image or video, and is, for example, a network camera such as an IP (Internet Protocol) camera or a digital camera installed at a work site. The imaging device is installed in a position that can capture the area where unloading or loading work is performed at the work site. In this embodiment, one imaging device 11 is installed, but the number of devices installed is not particularly limited. Multiple imaging devices may also be installed in an arrangement that provides different fields of view.
[0023] The behavior detection system 1 mainly consists of an image acquisition unit 21 that acquires images of the work site, a specific data creation unit 22 that creates multiple types of specific data related to the identification of people from the images of the work site, a detection processing unit 23 that acquires multiple types of specific data and performs processing to detect unsafe behavior, and a storage unit 24 that stores information for performing the detection processing.
[0024] The image acquisition unit 21 acquires images that make up the video footage taken at the work site via the network. In this embodiment, an image refers to an image obtained by dividing the video footage taken by the shooting device 11 into predetermined intervals (for example, 60 frames, 30 frames, or 15 frames per second) after it has been downloaded as a video file. Note that the image is not limited to images divided from video, but may also be other types of data such as video. Alternatively, the shooting device 11 may take a photograph, and the data of that photograph may be directly input to the image acquisition unit 21 as an image. The image acquisition unit 21 also determines whether or not a person is included in the acquired image, and sets a flag for images that are determined to include a person. The image acquisition unit 21 then stores the flagged images and the unflagged images in the storage unit 24.
[0025] The specific data creation unit 22 acquires flagged images from the image acquisition unit 21 either directly or via the storage unit 24, and creates multiple types of specific data related to the identification information of a person. Identification information refers to information used to identify a person, including information that identifies only the person's posture, information that identifies only the positional relationship between the person and construction equipment, and information that identifies both the person's posture and the positional relationship between the person and construction equipment. The multiple types of specific data are created by performing different processes on a single image (an image of a work site) to extract features related to the identification information of a person. Features are numerical representations of the degree of a feature. For example, a sufficient amount of data that should be judged as having high and low feature values is pre-trained, and the degree of the feature is determined by its similarity to the training data. Alternatively, actions that are not predetermined actions (unsafe actions) are pre-trained, and the degree of the feature is determined by its deviation from the training data. In this embodiment, the specific data creation unit 22 creates three types of specific data (first specific data 221, second specific data 222, and third specific data 223), but there are only two or more types of specific data, and the number of types of specific data is not particularly limited. Details of the specific data will be described later. The specific data creation unit 22 stores the feature quantities extracted from each type of specific data as feature quantity data in the storage unit 24.
[0026] The detection processing unit 23 includes a detection unit 231 that acquires multiple types of specific data from the specific data creation unit 22 and links them together to detect a predetermined action, and a post-processing unit 232 that acquires the detection results from the detection unit 231 and performs post-processing. Post-processing refers to processing the detection results after an unsafe action has been detected in a way that is easy for the user to handle. For example, this could include displaying the detection results on the display device 12 using an application, or rearranging the detection results to make them easy for the user to view or making them editable.
[0027] The detection unit 231 acquires multiple types of specific data from the specific data creation unit 22 and combines them to detect a predetermined behavior (unsafe behavior). The detection unit 231 extracts a score related to the predetermined behavior (unsafe behavior) based on numerical features derived from the specific data, and detects the unsafe behavior by comparing the extracted score with a set threshold. The score is a numerical value in the range of 0 to 1, and the higher the value, the more likely the predetermined behavior is to have occurred; it is also called the probability. In this embodiment, the detection unit 231 extracts a score related to unsafe behavior using three types of specific data and a multimodal model 244. The detection unit 231 stores the detection results in the storage unit 24.
[0028] The multimodal model 244 is configured to extract a score related to unsafe behavior by linking features extracted from multiple types of specific data via separate learning models in order to detect a predetermined behavior. For example, one could use three learning models. The multimodal model 244 is pre-trained using machine learning or the like with training data.
[0029] The multimodal model 244 is trained, for example, using a designated learning device. The trained multimodal model 244 generated by the learning device is downloaded from the learning device to the behavior detection system 1 via a network or a portable storage medium and stored in the storage unit 24.
[0030] The learning device consists of a personal computer, a server computer, etc., and includes a control unit, a memory unit, a communication unit, an input unit, a display unit, a reading unit, etc., and each of these units is interconnected via a bus. The memory unit of the learning device stores a control program and a feature extraction model, as well as a learning data database (DB) that stores learning data used to train the feature extraction model. The learning data DB stores learning data that associates information indicating a predetermined type of action (ground truth label) with images taken of a person performing each type of action. For example, if the unsafe action is jumping off the back of a truck, the learning data DB may consist of learning images (e.g., descending from the truck bed using a ladder) that include an image of a person jumping off the back of a truck taken at a work site and an image that does not include the person jumping off, and ground truth data indicating the presence or absence of the unsafe action in the learning images.
[0031] In the process of generating the multimodal model 244, the control unit of the learning device first retrieves training data from the training data database. The control unit then performs training processing using the retrieved training data. Training processing is performed on three types of training models: the first feature extraction model, the second feature extraction model, and the third feature extraction model.
[0032] The first feature extraction model takes images of the work site, processed to extract portions of people and construction equipment, as input. The first feature extraction model is then configured to output the feature quantities of the input images of people and construction equipment.
[0033] The second feature extraction model takes time-series posture data of a specific person from images of the work site as input. The second feature extraction model is then configured to output the features of the input time-series posture data.
[0034] The third feature extraction model extracts the areas of people and construction equipment from the captured images of the work site and inputs them as a map image. The third feature extraction model is then configured to output the features of the input map image.
[0035] The multimodal model 244 has thresholds set for the scores that determine a particular behavior.
[0036] Finally, the multimodal model 244 combines the features output from the first, second, and third feature extraction models to extract a score related to unsafe behavior and compares it to a threshold. The multimodal model 244 determines that a predetermined behavior has been performed under certain conditions.
[0037] The post-processing unit 232 performs post-processing on the detection results that have been detected. First, the post-processing unit 232 obtains the score of the detection result of unsafe behavior, distinguishes between sections where the score is above the threshold and sections where the score is not above the threshold, and extracts the sections where the score is above the threshold to create a detection section. Then, the post-processing unit 232 outputs the detection section to the display device 12.
[0038] The functional components of the aforementioned behavior detection system 1 (image acquisition unit 21, specific data creation unit 22, detection processing unit 23) are realized, for example, by a hardware processor such as a CPU (Central Processing Unit) executing a program (software). Some or all of these components may be realized by hardware such as a GPU (Graphics Processing Unit) or MPU (Micro Processing Unit), or by the cooperation of software and hardware. The program may be stored in advance in a storage device such as an HDD (Hard Disk Drive) or flash memory, or it may be stored in a removable storage medium such as a DVD or CD-ROM and installed when the storage medium is inserted into a drive device. However, the program does not necessarily have to be installed from a storage medium; it may also be downloaded from an external device via a network or the like.
[0039] The storage unit 24 includes RAM (Random Access Memory), flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), etc. The storage unit 24 stores programs for executing various processes of the behavior detection system 1, and temporary data used when performing these processes.
[0040] The memory unit 24 stores, for example, a video file 241, feature data 242, detection result data 243, and a multimodal model 244.
[0041] The display device 12 is, for example, a computer or smartphone used by the user, or a liquid crystal display or organic EL display installed in a control room for managing the work site, and displays the detection results. Details of the display method will be described later.
[0042] Conventionally, techniques for detecting specific actions from video footage captured by a camera have involved analyzing the captured images directly. This allows for the detection of, for example, a target person or the identification of the person's movements. However, such conventional techniques utilize information obtained from only one modality—the captured image—making it difficult to extract information such as the fluctuating relationships between a person, the objects they use, and objects surrounding them, which can be crucial in detecting unsafe behavior. In other words, simply detecting the positions of a person and objects is insufficient to determine if a person is engaging in unsafe behavior. Similarly, identifying a person's movements is also insufficient because it doesn't reveal the relationship between the person and objects, making it difficult to determine if unsafe behavior is occurring. Therefore, conventional methods of directly analyzing captured video footage have limitations in the amount of information extracted, or sometimes insufficient information, making it difficult to accurately detect unsafe behavior, and thus improving accuracy has been required.
[0043] For example, to detect unsafe behavior by analyzing only a single modality, one could consider using machine learning with a large amount of training data on unsafe behavior beforehand. However, creating training data by photographing unsafe behavior at each work site and then performing machine learning is time-consuming and impractical. Furthermore, it requires the construction of a large dataset, which increases costs. Therefore, there has been a need for a more efficient method to detect unsafe behavior.
[0044] Therefore, in this embodiment, the detection unit 231 is configured to detect unsafe behavior by linking multiple types of specific data using a multimodal model 244. As a result, since multiple types of specific data are used by the detection unit 231, a wealth of features can be obtained extracted from multiple types of specific data even with a small amount of data, and unsafe behavior can be detected with a certain level of accuracy. Thus, unsafe behavior can be detected efficiently. Furthermore, since it does not require an enormous amount of training time to train a large amount of training data, it is an action detection system with good learning efficiency and cost advantages. Moreover, the accuracy of detecting unsafe behavior is improved when the amount of training is the same as with a single modality.
[0045] Furthermore, in this embodiment, the detection unit 231 does not input the captured image directly into the learning model, but rather inputs multiple types of specific data created by performing different processes on the captured image, and connects them to detect unsafe behavior. By configuring it in this way, different features related to a person's movement are extracted from each of the different specific data, so a richer set of features can be extracted compared to analyzing only a single modality. Therefore, unsafe behavior can be detected efficiently.
[0046] Furthermore, in this embodiment, the detection unit is configured to extract a score related to a predetermined behavior based on numerically represented features from specific data, and to detect unsafe behavior by comparing the extracted score with a set threshold. As a result, since the features are used as numerical values in the detection unit, it becomes easier to detect unsafe behavior.
[0047] Next, we will explain the details of the specific data.
[0048] The specific data creation unit 22 acquires images of the work site from the image acquisition unit 21 and creates the first specific data 221, the second specific data 222, and the third specific data 223.
[0049] Figure 2 shows an example of an image acquired by the image acquisition unit, Figure 3 shows an example of an image from which the parts of a person and construction equipment have been cut out from the image shown in Figure 2, Figure 4 shows an example of posture data representing the posture of a person in chronological order, and Figure 5 shows an example of image data representing the region of a person and the region of construction equipment.
[0050] Referring to Figure 2, as an example of this embodiment, Figure 2 shows that two people 100 and 101 and one truck 102 (construction equipment) are detected, and frames are displayed surrounding each of the people 100 and 101 and the truck 102.
[0051] Referring to Figure 3, the first specific data 221 is an image obtained by processing an image of a work site to extract portions of people and construction equipment that are close to the construction equipment (i.e., processing to remove unnecessary parts such as roads that are not related to people or construction equipment). In Figure 2, because the positional relationship between the person 100 on the left and the truck 102 is close, the portion of the person 100 on the left and the truck 102 is extracted to create the image shown in Figure 3. In other words, the first specific data 221 is extracted only from the parts necessary for detecting unsafe behavior. By processing in this way, the detection process can be narrowed down to the target of analysis, so unsafe behavior can be efficiently detected when combined with other specific data described later. Furthermore, even if the location, background, and surrounding circumstances are different, unsafe behavior can be detected by combining it with other specific data described later, eliminating the need to perform machine learning on unsafe behavior for each work site and increasing the versatility of the behavior detection system.
[0052] Referring to Figure 4, the second specific data 222 is an image representing the posture of a person in close proximity to construction equipment in a time series, taken from images of the work site. For example, the second specific data 222 is an image that detects multiple body parts of a person and shows skeletal information indicating the position of each of these body parts. Examples of body parts include the head, right shoulder, left shoulder, right arm, left arm, right leg, left leg, back, waist, and equipment such as clothing and tools attached to the person. In Figure 4, skeletal information is displayed for every 6 frames out of 90 frames divided into 3 seconds. Note that in Figure 4, 3 out of 15 frames are omitted, so skeletal information for 12 frames (t=6, t=12, t=18, t=24, t=30, t=36, t=54, t=60, t=66, t=72, t=78, t=84) is displayed in time series. By processing the data in this way, it becomes possible to identify a person's posture based on the time-series movements of multiple body parts. This makes it easier to detect unsafe behavior when the first identification data 221 is linked to the third identification data 223, which will be described later. Furthermore, the second identification data 222 may be subject to additional tracking to identify that it is the same person.
[0053] Referring to Figure 5, the third specific data 223 is a map image representing the area of a person close to construction equipment and the area of the construction equipment, with each area color-coded. In Figure 5, person 100 is shown in red, truck 102 in blue, and ground 103 in green. By showing the two-dimensional area between the person and the construction equipment in a map image in this way, it becomes easier to extract information about the relationship between the person and the loading area of the construction equipment. Therefore, when the third specific data 223 is linked with other specific data, it becomes easier to detect unsafe behavior. In this embodiment, the area of ground 103 is also learned as a feature. With this configuration, in addition to the positional relationship between the person and the construction equipment, information about the relationship between the area of person 100 and the area of ground 103 can be obtained. Therefore, when detecting unsafe behavior such as "climbing onto and jumping off the truck bed," it is possible to determine whether the area of person 100 is far from the area of ground 103, making it easier to detect unsafe behavior.
[0054] Thus, since multiple types of specific data are created by performing different processing on a single captured image, features are extracted using multiple types of specific data different from the captured image. Therefore, unsafe behavior can be efficiently detected. Furthermore, by linking the first, second, and third types of specific data described above to detect unsafe behavior, multiple types of information regarding the relationship between people and construction equipment can be obtained. Therefore, unsafe behavior can be efficiently detected.
[0055] Here, we will describe the multimodal model 244 in this embodiment.
[0056] Figure 6 illustrates the multimodal model.
[0057] Referring to Figure 6, the multimodal model 244 comprises, for example, a feature extraction model, a feature-connected input layer 244D, and a fully connected layer 244E.
[0058] The feature extraction model takes specific data as input and extracts features. Specifically, features from the first specific data 221 are output from the first feature extraction model 244A. Features from the second specific data 222 are output from the second feature extraction model 244B. Features from the third specific data 223 are output from the third feature extraction model 244C. For example, known models such as ResNet18 can be used for the first feature extraction model, ST-GCN for the second feature extraction model, and ResNet18 for the third feature extraction model.
[0059] The feature concatenation input layer 244D is the input layer for the multimodal model 244, and receives features from the first specific data 221, the second specific data 222, and the third specific data 223, respectively. In this embodiment, the feature concatenation input layer 244D employs a method of concatenation to sequentially combine the vectors of the extraction results from the first feature extraction model 244A, the second feature extraction model 244B, and the third feature extraction model 244C in a single column. However, a method called Max Pooling may also be used to output the maximum value of each model and combine them. Furthermore, in this embodiment, features are extracted by the first feature extraction model 244A, the second feature extraction model 244B, and the third feature extraction model 244C, and then converted into vectors before being combined. However, it may also be configured to combine different data modalities at different times, or to combine data modalities that have not undergone conversion processing. For example, this could include structures that combine elements without converting them to vectors, or structures that combine elements that perform vector conversion with those that do not. Furthermore, the structure may be configured to repeatedly extract features and combine them in multiple layers.
[0060] The fully connected layer 244E is a layer that fully connects multiple types of features and outputs a score related to unsafe behavior. In this embodiment, the fully connected layer employs a logistic regression model, but other methods, such as a deep learning model using a deep neural network, may also be employed.
[0061] In this way, by combining multiple data modalities using the multimodal model 244, it is possible to extract a richer set of features and detect unsafe behaviors compared to analyzing from a single data modality (images of the work site in their original state). Therefore, it becomes unnecessary to build a large dataset, and unsafe behaviors can be detected efficiently. It should be noted that linking multiple types of features extracted from multiple types of specific data is also included as one form of linking multiple types of specific data.
[0062] Next, we will explain the behavior detection method.
[0063] Figure 7 is a flowchart showing a behavior detection method according to an embodiment of the present invention.
[0064] The behavior detection method may be implemented by the behavior detection system 1 described above. Specifically, the behavior detection method may be implemented by executing a program stored in the storage device of the behavior detection system 1 by a processor. Prior to the start of the process shown in Figure 6, the user specifies the unsafe behavior to be detected. In this embodiment, "climbing onto and jumping off the truck bed" is specified as the unsafe behavior, but other types of unsafe behaviors may also be specified. Furthermore, it can be applied to predetermined behaviors other than unsafe behaviors that may occur at the work site.
[0065] Referring to Figure 7, the image acquisition unit 21 acquires images that make up the video footage captured at the work site by the shooting device 11 (S11). In this embodiment, the shooting device 11 is installed so as to capture the field of view of the area where unloading and loading work is performed at the work site, and is set to take pictures during the time when the work is performed (for example, from 7:00 to 17:00). The shooting device 11 may be installed to capture other locations or fields of view. The shooting device 11 may also be set to take pictures at other times or 24 hours a day. The captured video footage is then transferred to the action detection system 1 at any arbitrary timing and input to the image acquisition unit 21 as segmented images.
[0066] Next, the image acquisition unit 21 determines the location of the work site based on the acquired image and determines whether or not it contains people or construction equipment (S12). At this time, the image acquisition unit 21 may refer to the feature data of people and construction equipment stored in the storage unit 24 and make the determination by comparing it with the feature data. When the image acquisition unit 21 determines that people and construction equipment are included, it detects people and construction equipment from the acquired image.
[0067] In this embodiment, the image acquisition unit 21 detects parts of objects such as people and construction equipment (e.g., trucks) from the acquired image and outputs the position of the detected part. Here, it is assumed that a rectangular detection frame enclosing the detected part is output as the position of the detected part of the object. Note that the detection frame may be other shapes such as circles or ellipses. In this embodiment, the area of the detected person is represented by a rectangle, and the area of the construction equipment (e.g., trucks) is represented by a rectangle, and it is determined whether the overlap ratio (IoU, Intersection over Union) between any rectangle of the person and any rectangle of the truck is greater than a preset threshold (e.g., 0.005). If the overlap ratio (IoU) between the rectangle of the person and all the rectangles of the scaffolding is equal to 0 (i.e., the rectangle of the person and all the rectangles of the scaffolding do not overlap), it is determined to be the image to be detected.
[0068] Subsequently, the image acquisition unit 21 flags images that are determined to contain people (S13) so that it can distinguish images from the acquired images for which specific data is to be created. If the overlap ratio (IoU) between any person's rectangle and all the other construction equipment (e.g., scaffolding) is not equal to 0, the image is not flagged as a target for detection.
[0069] Subsequently, the images processed by the image acquisition unit 21 (images with flags and images without flags) are stored in the storage unit 24.
[0070] Next, the identification data creation unit 22 acquires a list of flagged images from the image acquisition unit 21 either directly or via the storage unit 24, and creates multiple types of identification data related to the identification information of a person (S14). Specifically, first, the identification data creation unit 22 processes the detected person and construction equipment parts to create the first identification data 221. Then, the identification data creation unit 22 creates the second identification data 222, which represents the person's posture in chronological order, from the image of the work site transmitted from the image acquisition unit 21. Furthermore, the identification data creation unit 22 creates the third identification data 223, which represents the area of the person and the area of the construction equipment, from the image of the work site transmitted from the image acquisition unit 21. After that, the identification data creation unit 22 transmits the first identification data 221, the second identification data 222, and the third identification data 223 to the detection unit 231. In other words, in this embodiment, instead of transmitting images of the work site to the detection unit 231 in their original state as in the conventional method, the specific data creation unit 22 performs different processing on each image of the work site to create the first specific data 221, the second specific data 222, and the third specific data 223, which are then transmitted to the detection unit 231. For images that have not been flagged, the processing from the specific data creation step (S14) onward is not performed.
[0071] The detection unit 231 acquires multiple types of specific data (first specific data 221, second specific data 222, and third specific data 223) transmitted from the specific data creation unit 22 (S15), combines them (S16), and detects unsafe behavior (S17). Specifically, the detection unit 231 quantifies features from each type of specific data. Based on the quantified features, it extracts a score related to unsafe behavior and detects unsafe behavior by comparing the extracted score with a preset threshold. In this embodiment, the detection unit 231 extracts a first feature from the target image using the first specific data 221 and the first feature extraction model 244A. The detection unit 231 also extracts a second feature from the target image using the second specific data 222 and the second feature extraction model 244B. Furthermore, the detection unit 231 extracts a third feature from the target image using the third specific data 223 and the third feature extraction model 244C. Subsequently, the detection unit 231 uses the multimodal model 244 described above to extract scores related to unsafe behavior in the target image, and compares the extracted scores with a threshold to detect unsafe behavior. This completes the unsafe behavior detection process.
[0072] The behavior detection method described above allows for more efficient detection of unsafe behavior compared to detecting unsafe behavior using images taken at the work site as is, because the detection unit 231 uses multiple types of specific data. Furthermore, when a person is included in the image acquired by the image acquisition unit 21, the detection process is performed in the behavior detection step, and when a person is not included in the image, the process of creating specific data is skipped, thus making the detection process more efficient and reducing false positives.
[0073] Furthermore, the aforementioned behavior detection method includes a step of flagging images in which an object is detected. This allows for skipping the subsequent process of detecting unsafe behavior in images that are not flagged, thereby improving processing efficiency and reducing false positives. In addition, since only images corresponding to unsafe behavior can be extracted, it becomes easier to verify the detection results.
[0074] Next, we will explain the process of outputting the detection results of unsafe behavior to the display device 12.
[0075] The detection unit 231 transmits the detection result of unsafe behavior to the display device 12 via the post-processing unit 232.
[0076] Figure 8 is a flowchart showing the process of outputting detection results to a display device, and Figure 9 is a diagram explaining the post-processing.
[0077] The post-processing unit 232 obtains the detection results after the detection process has been completed (S21). Referring to Figure 9(1), a score (detection score) is given for images in which unsafe behavior is detected among the detection results. In this embodiment, for example, if there are multiple people to be detected in one image, the multimodal model outputs a score for each person. Therefore, the post-processing unit 232 obtains the maximum value from the scores of the multiple people in one image (S22). The maximum value is the highest numerical score among the scores related to unsafe behavior present for each person in one image. This maximum value is then treated as the score for that one image. At this time, images for which no detection results exist are interpolated with a detection score of 0.
[0078] Next, referring to Figure 9(2), the image intervals (frame intervals) in which detection scores above a threshold exist are extracted within a fixed window width (S23). For example, frames with detection scores above the threshold are represented by a dotted line with a value of 1, and frames without are represented by a value of 0. Then, referring to Figure 9(3), the extracted results are divided into consecutive frame intervals to create detection sections 31 and 32 (S24).
[0079] Next, the frame numbers with the highest scores 33 and 34 within the detection sections 31 and 32 are obtained (S25). The highest score refers to the highest numerical score among the scores related to unsafe behavior present in each image within the detection section. Then, the feature quantities for a frame of a certain width centered on the frame number with the highest score are referenced (S26). After that, the images of the frame of a certain width as detection results, along with their scores and feature quantities, are stored in the storage unit 24 (S27). This completes the post-processing.
[0080] Next, the post-processing unit 232 transmits the post-processed detection results to the display device 12. The detection results are then displayed on the display device 12.
[0081] In this way, by extracting frame intervals with detection scores above a threshold within a certain window width and creating detection sections 31 and 32, images with particularly high scores related to unsafe behavior can be extracted as representative detection results and displayed on the screen. Therefore, the number of images displayed is reduced compared to displaying all results above the threshold, reducing the burden on the user when determining whether or not unsafe behavior occurred. Consequently, the usability of the behavior detection system 1 is improved.
[0082] Figure 10 shows an example of a detection results list screen displayed on a display device.
[0083] The display device 12 displays the detection results (unsafe behavior detection results) in which unsafe behaviors have been detected. First, the display device 12 displays a list of detection results for each detection section. For example, if the user selects a specific day, the unsafe behavior detection results detected on that day will be displayed on the display device for each detection section. For example, in Figure 10, seven unsafe behavior detection results that may constitute unsafe behaviors are displayed.
[0084] Within each unsafe behavior detection result frame, a representative image 35 from the detection frame is displayed at the top, text information 36 is displayed in the middle, and a details button 37 is displayed at the bottom. The unsafe behavior detection score 38 is displayed in the upper left of the image 35 in the top row. In this embodiment, the unsafe behavior detection score 38 is expressed as 0 to 100%, and a higher score 38 means that the likelihood of unsafe behavior is judged to be high. In this embodiment, the system is set to display the detection result only when the unsafe behavior detection score 38 is 90% or higher. An icon 39 representing the result after checking the display of the detection result is displayed in the lower right of the image 35 in the top row. For example, icons 39 can be assigned to different cases such as "I checked the display of the detection result and unsafe behavior was indeed detected," "I checked the display of the detection result, but it is not possible to determine from the image whether unsafe behavior is occurring, and it is necessary to check the original video for analysis," "I checked the display of the detection result, but no unsafe behavior was occurring," or "The display of the detection result has not been checked." Subsequently, the user checks the detection results displayed on the display device 12 to confirm whether or not the unsafe behavior detected by the detection unit 231 actually occurred. By assigning icons 39 to the display device 12 in this way, it is possible to confirm whether or not the user has checked the display of the detection results for the unsafe behavior detection results, and if so, the user makes an actual judgment, thereby enabling more accurate confirmation of the unsafe behavior detection results.
[0085] The middle section displays the type of unsafe behavior, the location where it was filmed, and the date and time.
[0086] Clicking the details button 37 at the bottom will allow you to view detailed information only for the relevant unsafe behavior detection results.
[0087] Figure 11 shows an example of the detailed screen of the unsafe behavior detection result output to the display device, and Figure 12 shows an example of the images before and after the image displayed in Figure 11.
[0088] Referring to Figure 11, the left side of the details screen displays a representative image 35 from the detected frames, while the right side displays the unsafe behavior detection score 38 and an icon 39. In Figure 9, the unsafe behavior detection score is 100%, and the detection result display check also determined that the behavior was unsafe. Therefore, "Detected" is selected from the buttons 40 that represent the result after checking the detection result display. If the user has not completed the detection result display check, they can perform the check and then select the button 40 that represents the result after checking the detection result display. Referring to Figure 12, the user can also view images 41 and 42 before and after the representative image 35. The display device 12 can also display multiple images from the detection section as a video on the display screen. In this way, detection of unsafe behavior becomes easier because detection sections with a high probability of unsafe behavior occurring can be extracted based on the unsafe behavior detection score 38. Furthermore, because images 41 and 42 before and after the single image 35 can be viewed, it becomes possible to detect unsafe behavior more reliably. In this embodiment, one image is displayed before and after the first image, but the number of images before and after is not limited to one. For example, there may be multiple images before and after the first image; for instance, five images may be displayed before and after the first image, resulting in a slideshow of 11 images in total.
[0089] In the above embodiment, learning processing was performed for "climbing onto and jumping off the bed of a truck," which is one of the unsafe behaviors at a work site, as a predetermined action. However, learning processing may also be performed for other types of unsafe behaviors. For example, actions such as a person going under a suspended load during crane operation or standing on the top platform of a stepladder may be performed. Furthermore, learning processing may also be performed for predetermined actions other than unsafe behaviors. For example, work such as tying reinforcing bars, using a (concrete) vibrator, and chipping work may be performed. All of these involve using a specific tool and performing a specific action.
[0090] Furthermore, in the above embodiment, the detection unit in the behavior detection system and behavior detection method was configured to detect "climbing onto and jumping off the bed of a truck," which is one of the unsafe behaviors at the work site. However, it may also be configured to detect other types of unsafe behaviors as described above. Alternatively, it may be configured to detect predetermined behaviors other than the unsafe behaviors described above.
[0091] Furthermore, in the above embodiment, the specific data creation unit created specific data having a specific configuration, but it may also create data having other configurations as specific data. For example, it may include 3D skeletal information, text information, and audio information, or it may combine images and non-image information.
[0092] Furthermore, in the above embodiment, different processing was performed on the images of the work site for all of the multiple types of specific data. However, if at least one type of specific data is applicable, images of the work site that have not been processed may be used for the other types of specific data.
[0093] Furthermore, in the above embodiment, multiple types of specific data were linked using a multimodal model, but unsafe behavior may also be detected by linking the data using methods other than a multimodal model.
[0094] Furthermore, in the above embodiment, the image acquisition unit acquired images of the work site, but it may also acquire information other than images, as long as it uses at least one type of image.
[0095] Furthermore, in the above embodiment, the determination step flagged images in which a person was detected, but images detected by other methods may also be distinguishable. For example, they may be distinguishable by text, color, etc.
[0096] Furthermore, although the detection unit used a specific threshold in the above embodiment, it may be set to a threshold within a different numerical range.
[0097] Furthermore, in the above embodiment, the process proceeded to the specific data creation step only when it was determined that a person was included. However, the system may be configured to proceed to the specific data creation step regardless of whether a person is included or not. Also, the step for determining whether a person is included or not may be omitted.
[0098] Furthermore, in the above embodiment, both flagged and unflaged images were saved as detection results. However, it is also possible to extract only flagged images and save them as detection results. This configuration reduces the amount of data to be saved and makes the processing in the post-processing unit of the detection results more efficient.
[0099] Furthermore, in the above embodiment, after unsafe behavior was detected, it was displayed on a display device via a post-processing unit, but the post-processing unit is not necessary.
[0100] Furthermore, in the above embodiment, three types of specific data were linked simultaneously to detect a predetermined action, but this does not have to be done simultaneously. For example, the first specific data and the second specific data could be linked, and then the linked specific data and the third specific data could be linked to make a decision.
[0101] Furthermore, although the above embodiment was configured to output the detection result to the display device the morning after the day the image was taken, it is also possible to detect the detection while taking the image and communicate it immediately by voice or other means. Examples of communication methods include sounding an alert or sending a notification via an application.
[0102] Furthermore, in the above embodiment, the multimodal model included a first feature extraction model, a second feature extraction model, and a third feature extraction model, but these do not necessarily have to be included in the multimodal model. For example, a multimodal model consisting of a feature-connected input layer and a fully connected layer may be used to input features extracted using a known learning model, and a score related to unsafe behavior may be extracted. [Explanation of Symbols]
[0103] 1…Behavior detection system 21...Image acquisition unit 22...Specific Data Creation Department 221...First specific data 222...Second specific data 223...Third specific data 231...Detection unit In each figure, the same reference numeral indicates the same or corresponding part.
Claims
1. An action detection system for detecting predetermined actions in a work site where people and construction equipment are present, An image acquisition unit that acquires images constituting the video footage taken at the aforementioned work site, A data creation unit creates multiple types of data related to the identification information of the person from the aforementioned image, An action detection system comprising a detection unit that acquires multiple types of the aforementioned specific data and links them together to detect the predetermined action.
2. The behavior detection system according to claim 1, wherein multiple types of the aforementioned specific data are created by performing different processing on the image.
3. The aforementioned specific data includes first specific data that has been processed to extract the parts of the person and the construction equipment, Second specific data representing the posture of the aforementioned person over time, This includes a third specific data representing the area of the person and the area of the construction equipment, The behavior detection system according to claim 1 or claim 2, wherein the detection unit detects the predetermined behavior by linking the first specific data, the second specific data, and the third specific data.
4. The behavior detection system according to claim 1 or 2, wherein the detection unit extracts a score related to the predetermined behavior based on the numerically represented feature quantities from the specific data, and detects the predetermined behavior by comparing the extracted score with a set threshold.
5. A method for detecting predetermined actions in a work site where people and construction equipment are present, An image acquisition process to acquire images that make up the video footage taken at the aforementioned work site, A determination step of determining whether the aforementioned image includes the aforementioned person, A process for creating specific data, which involves creating multiple types of specific data related to the specific information of the person from the aforementioned image, The system includes an action detection step that acquires multiple types of the aforementioned specific data and combines them to detect the predetermined action, The behavior detection method wherein the process for creating specific data includes creating multiple types of specific data if it is determined that the person is included, and not creating multiple types of specific data if it is determined that the person is not included.
6. The behavior detection method according to claim 5, wherein the determination step includes setting a flag that distinguishes an image determined to contain the person as an image used to create the specific data.
Citation Information
Patent Citations
Hand-washing recognition system and hand-washing recognition method
JP7447998B2