Image processing device and image processing program

The image processing device uses skeletal coordinate analysis and directional fall detection to accurately identify falls in workers, addressing the challenge of distinguishing between standing and falling states when workers are positioned away from the surveillance camera.

JP7850378B2Active Publication Date: 2026-04-23SAXA
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
SAXA
Filing Date
2022-12-22
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing surveillance systems struggle to accurately detect falls in workers when they occur away from the surveillance camera, as the relative positions of body parts in images remain similar between standing and falling states, making it difficult to distinguish between the two.

Method used

An image processing device that determines a subject's posture using skeletal coordinates, employs omnidirectional and depth-direction fall determination mechanisms to identify falls in any direction, including the depth direction, by analyzing frame images from surveillance cameras.

Benefits of technology

The system effectively detects falls even when workers fall away from the camera by utilizing posture estimation and directional analysis, ensuring timely detection and appropriate action can be taken.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007850378000001
    Figure 0007850378000001
  • Figure 0007850378000002
    Figure 0007850378000002
  • Figure 0007850378000003
    Figure 0007850378000003
Patent Text Reader

Abstract

To properly detect a fall-down state even when a subject as a person falls down in a depth direction as a direction of separating from a monitor camera.SOLUTION: An attitude estimation unit 103 acquires skeleton coordinates of a subject from the latest frame image data to identify information on the subject. An omnidirectional fall-down determination unit 1041 determines whether or not the subject is in a fall-down state on the basis of the identified information. When the subject is determined not to be in the fall-down state, a depth direction fall-down possibility determination unit 1042 uses the skeleton coordinates of the subject in the frame image taken before N seconds and the latest skeleton coordinates to determine a possibility of fall-down in a depth direction. When determining that there is a possibility of the fall-down in the depth direction, a depth direction fall-down determination processing unit performs processing of determining whether or not the subject is in the fall-down state in the depth direction.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an apparatus and a program that analyze frame images constituting video data obtained by photographing a subject who is a person, and can detect a case where the subject has fallen.

Background Art

[0002] Patent Document 1 described later discloses a technique of generating a skeleton model of a subject by analyzing an image obtained by photographing a person riding in a vehicle as a subject, and distinguishing and determining whether the subject is in a standing state or a sitting state from the generated skeleton model. As a result, it is possible to take fall prevention measures such as prompting a person in a standing state to sit down. Further, Patent Document 1 describes that when it is determined from the generated skeleton model that the subject has fallen, it notifies the driver of the vehicle or the like or notifies the outside. As a result, it becomes possible to take measures such as quickly protecting a fallen person.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] For example, in workplaces handling hazardous chemicals, such as chemical plants, accidents can occur due to slipping on spilled liquids or losing consciousness from inhaling toxic gases. In such cases, if another worker is nearby, a quick response is possible, but often, due to mechanization or other factors, there are no other workers nearby. It is also conceivable that surveillance cameras could be used to have a supervisor monitor the footage in real time to quickly detect the occurrence of an accident. However, if the supervisor leaves their post for any reason and an accident occurs while monitoring is not taking place, the detection of the accident will be delayed.

[0005] Therefore, there is a need to enable unmanned monitoring of the work site without relying on human intervention, and to detect the occurrence of falls without delay and take appropriate action. Thus, as disclosed in Patent Document 1 mentioned above, it is conceivable to analyze images obtained by photographing a worker to generate a skeletal model of the worker, understand the worker's posture, and automatically detect falls. More specifically, it is conceivable to analyze the image data obtained by photography to identify the skeletal coordinates of each part of the worker (head, shoulders, waist, knees, ankles, etc.), and distinguish between a standing position and a fallen position based on the positional relationship of each part.

[0006] However, if a worker falls in the direction away from the surveillance camera (depth), the relative positions of the worker's body parts will appear the same as when standing, making it difficult to distinguish between a fallen and a standing state. For example, in an image of a worker, when the worker is standing, the coordinates of the body parts detected are arranged from top to bottom in the order of head, shoulders, hips, knees, and ankles. Now consider the case where the worker falls towards the surveillance camera, or in other words, falls so that their head is close to the camera. In this case, the coordinates of the body parts detected in the image of the worker are arranged from top to bottom in the order of ankles, knees, hips, shoulders, and head, which is the reverse of the standing state, allowing for proper detection of a fallen state.

[0007] Furthermore, consider the case where the worker falls sideways relative to the surveillance camera, in other words, where the worker's head is facing to the right or left relative to the surveillance camera. In this case, the coordinates of each part detected in the image of the worker will be in the order of head, shoulders, waist, knees, and ankles in the horizontal direction of the image, which is perpendicular to the order in the standing position, so the fall can be appropriately detected. Now, consider the case where the worker falls in the depth direction, away from the surveillance camera, that is, where the worker's head falls away from the surveillance camera. In this case, the coordinates of each part detected in the image of the worker will be in the order of head, shoulders, waist, knees, and ankles from top to bottom of the image, which is the same order as in the standing position. Therefore, when the worker falls in the depth direction, away from the surveillance camera, it is not possible to distinguish between a fall and a standing position.

[0008] In view of the above points, the present invention aims to enable the proper detection of a fall even when a subject, which is a person, falls in the depth direction, which is the direction away from the surveillance camera. [Means for solving the problem]

[0009] To solve the above problems, the image processing apparatus of the invention described in claim 1 is: An image processing device that determines whether or not a subject is in a fallen state using frame images that constitute video data obtained by continuously photographing a predetermined person as the subject, A posture estimation means that obtains the skeletal coordinates of the subject from the latest frame image, identifies information about the subject, and enables the estimation of the subject's posture, An all-directional fall determination means that determines whether or not the subject is in a fallen state based on the information identified by the posture estimation means, In the omnidirectional tipping determination means, if it is determined that the subject is not in a tipping state, a depth-direction tipping possibility determination means determines the possibility of tipping in the depth direction using the skeletal coordinates of the subject obtained from the frame image N seconds ago and the skeletal coordinates of the subject from the latest frame image obtained by the posture estimation means, The depth-direction tipping possibility determination means, when it is determined that there is a possibility of tipping in the depth direction, performs a depth-direction tipping determination processing means to determine whether or not the subject is in a state of tipping in the depth direction. It is characterized by being equipped with [the following features].

[0010] According to the image processing apparatus of the invention described in claim 1, the posture estimation means obtains the skeletal coordinates of the subject from the latest frame image, identifies information about the subject, and makes it possible to estimate the posture of the subject. Based on the information identified by the posture estimation means, the omnidirectional tipping determination means determines whether or not the subject is in a tipping state. If the omnidirectional tipping determination means determines that the subject is not in a tipping state, the depth direction tipping possibility determination means uses the skeletal coordinates of the subject from the frame image N seconds ago and the latest skeletal coordinates to determine the possibility of tipping in the depth direction. If the depth direction tipping possibility determination means determines that there is a possibility of tipping in the depth direction, the depth direction tipping determination processing means performs a process to determine whether or not the subject is in a tipping state in the depth direction. [Effects of the Invention]

[0011] According to this invention, even if a subject, such as a person, falls in the depth direction, which is away from the surveillance camera, the falling state can be appropriately detected. [Brief explanation of the drawing]

[0012] [Figure 1] This is a block diagram illustrating an example configuration of an image processing system and image processing apparatus according to an embodiment. [Figure 2]This diagram illustrates the difference between images of a worker standing and images of a worker lying down, when the worker's image is located in the central part of the captured image or above the central part. [Figure 3] This diagram illustrates the difference between an image of a worker standing and an image of a worker lying down, when the worker's image is located below the center of the captured image. [Figure 4] This diagram illustrates the calculation of the angle (predicted angle) between the floor depth direction (Z-axis) and the extension line from the ankle to the shoulder, which is used to determine whether the subject is standing or lying down. [Figure 5] This is a flowchart illustrating the processes performed by the image processing apparatus of the embodiment. [Figure 6] This is a flowchart following Figure 5. [Figure 7] This diagram illustrates the timing of the process for detecting a fall. [Modes for carrying out the invention]

[0013] The following describes one embodiment of the apparatus and program according to this invention, with reference to the figures. In the embodiment described below, for the sake of simplicity, we will use as an example the case in which the condition of workers in a workplace handling hazardous chemicals, such as a chemical plant, is monitored via a surveillance camera. However, this invention is not limited to workplaces handling hazardous chemicals, such as chemical plants, but is applicable to various situations in which it is necessary to quickly and appropriately detect when a person falls.

[0014] [Examples of image processing system and image processing device configurations] Figure 1 is a block diagram illustrating an example configuration of an image processing system and image processing apparatus according to an embodiment. As shown in Figure 1, a surveillance camera 2 is connected to the image processing apparatus 1 as an input device, and a monitor device (display device) 3 is connected to it as an output device.

[0015] The monitoring camera 2 is installed on the ceiling or the like at the work site such as a chemical plant, and photographs workers working at the work site from above as a subject to be monitored, and is an existing commercially available product. The monitoring camera 2 of this embodiment captures frame images at 30 frames per second and provides continuous video data to the image processing device 1. The monitor device 3 includes a thin display element such as an LCD (Liquid Crystal Display), and displays a video corresponding to the video signal output from the image processing device 1, and is, for example, an existing commercially available product for personal computers. Although omitted in FIG. 1, a speaker can also be connected to the image processing device 1 as an output device to output voice information.

[0016] As shown in FIG. 1, the image processing device 1 includes an input terminal 101T for video data, a video input unit 101, a video storage unit 102, a posture estimation unit 103, a fall determination unit 104, a determination result output unit 105, an output terminal 105T for video signals, a control unit 110, and a storage device 111. The control unit 110 is a microprocessor configured by connecting a CPU (Central Processing Unit), a ROM (Read Only Memory), a RAM (Random Access Memory), a non-volatile memory, etc., which are not shown, and realizes a function of controlling each part of the image processing device 1.

[0017] The storage device 111 is, for example, a device unit composed of a recording medium such as an SSD (Solid State Drive) and its driver, and performs recording of various data and the like on the recording medium, reading, changing, deleting, etc. of data and the like recorded on the recording medium. The storage device 111 stores and holds necessary data and programs, and is also used as a work area for temporarily storing intermediate data generated in various processes.

[0018] The input terminal 101T for video data constitutes the connection end with the surveillance camera 2. The video input unit 101 receives video data from the surveillance camera 2 connected through the input terminal 101T, converts it into data in a format that can be processed by the device itself, captures it, and performs the process of recording it in the video storage unit 102. In this embodiment, the video storage unit 102 is a device unit composed of an SSD and its driver, and stores the video data from the video input unit 101. Regarding the video data stored in the video storage unit 102, for example, frame images at points in time as needed, such as the frame image one second before a predetermined time point, the frame image two seconds before, etc., can be read out and utilized from the frame images constituting the video data.

[0019] The posture estimation unit 103 analyzes the latest frame image (frame image data) of the video data stored in the video storage unit 102, and acquires various information regarding the posture of the image of the operator (the subject to be monitored) included in the frame image. Specifically, the posture estimation unit 103 acquires the skeletal coordinates of the image portion of the operator included in the frame image. Using the acquired skeletal coordinates, the posture estimation unit 103 obtains the minimum X coordinate, minimum Y coordinate, maximum X coordinate, and maximum Y coordinate of each part of the image portion of the operator, specifies a rectangular area surrounding the entire image portion of the operator in the frame image, and specifies the vertical length and horizontal length of the rectangular area. The rectangular area is a rectangular area specified by both ends of the diagonal line, that is, the upper left is (the minimum X coordinate, the minimum Y coordinate), and the lower right is (the maximum X coordinate, the maximum Y coordinate).

[0020] Furthermore, the posture estimation unit 103 identifies the orientation and order of each part of the acquired image of the worker from the skeletal coordinates of each part. The posture estimation unit 103 also identifies the length from the shoulder to the waist and the length from the waist to the ankle of the image of the worker from the skeletal coordinates of each part of the acquired image of the worker. Based on the rectangular area surrounding the image of the worker identified by the posture estimation unit 103, the orientation and order of each part of the image of the worker, the length from the shoulder to the waist, and the length from the waist to the ankle of the image of the worker, the posture of the image of the worker in that frame can be estimated. Note that the acquisition of skeletal coordinates can be performed using various methods such as background subtraction, mean shift, and pattern matching, and since these methods are known techniques, a detailed explanation will be omitted.

[0021] The fall detection unit 104 determines whether the worker's image portion in the latest frame image is in a fallen state, based on the information identified by the posture estimation unit 103, and also using past frame images if necessary. In this embodiment, the fall detection unit 104 performs fall detection for the worker in all directions using the positional relationships of each part of the worker's image portion. However, if the worker falls in the depth direction, which is the direction away from the surveillance camera 2, it is impossible to distinguish between the standing state and the fallen state because the positional relationships of each part of the worker's image portion, identified by the skeletal coordinates of the worker's image portion, appear the same in both the standing state and the fallen state.

[0022] Therefore, the fall detection unit 104 of the image processing device 1 in this embodiment is designed to appropriately determine that a fall has occurred, even if the worker falls in the depth direction, away from the surveillance camera 2. For this reason, as shown in Figure 1, the fall detection unit 104 includes an all-directional fall detection unit 1041, a depth direction fall possibility determination unit 1042, an image vertical position determination unit 1043, a lower position fall detection unit 1044, a middle-upper position fall possibility determination unit 1045, and a depth direction fall confirmation unit 1046. The following describes each part that constitutes the fall detection unit 104.

[0023] The omnidirectional fall detection unit 1041 determines, based on the information identified by the posture estimation unit 103, whether or not the worker has fallen in any direction (is in a fallen state) at the work site. Specifically, the omnidirectional fall detection unit 1041 makes its determination using the following three pieces of information. First, as described above, the omnidirectional fall detection unit 1041 determines whether the vertical length of the rectangular area surrounding the entire image portion of the worker in the frame image identified by the posture estimation unit 103 is longer than the horizontal length. This is because if the image portion of the worker is in an upright position, the vertical length of the rectangular area will be longer than the horizontal length.

[0024] Secondly, the all-directional fall determination unit 1041 determines, as described above, whether the order of the body parts in the image portion of the worker in the frame image identified by the posture estimation unit 103 is in the order of shoulder → waist → knee → ankle from top to bottom of the frame image. This is because if the worker is not in an upright position, the order of the body parts in the image portion of the worker will not be in this order.

[0025] Thirdly, the all-around fall detection unit 1041 determines, as described above, whether there is any bias in the length from the shoulder to the waist and the length from the waist to the ankle in the image portion of the worker identified by the posture estimation unit 103. This is because, when standing, the length from the shoulder to the waist and the length from the waist to the ankle are approximately the same, but when squatting or sitting, the length from the shoulder to the waist becomes longer and the length from the waist to the ankle becomes significantly shorter. Also, when falling, there may be a bias between the length from the shoulder to the waist and the length from the waist to the ankle.

[0026] The all-directional fall detection unit 1041 determines that the worker is in an upright position if all of the first to third conditions described above are met, and that the worker is in a fallen position if even one condition is not met. Specifically, it determines that the worker is in a fallen position if the vertical length of the rectangular area surrounding the entire image portion of the worker is shorter than the horizontal length. It also determines that the worker is in a fallen position if the arrangement of the various body parts of the subject is not in the order of shoulder → waist → knee → ankle from top to bottom. Furthermore, it determines that the worker is in a fallen position if there is a bias in the length from the shoulder to the waist and the length from the waist to the ankle. However, even if the all-directional fall detection unit 1041 determines that the worker is in an upright position, as mentioned above, the worker may have fallen in the depth direction, which is the direction away from the surveillance camera 2. In this case, the following parts come into play.

[0027] The depth-direction tipping possibility determination unit 1042 refers to the image storage unit 102 and extracts a frame image from N seconds ago (where N is an integer greater than or equal to 1, for example, 1 second ago or 2 seconds ago), and obtains the skeletal coordinates of the worker's image portion included in that frame image. The acquisition of skeletal coordinates is performed using the method used in the posture estimation unit 103 described above. Based on the skeletal coordinates of the worker's image portion obtained from the frame image from N seconds ago, the depth-direction tipping possibility determination unit 1042 identifies the positions of the shoulders and ankles of the worker's image portion. The depth-direction tipping possibility determination unit 1042 also identifies the positions of the shoulders and ankles of the worker's image portion in the latest frame image obtained by the posture estimation unit 103. The depth-direction tipping possibility determination unit 1042 compares the positions of the shoulders and ankles of the worker's image portion in the frame image from N seconds ago with the positions of the shoulders and ankles of the worker's image portion in the latest frame image.

[0028] As a result of this comparison, the depth-direction fall possibility determination unit 1042 determines that the worker may have fallen in the depth direction if a change occurs in only one of the two directions. For example, if a worker loses consciousness and falls headfirst in the depth direction, the position of the ankles may not change much between the two frame images, while the position of the shoulders will change significantly. Also, for example, if a worker slips and falls in the depth direction, the ankles may move significantly in the depth direction, while the position of the shoulders may not change much between the two frame images. However, in the case of walking in the depth direction, even if the worker is standing, the depth-direction fall possibility determination unit 1042 may determine that there is a possibility of a fall in the depth direction based on the positional relationship between the surveillance camera 2 and the subject (worker).

[0029] Therefore, the image vertical position determination unit 1043 determines, for example, whether the worker's image portion is located in the vertical center or above the center within the latest frame image, based on the position of the worker's image portion identified by the posture estimation unit 103. This is because the appearance of the worker's image portion differs depending on whether it is located in the vertical center or above the center within the frame image or below the center.

[0030] Figure 2 illustrates the difference between images of a worker standing and images of a worker lying down when the worker's image is located in the center or upper part of the captured frame image. Figure 3 illustrates the difference between images of a worker standing and images of a worker lying down when the worker's image is located below the center of the captured image. First, using Figure 2, we will explain the case where the worker's image is located in the vertical center of the frame image.

[0031] Surveillance camera 2 is installed at the work site to capture the entire work area where the worker being monitored (the target for fall detection) is working. In this embodiment, it is assumed to be installed on the ceiling or wall above and to the left of the worker at the work site. Furthermore, various conditions are set for surveillance camera 2 such that the image portion of the standing worker is located in the center of the frame image, and the vertical length of the image portion of the worker occupies 30% or more of the vertical length of the frame image. These conditions include, for example, the distance to the worker, the optical axis direction (orientation of surveillance camera 2), and the zoom magnification of surveillance camera 2. Of course, the worker may move within the work area, but the location should be determined based on where the worker is most frequently located.

[0032] Figure 2(A) shows a worker PS in a standing position and a worker PL in a fallen position. For simplicity of explanation, the ankles of worker PS and worker PL are assumed to be in the same position. In Figure 2(A), the area indicated by the dotted ellipse represents the minimum shooting range within the field of view of surveillance camera 2 that includes both the image portion of worker PS in a standing position and the image portion of worker PL in a fallen position. This range is set at a predetermined distance from surveillance camera 2.

[0033] Furthermore, in Figure 2(A), the diagonal line extending in the direction of installation of the surveillance camera 2 connects the plane containing the lens surface of the surveillance camera 2 and perpendicular to the optical axis to the worker's feet (toes) or the worker's head via the shortest distance. Therefore, the lines connecting the plane containing the lens surface of the surveillance camera 2 and the worker's toes or head are parallel to each other, and the angle θ between these lines and the floor is less than 45 degrees in the case of Figure 2(A). In other words, within the range where the angle θ is less than 45 degrees, the image portion of the worker will be located in the center of the frame image in the vertical direction, or above the center (upper side).

[0034] Therefore, in the example shown in Figure 2(A), that is, when the worker's image is located in the vertical center of the frame image, the vertical length HS of the standing worker PS image is longer than the vertical length HL of the fallen worker PL image. Figure 2(B) is an extracted image of the standing worker PS when the state in Figure 2(A) was photographed, and Figure 2(C) is an extracted image of the fallen worker PL when the state in Figure 2(A) was photographed. As can be seen by comparing Figure 2(B) and Figure 2(C), the vertical length HS of the standing worker PS image is longer than the vertical length HL of the fallen worker PL image. However, since the order of the skeletal coordinates does not change, it is difficult to distinguish between the standing worker PS image and the fallen worker PL image from the standing worker PS image.

[0035] Here, we used Figure 2 to explain the case where the subject (worker) image is located in the center of the frame image in the vertical direction. The vertical length relationship between the image portion of the standing worker PS and the image portion of the fallen worker PL, as explained using Figure 2, also holds true when the worker's image portion is located above the center of the frame image. However, the vertical length relationship between the image portion of the standing worker PS and the image portion of the fallen worker PL, as explained using Figure 2, does not hold true when the worker's image portion is located below the center of the frame image. Specifically, we will use Figure 3 to explain the case where the worker's image portion is located in the lower part of the frame image in the vertical direction.

[0036] The installation location and conditions for surveillance camera 2 are as described above. That is, even when explaining using Figure 3, it is installed in the work site to capture the entire work area where the worker subject to fall detection is working, and is installed on the ceiling or wall above and to the left of the worker in the work site. Furthermore, the distance to the worker and the optical axis direction (orientation of surveillance camera 2) are determined so that the image portion of the standing worker is located in the center of the frame image, and the vertical length of the image portion of the worker occupies 30% or more of the vertical length of the frame image.

[0037] In Figure 3(A), as in Figure 2(A), a standing worker PS and a fallen worker PL are shown, and for the sake of simplicity, the ankles of worker PS and worker PL are assumed to be in the same position. Also in Figure 3(A), the area indicated by the dotted ellipse represents the minimum shooting range within the field of view of the surveillance camera 2 that includes both the image portion of worker PS in the standing position and the image portion of worker PL in the fallen position, and is set at a predetermined distance from the surveillance camera 2.

[0038] Furthermore, in Figure 3(A), as in Figure 2(A), the diagonal line extending in the direction in which the surveillance camera 2 is installed connects the plane containing the lens surface of the surveillance camera 2 and perpendicular to the optical axis to the worker's feet (toes) or the worker's head via the shortest distance. Therefore, the lines connecting the plane containing the lens surface of the surveillance camera 2 and the worker's toes or head are parallel to each other, and the angle θ between these lines and the floor is 45 degrees or more in the case of Figure 3(A). In other words, within the range where the angle θ is 45 degrees or more, the image portion of the worker will be located below (below) the vertical center of the frame image.

[0039] Therefore, in the example shown in Figure 3(A), that is, when the worker's image is located below the center of the frame image in the vertical direction, the vertical length HL of the fallen worker PL image is longer than the vertical length HS of the standing worker PS image. Figure 3(B) is an extracted image of the standing worker PS when the state in Figure 3(A) was photographed, and Figure 3(C) is an extracted image of the fallen worker PL when the state in Figure 3(A) was photographed. As can be seen by comparing Figure 3(B) and Figure 3(C), the vertical length HS of the standing worker PS image is shorter than the vertical length HL of the fallen worker PL image. However, since the order of the skeletal coordinates does not change, it is difficult to distinguish between the standing worker PS image and the fallen worker PL image.

[0040] Furthermore, as can be seen by comparing Figure 2 and Figure 3, the appearance of the image of the worker is reversed depending on whether the worker's image is located in the center or above the center of the frame image, or below the center. Therefore, if it is not possible to determine the position of the worker's image in the vertical direction of the frame image, it will not be possible to properly detect that the worker is in a fallen state when they fall in the depth direction, which is away from the surveillance camera.

[0041] Therefore, the image vertical position determination unit 1043 determines the position of the subject's image within the latest frame image based on the skeletal coordinates of the worker's image portion and the position of the contour (edge) of the worker's image portion, which were identified by the posture estimation unit 103. Specifically, the image vertical position determination unit 1043 determines whether the worker's image portion is located in the vertical center of the frame image, above the center, or below the center.

[0042] The downward position fall detection unit 1044 functions when the vertical position detection unit 1043 determines that the worker's image portion is located in the lower part of the latest frame image (below the center). The downward position fall detection unit 1044 compares the length from the shoulder to the ankle of the worker's image portion in the frame image N seconds ago with the length from the shoulder to the ankle of the worker's image portion in the latest frame image.

[0043] This is because, as explained using Figure 3, when the worker's image is located below the vertical center of the captured image, the vertical length of the image of the worker in a standing position will be shorter than the vertical length of the image of the worker in a fallen position. Therefore, if the length from the shoulder to the ankle of the worker's image in the most recent frame image is longer than the length from the shoulder to the ankle of the worker's image in the frame image N seconds earlier, it can be determined that the worker may be in a fallen state.

[0044] Furthermore, the lower position fall detection unit 1044 identifies the length from the waist to the knee and the length from the knee to the ankle in the image portion of the worker in the most recent frame image, and checks whether there is any bias in the two (whether the ratio of the two is skewed). This is because if the worker is simply squatting, there will be a bias in the length from the waist to the knee and the length from the knee to the ankle. As shown in Figures 3(B) and (C), there is no bias in the length from the waist to the knee and the length from the knee to the ankle in the image portion of the worker between the standing position and the fallen position.

[0045] Furthermore, the length from the shoulder to the ankle of the worker's image portion in the frame image from N seconds ago can be determined using the skeletal coordinates of the worker's image portion identified by the depth-direction fall possibility determination unit 1042 described above. Similarly, the length from the shoulder to the ankle of the worker's image portion in the most recent frame image can be determined using the skeletal coordinates of the worker's image portion identified by the posture estimation unit 103 described above. Likewise, the length from the waist to the knee and the length from the knee to the ankle of the worker's image portion in the most recent frame image can be determined using the skeletal coordinates of the worker's image portion identified by the posture estimation unit 103 described above.

[0046] The lower position fall determination unit 1044 determines that the length from the shoulder to the ankle of the worker's image in the most recent frame image is longer, and that there is no bias between the length from the waist to the knee and the length from the knee to the ankle of the worker's image in the most recent frame image. In this case, the lower position fall determination unit 1044 determines that the worker, who is the subject of the image, is in a fallen state. If either of the above conditions is not met, the lower position fall determination unit 1044 determines that the worker is in an upright state.

[0047] Furthermore, even if the worker moves on foot towards the surveillance camera 2, the relationship shown in Figure 3 between the standing and fallen positions will never be broken. In other words, if the worker's image is located in the lower part of the frame (below the center), the length from the shoulder to the ankle of the worker's image in the most recent frame will never be longer than the length in the standing position. Therefore, the lower position fall determination unit 1044 can determine that the worker, who is the subject of the image, is in a fallen position if the above conditions are met.

[0048] In contrast, the upper-middle position fall possibility determination unit 1045 cannot determine on its own that the worker is in a fallen state, and therefore only determines that there is a possibility that the worker is in a fallen state. The upper-middle position fall possibility determination unit 1045 functions when the image vertical position determination unit 1043 determines that the image portion of the worker is located in the vertical center of the latest frame image or above the center.

[0049] The upper-middle position fall possibility determination unit 1045 compares the length from the shoulder to the ankle of the worker's image portion in the frame image N seconds ago with the length from the shoulder to the ankle of the worker's image portion in the most recent frame image. This is because, as explained using Figure 2, when the worker's image is located in or above the center of the captured image in the vertical direction, the vertical length (up and down direction) of the image of the worker in a standing position will be longer than the vertical length of the image of the worker in a fallen position. Therefore, if the length from the shoulder to the ankle of the worker's image portion in the most recent frame image is shorter than the length from the shoulder to the ankle of the worker's image in the frame image N seconds ago, it can be determined that the worker may be in a fallen state.

[0050] Furthermore, the upper-middle position fall possibility determination unit 1045 identifies the length from the waist to the knee and the length from the knee to the ankle in the image portion of the worker in the most recent frame image, and checks whether there is any bias between the two (whether the ratio between the two is skewed). This is because if the worker is simply squatting, there will be a bias between the length from the waist to the knee and the length from the knee to the ankle. As shown in Figures 2(B) and (C), there is no bias between the length from the waist to the knee and the length from the knee to the ankle in the image portion of the worker in a standing position and a fallen position.

[0051] In this case as well, the length from the shoulder to the ankle of the worker's image portion in the frame image from N seconds ago can be determined using the skeletal coordinates of the worker's image portion identified by the depth-direction fall possibility determination unit 1042 described above. Similarly, the length from the shoulder to the ankle of the worker's image portion in the most recent frame image can be determined using the skeletal coordinates of the worker's image portion identified by the posture estimation unit 103 described above. Likewise, the length from the waist to the knee and the length from the knee to the ankle of the worker's image portion in the most recent frame image can be determined using the skeletal coordinates of the worker's image portion identified by the posture estimation unit 103 described above.

[0052] The upper-middle position fall possibility determination unit 1045 determines that the length from the shoulder to the ankle of the worker's image in the most recent frame image is shorter, and that there is no bias between the length from the waist to the knee and the length from the knee to the ankle of the worker's image in the most recent frame image. In this case, the lower position fall determination unit 1044 determines that the worker, who is the subject of the image, may be in a fallen state. If either of the above conditions is not met, the lower position fall determination unit 1044 determines that the worker is in an upright state.

[0053] As described above, the upper-middle position fall possibility determination unit 1045 cannot determine that the subject, the worker, is in a fallen state, as can be done in the lower position fall determination unit 1044. The upper-middle position fall possibility determination unit 1045 merely determines that the subject, the worker, may be in a fallen state. This is because when a worker moves on foot in the depth direction, away from the surveillance camera 2, they may be in a state similar to that of a fallen state, even if they are standing. In other words, when a worker moves on foot in the depth direction, the length from the shoulder to the ankle of the worker's image in the latest frame image may be shorter, and there may be no bias between the length from the waist to the knee and the length from the knee to the ankle of the worker's image in the latest frame image.

[0054] Therefore, if the upper-middle position fall possibility determination unit 1045 determines that there is a possibility that the worker is in a fallen state, the depth direction fall confirmation unit 1046 functions. The depth direction fall confirmation unit 1046 places the shoulder position of the worker's image portion in the latest frame image onto a circle whose radius is the length from the ankle to the shoulder, with the ankle position of the worker's image portion in the frame image from N seconds ago as the center, and calculates the angle between the floor surface and the radius corresponding to the shoulder position. This angle becomes the predicted angle between the straight line (radius) corresponding to the body of the worker in a fallen state and the floor surface, and is the predicted angle calculated according to the shoulder position of the latest image portion of the worker. This predicted angle is compared with a predetermined threshold (for example, 30 degrees), and if it is smaller than the threshold, it is confirmed that the worker is in a fallen state.

[0055] The calculation of the predicted angle will be explained in detail. Figure 4 is a diagram illustrating the calculation of the angle (predicted angle) between the floor and the extension line from the ankle to the shoulder, which is used to determine whether the worker is standing or lying down. Figure 4(A) shows how the images of the worker in a standing position and lying down position appear in front of the surveillance camera 2 when the worker's image portion is located at or above the center position in the vertical direction of the frame image. In other words, the images of the standing position and lying down position in Figure 4(A) are the same as those shown in Figures 2(B) and (C).

[0056] As shown in Figure 4(A), the length from the ankle AL to the shoulder SL in the image portion of a worker in a fallen position is shorter than the length from the ankle AS to the shoulder SS in the image portion of a worker in a standing position. However, if the position of the ankle AL in the image portion of the worker in a fallen position is aligned with the position of the ankle AS in the image portion of the worker in a standing position, then the position of the shoulder SL in the image portion of the worker in a fallen position and the position of the shoulder SS in the image portion of the worker in a standing position are actually located on the same circumference. Therefore, if the angle (predicted angle) between the line connecting the ankle and shoulder in the image portion of the worker and the floor is known, it is possible to objectively determine whether the worker is in a standing or fallen position according to the predicted angle.

[0057] Therefore, if the above-mentioned upper-middle position fall possibility determination unit 1045 determines that there is a possibility that the worker has fallen, the image portion of the worker in the frame image from N seconds ago and the latest frame image can be identified as being in the following states (1) to (3): (1) The position of the ankle has not changed significantly. (2) The length from the ankle to the shoulder has shortened. (3) The position of the shoulder has dropped. Based on these states (1) to (3), the predicted angle is calculated.

[0058] Figure 4(B) shows the state in which surveillance camera 2 is photographing a worker, viewed from the side (from a direction perpendicular to the optical axis of surveillance camera 2). Therefore, the Y-axis direction is the height direction, and the Z-axis direction is the depth direction. In this case, the Z-axis corresponds to the floor surface. Also, the center (origin) of the circle in Figure 4(B) corresponds to the positions of the ankles AS and AL in the image portion of the worker. Furthermore, in Figure 4(B), the intersection of the Y-axis and the circumference corresponds to the position of the shoulder SS in the image portion of the worker in a standing position. In addition, in Figure 4(B), the position shown on the circumference near the Z-axis corresponds to the position of the shoulder SL in the image portion of the worker in a fallen position.

[0059] Therefore, when the worker's image is in an upright position, the line connecting the ankle AS and shoulder SS is located on the Y-axis, and the angle between the Y-axis and the Z-axis is, needless to say, 90 degrees (right angle). In contrast, when the worker's image is in a fallen position, the position of the shoulder SL moves significantly closer to the Z-axis (floor), so the angle (predicted angle) α between the line connecting the ankle AL and shoulder SL and the Z-axis will be 30 degrees or less in this example. The reason for setting the threshold for the predicted angle to 30 degrees or less is that when a worker falls, there may be errors such as slight movement in the depth direction. For this reason, the range of angles that the line connecting the ankle and shoulder in the worker's image and the Z-axis can take is 0 to 90 degrees.

[0060] Based on the above, here is an example of a method for calculating the predicted angle. As preparation, perform the following steps: (1) With the worker's image positioned in the vertical center of the frame image, determine the length LS from the ankle AS to the shoulder SS for the standing worker's image and the length LL from the ankle AL to the shoulder SL for the fallen worker's image. (2) Divide the length LL when the worker is fallen by the length LS when the worker is standing to find the ratio RT of the length LL when the worker is fallen to the length LS when the worker is standing.

[0061] (3) By subtracting the ratio RT from 1 and multiplying by 100, the ratio RG of the change range in the length from the ankle to the shoulder when the worker's image changes from a standing position to a fallen position can be obtained. This change range ratio RG corresponds to a range from 0 to 90 degrees and is information that indicates the change range in the worker's image from a standing position to a fallen position. Therefore, by dividing 90 degrees by the change range RG, the unit change angle UA in the change range RG can be obtained. For this reason, the depth direction fall determination unit 1046 records and stores the change range ratio RG and the unit change angle UA, which are information that indicates the change range, in its own memory or the non-volatile memory of the control unit 110, and makes them available at any time.

[0062] Next, when actually determining the state of the fall, the depth-direction fall determination unit 1046 performs the following processing. First, the depth-direction fall determination unit 1046 determines the length LS1 from the ankle AS1 to the shoulder SS1 of the worker's image portion in the frame image N seconds ago, and the length LL1 from the ankle AL1 to the shoulder SL1 of the worker's image portion in the latest frame image. Next, it divides the length LL1 in the worker's image portion of the latest frame image by the length LS1 in the worker's image portion of the latest frame image to obtain the ratio RT1. Then, by subtracting the ratio RT1 from 1 and multiplying by 100, the ratio RG1 of the change range, which indicates how much the worker's image portion in the latest frame has changed from the worker's image portion in the frame image N seconds ago, can be obtained.

[0063] Therefore, by subtracting the newly calculated change range ratio RG1 from the already calculated and stored change range ratio RG, we can determine the changeable amount CH, which indicates how much the worker's image portion has changed (how much it has fallen) within that change range, and how much more change is possible, i.e., how much more it can fall in the Z-axis direction. By multiplying this changeable amount CH by the unit change angle UA, we can calculate the changeable angle, i.e., the predicted angle α. If this predicted angle α is 30 degrees or less, the worker can be identified as being in a fallen state.

[0064] The following explains a specific example of the method for calculating the predicted angle α described above. Here, (1) under the condition that the worker's image portion is located in the vertical center of the frame image, the length LS from the ankle AS to the shoulder SS of the standing worker's image portion is a "value of 10", and the length LL from the ankle AL to the shoulder SL of the fallen worker's image portion is a "value of 7". These values ​​can be determined according to the skeletal coordinates of the worker's image portion in the frame image, as described above. In this case, the coordinates may be set by setting a coordinate system in which multiple pixel units are one scale unit vertically and horizontally for the frame image, or the vertical and horizontal pixel positions may be used as coordinates. Of course, other coordinates may also be set. (2) The length LL "value of 7" when the worker is in a fallen state is divided by the length LS "value of 10" when the worker is in a standing state to find the ratio RT of the length LL when the worker is in a fallen state to the length LS when the worker is in a standing state. In this example, the ratio RT = 7 / 10 = "value of 0.7".

[0065] (3) Subtracting the ratio RT "value 0.7" from 1 and multiplying by 100 gives the ratio RG "30%" of the range of change in the length from ankle to shoulder when the worker's image changes from a standing position to a fallen position. This ratio RG "value 30%" corresponds to a range from 0 to 90 degrees and represents information indicating the range of change in the worker's image from a standing position to a fallen position. Therefore, dividing 90 degrees by the range of change RG "value 30%" gives the unit change angle UA "value 3 degrees" in the range of change RG. For this reason, the depth direction fall determination unit 1046 records and stores the ratio RG "value 30%" and the unit change angle UA "value 3 degrees", which represent information indicating the range of change, in a predetermined memory and makes them available at any time.

[0066] Next, when actually determining the fall state, the depth-direction fall determination unit 1046 performs the following processing. First, the depth-direction fall determination unit 1046 determines the length LS1 from the ankle AS1 to the shoulder SS1 of the worker's image portion in the frame image from N seconds ago, and the length LL1 from the ankle AL1 to the shoulder SL1 of the worker's image portion in the latest frame image. In this example, assume that the length LS1 from the ankle AS1 to the shoulder SS1 of the worker's image portion in the frame image from N seconds ago is "value 8", and the length LL1 from the ankle AL1 to the shoulder SL1 of the worker's image portion in the latest frame image is "value 6".

[0067] Next, the length LL1 (value 6) of the worker's image portion in the latest frame image is divided by the length LS1 (value 8) of the worker's image portion in the latest frame image to obtain the ratio RT1. In this case, the ratio RT1 = 6 / 8 = 0.75. Next, subtracting the ratio RT1 (value 0.75) from 1 and multiplying by 100 gives the ratio RG1 of the change range, which indicates how much the worker's image portion in the latest frame has changed (how much it has fallen) from the worker's image portion in the frame image N seconds ago. In this example, the ratio RG1 of the change range = (1 - 0.75) × 100 = 25% can be calculated.

[0068] Therefore, by subtracting the newly calculated change range percentage RG1 ("25%") from the already calculated and stored change range percentage RG ("value 30%"), we can determine the changeable amount CH, which indicates how much the worker's image changes (how much it falls) within that change range, and how much change is possible, i.e., how much it can fall in the Z-axis direction. In this example, the changeable amount CH = 30% - 25% = 5%. Therefore, by multiplying the stored unit change angle UA ("value 3 degrees") and the changeable amount CH ("value 5%"), we can determine the angle that indicates how much it can fall in the Z-axis direction, i.e., the predicted angle α = 15 degrees. In this case, since the predicted angle is 30 degrees or less, it can be determined that the person is in a fallen state.

[0069] In this way, the depth-direction fall determination unit 1046 calculates a predicted angle α, which is the angle between the straight line connecting the ankle and shoulder of the worker's image in the latest frame image and the floor. The depth-direction fall determination unit 1046 identifies the worker as being in a fallen state if the calculated predicted angle α is less than 30 degrees, and identifies the worker as being in an upright state if the predicted angle α is 30 degrees or more. In this way, if the fall determination unit 104 determines that the worker is in a fallen state, the determination result output unit 105 functions under the control of the control unit 110, and displays a notification that the worker is in a fallen state through the monitor device 3. In this example, only a display output is provided, but if the monitor device 3 is equipped with a speaker, or if a separate speaker is connected, the worker may be notified by voice that they have fallen.

[0070] Furthermore, the threshold for comparison with the predicted angle α is not always 30 degrees. It should be set according to the work site, taking into account the installation conditions of the surveillance camera 2 (angle of the optical axis and distance to the worker), as well as the position of the shoulders in the image portion of a standing worker and the position of the shoulders in the image portion of a fallen worker. In this example, we used the frame image from N seconds ago and the latest frame image to determine whether or not the worker is in a fallen state. However, it is also possible to perform multiple determination processes, such as determining whether or not the worker is in a fallen state using the frame image from 1 second ago and the latest frame image, and then determining whether or not the worker is in a fallen state using the frame image from 2 seconds ago and the latest frame image.

[0071] [Summary of processing performed by the image processing device 1] Figures 5 and 6 are flowcharts illustrating the processes performed by the image processing apparatus of the embodiment. The processes shown in the flowcharts of Figures 5 and 6 are mainly performed by the posture estimation unit 103 and the fall detection unit 104 under the control of the control unit 110. The control unit 110 executes the processes shown in the flowcharts of Figures 5 and 6 at predetermined timings, such as for each frame image or for multiple frames of images.

[0072] When the processes shown in the flowcharts of Figures 5 and 6 are executed by the control unit 110, the posture estimation unit 103 functions under the control of the control unit 110. The posture estimation unit 103 reads the latest frame image of the video data from the surveillance camera 2 stored in the video storage unit 102 and obtains the skeletal coordinates of each part in the image portion of the person (worker) (step S101). Next, under the control of the control unit 110, the posture estimation unit 103 functions to identify a rectangular region surrounding the entire image portion of the worker in the read latest frame image and check whether the vertical length of the rectangular region is longer than the horizontal length (step S102). The result of the check in step S102 is notified to the omnidirectional fall determination unit 1041 of the fall determination unit 104.

[0073] Furthermore, the posture estimation unit 103 determines the positional relationship of each part in the image portion of the person (worker) based on the skeletal coordinates of each part of the worker's image portion identified in step S101, and checks whether they are arranged in the order of shoulder, waist, knee, and ankle from top to bottom (step S103). The results of the check in step S103 are notified to the omnidirectional fall detection unit 1041 of the fall detection unit 104. In addition, the posture estimation unit 103 determines the length from the shoulder to the waist and the length from the waist to the ankle based on the skeletal coordinates of each part of the image portion of the worker's image portion identified in step S101, and checks whether there is any bias (step S104). The results of the check in step S104 are notified to the omnidirectional fall detection unit 1041 of the fall detection unit 104.

[0074] After this, the all-directional fall detection unit 1041 determines whether all the results of steps S102 to S104 are satisfied (step S105) and notifies the control unit 110 of the determination result. The all-directional fall detection unit 1041 determines whether all of the following (1) to (3) are satisfied: (1) The vertical length of the rectangular area surrounding the entire image portion of the worker is longer than the horizontal length. (2) The parts of the image portion of the worker are arranged from top to bottom in the order of shoulder, waist, knee, and ankle. (3) There is no bias in the length from the shoulder to the waist and the length from the waist to the ankle of the image portion of the worker.

[0075] If the all-directional fall detection unit 1041 determines that any of the above conditions (1) to (3) are not satisfied, it determines that the worker being judged at the work site is in a fall state (step S106). Upon receiving notification of the determination result, the control unit 110 outputs the determination result to the monitoring device 3 through the determination result output unit 105, and the monitor device 3 notifies the supervisor and manager of the work site of the worker's fall via its display output (step S107). After this, the control unit 110 completes the process shown in Figures 5 and 6 and waits for the next execution timing.

[0076] In the determination process of step S105, it is determined that all of the above conditions (1) to (3) are satisfied. In this case, there is still a possibility that the worker is falling in the depth direction, which is away from the surveillance camera 2. For this reason, the depth direction fall possibility determination unit 1042 functions under the control of the control unit 110. The depth direction fall possibility determination unit 1042 reads the frame image from the video storage unit 102 from N seconds ago and obtains the skeletal coordinates of the image portion of the person (worker) using the same method as the posture estimation unit 103 in step S101 (step S108).

[0077] Next, the depth-direction fall possibility determination unit 1042 compares the positions of the shoulders and ankles of the worker's image portion N seconds ago with the positions of the shoulders and ankles of the worker's image portion in the latest frame image acquired in step S101, according to the skeletal coordinates of the worker's image portion (step S109). From the comparison result in step S109, the depth-direction fall possibility determination unit 1042 determines whether only one of the positions of the shoulders and ankles of the worker's image portion has changed between N seconds ago and the latest frame image (step S110), and notifies the control unit 110 of the determination result. If the determination process in step S110 determines that only one of them has not changed, the depth-direction fall possibility determination unit 1042, under the control of the control unit 110, determines that the worker is in an upright position (step S111). After this, the control unit 110 terminates the process shown in Figures 5 and 6 and waits for the next execution timing.

[0078] In the determination process of step S110, if it is determined that only one of the two has changed, the possibility that the worker has fallen in the depth direction cannot be ruled out, so the process proceeds to step S112 in Figure 6. In this case, under the control of the control unit 110, the image vertical position determination unit 1043 functions and identifies the position of the person (image portion of the worker) within the latest frame image read in step S101 (step S112). Based on the identification result of step S112, the image vertical position determination unit 1043 determines whether the position of the person (image portion of the worker) is located in the central part or above the central part in the vertical direction within the frame image (step S113). The determination result is notified to the control unit 110.

[0079] In the determination process of step S113, it is determined that the position of the worker's image portion is not located in or above the vertical center of the frame image, that is, it is located below the vertical center of the frame image. In this case, the downward position fall determination unit 1044 functions under the control of the control unit 110. The downward position fall determination unit 1044 compares the length from the worker's shoulder to the ankle, according to the skeletal coordinates of the worker's image portion in the frame image from N seconds ago, with the length from the worker's shoulder to the ankle, according to the skeletal coordinates of the worker's image portion in the most recent frame image (step S114).

[0080] Furthermore, the downward position fall detection unit 1044 identifies the length from the waist to the knee and the length from the knee to the ankle of the worker's image portion in the latest frame image, according to the skeletal coordinates of the worker's image portion in the latest frame image (step S115). Note that in step S114, the skeletal coordinates of the worker's image portion in the frame image N seconds ago can be those obtained in step S108, and the skeletal coordinates of the worker's image portion in the latest frame image can be those obtained in step S101. Similarly, in step S115, the skeletal coordinates of the worker's image portion in the latest frame image can be those obtained in step S101.

[0081] After this, the lower position fall determination unit 1044 performs a determination process based on the comparison result in step S114 and the identification result in step S115 (step S116), and notifies the control unit 110 of the determination result. In step S116, it is determined whether the following two conditions are satisfied: that the length from the shoulder to the ankle of the worker's image portion in the latest frame image is longer than the length from the shoulder to the ankle of the worker's image portion in the frame image N seconds ago; and that the ratio of the lower body portion of the worker's image portion in the latest frame image is appropriate. The determination is made whether both of these conditions are satisfied. If the determination process in step S116 determines that neither condition is satisfied, the lower position fall determination unit 1044 determines that the worker is in a standing position (not in a fallen position) (step S117). After this, the control unit 110 terminates the process shown in Figures 5 and 6 and waits for the next execution timing.

[0082] On the other hand, if the determination process in step S116 determines that both conditions are met, the lower position fall determination unit 1044 determines that the worker is in a fallen state (step S118). The control unit 110 outputs the determination result to the monitoring device 3 through the determination result output unit 105, and the monitor device 3 notifies the supervisor and manager of the work site of the worker's fall through its display output (step S119). After this, the control unit 110 completes the process shown in Figures 5 and 6 and waits for the next execution timing.

[0083] In the determination process of step S113, it is determined that the position of the worker's image portion is located in the central part of the frame image in the vertical direction or above the central part. In this case, the upper-middle position fall possibility determination unit 1045 functions under the control of the control unit 110. The upper-middle position fall possibility determination unit 1045 compares the length from the worker's shoulder to the ankle, according to the skeletal coordinates of the worker's image portion in the frame image N seconds ago, with the length from the worker's shoulder to the ankle, according to the skeletal coordinates of the worker's image portion in the most recent frame image (step S120). The process in step S120 is the same as the process in step S114 described above.

[0084] Furthermore, the upper-middle position fall possibility determination unit 1045 identifies the length from the waist to the knee and the length from the knee to the ankle of the worker's image portion in the latest frame image, according to the skeletal coordinates of the worker's image portion (step S121). The processing in step S121 is the same as the processing in step S115 described above. After this, the upper-middle position fall possibility determination unit 1045 performs a determination process based on the comparison result in step S120 and the identification result in step S121 (step S122), and notifies the control unit 110 of the determination result. In step S116, it is determined whether the following two conditions are satisfied: that the length from the shoulder to the ankle of the worker's image portion in the frame image N seconds ago is longer than the length from the shoulder to the ankle of the worker's image portion in the latest frame image; and that the ratio of the lower body portion of the worker's image portion in the latest frame image is appropriate.

[0085] In the determination process of step S122, if it is determined that both conditions are met, the depth-direction tilt determination unit 1046 functions under the control of the control unit 110. As explained using Figure 4, the depth-direction tilt determination unit 1046 performs a process to calculate the angle between the worker's image and the floor surface (predicted tilt angle α in the depth direction of the worker's image) (step S123). After this, the depth-direction tilt determination unit 1046 determines whether the calculated predicted angle α is smaller than a threshold (step S124) and notifies the control unit 110 of the determination result.

[0086] In the determination process of step S124, if it is determined that the predicted angle α is smaller than the threshold (the predicted angle α is less than the threshold), the depth-direction tipping confirmation unit 1046 determines (confirms) that the worker is in a tipping state (step S126) and notifies the control unit 110 of the determination result. Upon receiving this notification, the control unit 110 outputs the determination result to the monitoring device 3 through the determination result output unit 105, and the monitor device 3 notifies the supervisor and manager of the work site of the worker's tipping (step S127). After this, the control unit 110 completes the process shown in Figures 5 and 6 and waits for the next execution timing.

[0087] Furthermore, if the determination process in step S122 determines that neither condition is met, or if the calculated predicted angle is greater than or equal to a threshold in step S124, the depth-direction tipping determination unit 1046 determines that the worker is in a standing position (step S125). After this, the control unit 110 completes the processes shown in Figures 5 and 6 and waits for the next execution timing.

[0088] [Effects of the embodiment] The image processing device 1 in this embodiment can use video data from a surveillance camera 2 to quickly and appropriately determine if a worker at a work site has fallen, and notify the supervisor or manager of the work site. In particular, it can quickly and appropriately determine if a worker has fallen in the depth direction, which is away from the surveillance camera 2. In other words, it can prevent situations where a fall is not detected or is detected late. This prevents inconveniences such as delays in detecting a worker's fall and inability to take appropriate action quickly.

[0089] Furthermore, high-precision surveillance cameras are not required; a so-called monocular camera that functions with a single camera lens can be used as surveillance camera 2, making it easy to implement and reducing implementation costs.

[0090] [Differentiation] Figure 7 is a diagram illustrating the timing of the execution of the process for detecting a fall. As described above, the worker fall detection process in the image processing device 1 (Figures 5 and 6) is performed for each frame image or for multiple frames images. For example, as shown in Figure 7(A), if the worker fall detection process is performed every three frames, it is possible to set it so that a fall is not detected unless it is detected as a fall five times in a row. This makes it possible to avoid detecting minor falls, such as tripping, where the person can quickly return to an upright position, as a fall.

[0091] However, as shown in Figure 7(B), in the case of the image processing device 1 of the above embodiment, when determining whether a worker has fallen in the depth direction, past frame images from N seconds prior are also used. For this reason, when determining whether a worker has fallen in the depth direction, for example, suppose it is determined that the worker is in a fallen state using a frame image from 1 second ago. In this case, it is also possible to perform fall detection using frame images from 2 seconds and 3 seconds ago, and if it is determined that the worker is in a fallen state in each of those cases, then it is determined that the worker is in a fallen state.

[0092] The method for calculating the predicted angle in the embodiment described above is just one example, and various other calculation methods can be used.

[0093] Furthermore, in the embodiment described above, the calculation of the predicted angle α for detecting the tipping state in the depth direction was explained as being performed using the frame image from N seconds ago and the latest frame image. However, this is not the only way. For example, the predicted angle may be calculated first using the frame image from 1 second ago and the latest frame image, and if the result is an upright state, the predicted angle may be calculated using the frame image from 2 seconds ago and the latest frame image, and the determination of whether it is an upright or tipping state may be made again. In this case, the time to go back in time should be predetermined, for example, up to 3 seconds ago.

[0094] Furthermore, while the image processing apparatus of the above-described embodiment was applied to monitoring workers in a chemical plant, it is not limited to this. It can be applied to various situations where it is necessary to quickly and accurately detect when a worker falls.

[0095] [others] As can be seen from the description of the above embodiment, the function of the pose estimation means of the claim is realized by the pose estimation unit 103 of the image processing device 1 of the embodiment, and the function of the omnidirectional tipping determination means of the claim is realized by the omnidirectional tipping determination unit 1041 of the image processing device 1. Furthermore, the function of the depth direction tipping possibility determination means of the claim is realized by the depth direction tipping possibility determination unit 1042 of the image processing device 1. Furthermore, the function of the depth direction tipping determination processing means of the claim is realized by the image vertical position determination unit 1043, the lower position tipping determination unit 1044, the middle upper position tipping possibility determination unit 1045, and the depth direction tipping confirmation unit 1046 of the image processing device 1.

[0096] Furthermore, the function of the vertical position determination means of the claim is realized by the vertical position determination unit 1043 of the image processing device 1, and the function of the downward position tilt determination means of the claim is realized by the downward position tilt determination unit 1044 of the image processing device 1. Furthermore, the function of the upper-middle position tilt possibility determination means of the claim is realized by the upper-middle position tilt possibility determination unit 1045 of the image processing device 1, and the function of the depth direction tilt determination means of the claim is realized by the depth direction tilt determination unit 1046 of the image processing device 1.

[0097] Furthermore, the program that executes the processes shown in the flowcharts of Figures 5 and 6 of the above-described embodiment is an example to which one embodiment of the image processing program according to this invention is applied. Therefore, the image processing apparatus 1 of this invention can be realized by installing and running the image processing program that executes the processes shown in Figures 5 and 6 on a personal computer. In other words, the functions of the five determination units and one confirmation unit that constitute the posture estimation unit 103 and the fall determination unit 104 shown in Figure 1 can be realized by software executed in the control unit 110. [Explanation of Symbols]

[0098] 1…Image processing device, 101T…Input terminal, 101…Video input unit, 102…Video storage unit, 103…Posture estimation unit, 104…Tipping detection unit, 1041…Omnidirectional tipping detection unit, 1042…Depth direction tipping possibility detection unit, 1043…Image vertical position detection unit, 1044…Lower position tipping detection unit, 1045…Mid-upper position tipping possibility detection unit, 1046…Depth direction tipping confirmation unit, 105…Determination result output unit, 105T…Output terminal, 2…Surveillance camera, 3…Monitor device

Claims

1. An image processing device that determines whether or not a subject is in a fallen state using frame images that constitute video data obtained by continuously photographing a predetermined person as the subject, A posture estimation means that obtains the skeletal coordinates of the subject from the latest frame image, identifies information about the subject, and enables the estimation of the subject's posture, An all-directional fall determination means that determines whether or not the subject is in a fallen state based on the information identified by the posture estimation means, In the omnidirectional tipping determination means, if it is determined that the subject is not in a tipping state, a depth-direction tipping possibility determination means determines the possibility of tipping in the depth direction using the skeletal coordinates of the subject obtained from the frame image N seconds ago and the skeletal coordinates of the subject from the latest frame image obtained by the posture estimation means, The depth-direction tipping possibility determination means, when it is determined that there is a possibility of tipping in the depth direction, performs a depth-direction tipping determination processing means to determine whether or not the subject is in a state of tipping in the depth direction. An image processing apparatus characterized by comprising:

2. An image processing apparatus according to claim 1, The depth-direction tipping determination processing means is A vertical position determination means for determining the vertical position of the subject in the latest frame image, The downward position fall determination means determines that the subject is in a fallen state if the vertical position determination means determines that the subject is located at the bottom of the latest frame image, and the length from the shoulder to the ankle in the frame image N seconds ago is shorter than the length from the shoulder to the ankle in the latest frame image, and there is no bias in the ratio of the length from the waist to the knee to the length from the knee to the ankle in the latest frame image. An image processing apparatus characterized by comprising:

3. An image processing apparatus according to claim 1, The depth-direction tipping determination processing means is A vertical position determination means for determining the vertical position of the subject in the latest frame image, If the vertical position determination means determines that the subject is located in the center or upper part of the latest frame image, and the length from the shoulder to the ankle in the frame image N seconds ago is longer than the length from the shoulder to the ankle in the latest frame, and there is no bias in the ratio of the length from the waist to the knee to the length from the knee to the ankle in the latest frame image, then the upper-middle position fall possibility determination means determines that the subject may be in a fallen state. The above-mentioned upper-middle position tipping possibility determination means is a means that functions when it is determined that the subject may be in a tipping state, and calculates a predicted tilt angle in the depth direction, and if the calculated predicted tilt angle is smaller than a threshold, it is determined that the subject is in a tipping state, and is a depth direction tipping determination means An image processing apparatus characterized by comprising:

4. An image processing apparatus according to any one of claim 1, claim 2, or claim 3, The posture estimation means obtains the minimum X coordinate, minimum Y coordinate, maximum X coordinate, and maximum Y coordinate of each part from the acquired skeletal coordinates, and identifies a rectangle that encloses the entire subject. The aforementioned all-directional tipping detection means is: The subject is determined to be in a fallen state if one or more of the following conditions are not met: the vertical length of the rectangle is longer than the horizontal length; the obtained skeletal coordinates are arranged in the order of shoulder, waist, knee, and ankle from top to bottom in the latest frame image; and there is no bias in the length from shoulder to waist and the length from waist to ankle, which are determined based on the obtained skeletal coordinates. An image processing apparatus characterized by the following:

5. An image processing program performed by a computer that performs image processing to determine whether or not a subject is in a fallen state, using frame images that constitute video data obtained by continuously photographing a predetermined person as the subject, A posture estimation step which involves obtaining the skeletal coordinates of the subject from the latest frame image, identifying information about the subject, and enabling the estimation of the subject's posture, A fall detection step that determines whether the subject is in a fallen state based on the information identified in the posture estimation step, In the all-directional tipping determination step, if it is determined that the subject is not in a tipping state, a depth-direction tipping possibility determination step is performed to determine the possibility of tipping in the depth direction, using the skeletal coordinates of the subject obtained from the frame image N seconds prior and the skeletal coordinates of the subject from the latest frame image obtained in the posture estimation step. In the depth-direction tipping possibility determination step, if it is determined that there is a possibility of tipping in the depth direction, a depth-direction tipping determination process is performed to determine whether or not the subject is in a state of tipping in the depth direction. An image processing program characterized by performing the following.

Citation Information

Patent Citations

  • Methods and apparatus for fall prevention and detection

    JP2006522959A

  • Person detection device, method, and program

    JP2021033343A

  • Casualty detection device, casualty detection system, and casualty detection method

    JP2022016979A

  • Fall prevention system

    JP7091007B2

  • Method and device for fall prevention and detection

    US20060145874A1