Image processing device and computer program
The image processing apparatus enhances detection accuracy by using position-based reliability determination to correct for misclassifications in image areas where the camera-target relationship affects state detection, particularly for behaviors like sitting or angry postures.
Patent Information
- Application Number
- JP2023223284
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
The detection accuracy of a detection target's state in images can decrease due to the relative positional relationship between the camera and the detection target, leading to misclassification of behaviors such as angry or sitting postures.
An image processing apparatus that includes an image acquisition unit, a state detection unit, an image position detection unit, and an output determination unit to determine the reliability of detected states based on the image position, adjusting confidence thresholds and determination criteria according to the image area and camera installation conditions.
The apparatus suppresses decreases in detection accuracy by accurately determining the reliability of detected states, particularly in specific image areas where misclassification is likely, thereby improving overall detection precision.
Smart Images

Figure 2025105026000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image processing apparatus and a computer program.
Background Art
[0002] Conventionally, techniques for detecting the state of a detection target shown in an image have been proposed. For example, Patent Document 1 below discloses a system for detecting states such as the posture and behavior of a person from image data obtained by photographing a person who is the detection target.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, depending on the relative positional relationship between the camera that photographs the detection target and the detection target, there is a risk that the detection accuracy of the state of the detection target may decrease. For example, it may fail to detect an angry behavior (such as a destructive behavior or a fight) that is an angry behavior of the detection target person, or it may erroneously detect a sitting behavior (such as bending, crouching, or kneeling on the ground) even though the person is standing upright. An object of the present invention is to suppress a decrease in detection accuracy due to the relative positional relationship between a camera that photographs an object and a detection target when detecting the state of the detection target shown in an image.
Means for Solving the Problems
[0005] An image processing apparatus according to an aspect of the present invention includes an image acquisition unit that acquires an input image captured by a camera, a state detection unit that detects a state type of a moving object shown in the input image, an image position detection unit that detects an image position that is the position of the moving object on the input image, and a determination unit that determines the reliability of a detection result of a predetermined state by the state detection unit based on the image position detected by the image position detection unit.
Effect of the Invention
[0006] According to the present invention, when detecting the state of a detection target shown in an image, it is possible to suppress a decrease in detection accuracy due to the relative positional relationship between the camera that captures the target object and the detection target.
Brief Description of the Drawings
[0007]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Embodiments for Carrying Out the Invention
[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments of the present invention shown below illustrate apparatuses and methods for embodying the technical idea of the present invention, and the technical idea of the present invention does not specify the structure, arrangement, etc. of the components as follows. The technical idea of the present invention can be variously modified within the technical scope defined by the claims described in the claims.
[0009] FIG. 1 is a schematic configuration diagram of an example of an image processing apparatus according to an embodiment. The image processing apparatus 10 detects the type of state of a detection target in the monitoring space 1. In this specification, the case of detecting the type of state of a person 2 as a detection target existing in the monitoring space 1 will be exemplified.
[0010] The image processing apparatus 10 detects the posture and behavior of the person 2 as the type of state of the person 2. The types of states of the detection target include those detected by the posture and behavior of the person 2 alone and those detected by the postures and behaviors of a plurality of persons. The image processing apparatus 10 detects an abnormal posture and abnormal behavior of the person 2 as an abnormal state, which is a state of the person 2 notifying the monitor etc. of the monitoring space 1 that an abnormal situation has occurred. The postures and behaviors of the detection target as the abnormal state are, for example, "intimidation (a person waving their arms etc.)", "fight (a person punching etc. another person)", "destruction (a person swinging their arms etc. downwards)", "fall (a person whose standing posture has changed to a collapsed and lying sideways posture)", "hold-up (a person continuously raising both hands due to being threatened etc.)", "thrust (a person stretching their arm towards another person (for example, a person thrusting a weapon etc. at another person))", "bending (a person continuously crouching with bent knees or waist)", "kowtow (a person prostrating themselves to show respect to another person)", "crawling (a person lying on their stomach)", "lying down (a person lying horizontally and continuing to sleep)", etc. Further, states other than the abnormal state may be detected. For example, postures such as "standing position", "standing upright", "sitting position", etc. and behaviors such as "walking", "running", etc. may be detected. Note that the image processing apparatus 10 may detect the state types of detection targets other than the person 2. For example, it may detect the state types of a machine having a manipulator or a non-humanoid robot as detection targets. For example, the image processing apparatus 10 may detect postures such as the orientation and flexion state of the manipulator as state types, and may also detect operations that are changes in these postures as state types.
[0011] The image processing apparatus 10 includes a camera 11, an input unit 12, a storage unit 13, a control unit 14, and an output unit 15. Among these, the storage unit 13 and the control unit 14 may be realized by a so-called computer, and the input unit 12 and the output unit 15 may be realized as peripheral devices of the computer. The camera 11 is disposed in the monitoring space 1 to generate an image of an object existing in the monitoring space 1. The camera 11 may be, for example, a single-direction type camera with a horizontal viewing angle and a vertical viewing angle of about 90 degrees. The camera 11 may also be an omnidirectional camera that takes the entire 360 degrees as the imaging area (monitoring area). Note that the image processing apparatus 10 may not include the camera 11 and may acquire video from an external imaging device.
[0012] The input unit 12 includes user interfaces such as a keyboard and a mouse that are operated by a user and used for data input and the like. The input unit 12 is connected to the control unit 14, converts the user's operation into an operation signal, and outputs it to the control unit 14. Further, the input unit 12 may include a DVD (Digital Versatile Disc) drive and a USB (Universal Serial Bus) interface. The input unit 12 inputs data to the control unit 14 as a file and outputs data from the control unit 14 as a file. The storage unit 13 is a memory device such as a ROM (Read Only Memory) and a RAM (Random Access Memory), and stores various programs and various data. The storage unit 13 is connected to the control unit 14 and inputs and outputs this information to and from the control unit 14.
[0013] The control unit 14 is composed of arithmetic units such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), and an MCU (Micro Control Unit). The control unit 14 is connected to the storage unit 13, reads and executes a computer program from the storage unit 13, operates as various processing units, stores various data in the storage unit 13, and reads it out. The functions of the image processing apparatus 10 described below are realized by the control unit 14 executing the computer program stored in the storage unit 13.
[0014] The output unit 15 is connected to the control unit 14 and outputs the detection result of the state type by the control unit 14. The output unit 15 may include a display device such as a liquid crystal display or a CRT (Cathode Ray Tube) display. The output unit 15 may be provided with an audio signal output device such as a speaker or a buzzer. The output unit 15 may be provided with a network interface or the like that transmits and receives data between the image processing apparatus 10 and an external device by wired communication or wireless communication.
[0015] The control unit 14 acquires a captured image generated by the camera 11 capturing a person 2 existing in the monitoring space 1, and detects the state type of the person 2 shown in the captured image. At this time, the position on the captured image where the person 2 appears in the captured image (hereinafter referred to as "image position") is determined according to the relative positional relationship between the camera 11 and the person. Depending on the image position of the person 2, there is a risk that the detection accuracy of the state type of the person 2 may decrease. Now, as shown in FIG. 1, assume a case where the person 2 is located near the area directly below the camera 11. FIG. 2 is a schematic diagram of a captured image of the person 2 captured by the camera 11.
[0016] In the case of the captured image 3 in FIG. 2, a person 2 is shown in an image area (which may be referred to as the "directly below image area" in the following description) near the area directly below the camera 11. When the camera 11 is a single-direction type camera, if the range of the captured image 3 in which a subject above the optical axis of the camera 11 is shown is defined as the upper side and the range in which a subject below the optical axis is shown is defined as the lower side, the directly below image area is an area near the lower end of the captured image 3.
[0017] When the person 2 is shown in the directly below image area, the lower body is hidden by the person 2's own body, making it easy to misdetect the posture of the person 2. Also, since the movement of each part of the person 2 is difficult to detect, the behavior of the person 2 is also likely to be misdetected. Therefore, the image processing apparatus 10 of the embodiment detects the state type of the person 2 shown in the captured image 3, detects the image position which is the position of the person 2 on the captured image 3, and determines the reliability of the detected state type based on the detected state type and the image position.
[0018] FIG. 3 is a block diagram showing an example of the functional configuration of the image processing apparatus 10 of the embodiment. The image processing apparatus 10 includes the above-described camera 11 and output unit 15, an image acquisition unit 20, a state detection unit 21, an image position detection unit 22, and an output determination unit 23. The image acquisition unit 20 acquires, as an input image, an image generated by the camera 11 photographing an object existing in the monitoring space 1.
[0019] The state detection unit 21 detects the state type of the person 2 shown in the input image acquired by the image acquisition unit 20. For example, the state detection unit 21 may detect the type of the behavior of the person 2 as the state type. In this case, the state detection unit 21 first detects the posture of the person 2 shown in the input image. For example, the state detection unit 21 may input the input image into a learning model (which may be hereinafter referred to as "posture detection AI (Artificial Intelligence)") generated by machine learning that takes image data as input and outputs the posture of the person shown in the image data in the image data, and estimate the posture of the person 2. For example, the posture detection AI may identify the positions of the joint points of the person 2 shown in the input image and estimate the posture of the person 2 based on the feature amounts of the positions of the joint points.
[0020] Also, for example, the state detection unit 21 may detect the posture of the person 2 by means other than the learning model (for example, rule-based, etc.). For example, the state detection unit 21 may extract the silhouette image of the person 2 by the background difference method, and obtain a score from the similarity between the extracted silhouette image and the silhouette images of each posture registered in advance to detect the posture, or extract the positions of the joints of the person from the silhouette of the person extracted by the background difference method, and obtain a score from the similarity between the positions of the joints of each posture registered in advance and the positions of the joints of the extracted person to detect the posture. Next, the state detection unit 21 determines the action of the person 2 based on the time-series change information of the detected posture. For example, the state detection unit 21 may determine the action of the person 2 based on a rule-based approach based on the time-series change information of the detected posture. For example, when the posture of the person 2 changes in the order of standing upright, falling down, and lying down, it may be determined that the type of action of the person 2 is a falling action. The state detection unit 21 may output the "confidence level", which is the correct probability of the detected posture. For example, the type of the detected action and its confidence level may be output, such as the probability of a falling action being 80%.
[0021] Also, for example, for each frame of the input image, the state detection unit 21 determines the action based on the time-series change information of the posture between the target frame and the previous frame, and may detect the type of action of the person 2 by a majority vote of the postures detected in each frame for the same person existing in the images of multiple frames. Also, the state detection unit 21 may detect the type of action of the person 2 using a learning model generated by machine learning that takes video or time-series information of joints as input and outputs data on the action of the person.
[0022] The state detection unit 21 may output not only the type of a single state, but also the types of multiple states and their confidence levels (for example, "standing upright" 90%, "walking" 80%, etc.). For example, the state detection unit 21 may detect the type of the posture of person 2 as the state type. In this case, the state detection unit 21 may input the input image to the posture detection AI to estimate the posture of person 2, or may detect the posture of person 2 by means other than the posture detection AI (for example, rule-based, etc.).
[0023] The image position detection unit 22 detects the image position which is the position on the input image of person 2. Note that the state detection unit 21 may be a part of the image detection unit 22, and the state detection unit 21 may detect the state of the person and also detect the position on the input image of person 2. The output determination unit 23 determines the reliability of the state type detected by the state detection unit 21 based on the state type detected by the state detection unit 21 and the image position detected by the image position detection unit 22. As described above, depending on the image position of person 2, there is a possibility that the detection accuracy of the state type by the state detection unit 21 may decrease. For example, in the directly below image area where the area near the directly below area of the camera 11 is captured, there is a possibility that the detection accuracy of the state type by the state detection unit 21 may decrease.
[0024] FIG. 4(a) is a schematic diagram of an example of the optical axis OA direction and the imaging angle θf of the camera when the camera 11 is a single-direction type camera, and FIG. 4(b) is a schematic diagram of the directly below image area 3D in the captured image 3 of the single-direction type camera 11. When the camera 11 is a single-direction type camera, the vertical imaging angle θf of the camera 11 is about 70 degrees to 90 degrees, and it is installed above the monitoring space 1 and the optical axis OA direction is set to have a depression angle θd. In the captured image 3 of such a single-direction type camera 11, the directly below image area 3D is located at the lower end of the captured image 3.
[0025] FIG. 5(a) is a schematic diagram of an example of the optical axis OA direction and the imaging range of the camera when the camera 11 is an omnidirectional camera, and FIG. 5(b) is a schematic diagram of the directly below image area 4D in the captured image 4 of the omnidirectional camera 11. The hatched area RF in FIG. 5(a) indicates the imaging range of the omnidirectional camera 11. The omnidirectional camera 11 takes the entire range (360 degrees) as the imaging area, and is installed above the monitoring space 1 and the optical axis OA direction is directed straight down. In the captured image 4 of such an omnidirectional camera 11, the overall image is circular, and the directly below image area 4D is located at the center of the captured image 4.
[0026] When the person 2 is shown in the directly below image areas 3D and 4D, the lower body is hidden by the person 2's own body, making it easy to misdetect the posture of the person 2. For example, when detecting the posture of a person from the positions of each joint point using posture detection AI, the posture detection AI infers the joint points that cannot be seen and detects the posture including the inferred joint points. When the person 2 is shown in the directly below image areas 3D and 4D, it often infers that the joint points of the lower body are bent, and misdetects sitting behavior even though the person is not performing sitting behavior. That is, it becomes easy to misdetect the mode type that emphasizes the joint information of the lower body of a person among the mode types. Specifically, it becomes easy to misdetect the mode type of bending the knees or the mode type of sitting on the floor with the knees bent.
[0027] Therefore, when the mode type detected by the mode detection unit 21 is sitting behavior, the output determination unit 23 determines the reliability of the mode type detected by the mode detection unit 21 more strictly when the person 2 is shown in the directly below image areas 3D and 4D than when the person 2 is shown in an area other than the directly below image areas 3D and 4D. Strict determination means determining the reliability of the detection result of the mode detection unit 21 to be low for the mode type that is likely to be misdetected in a predetermined image area. By determining the reliability to be low, it is possible to suppress the situation where it is determined that the mode type is a certain mode type even though it is not a mode type that is easily detected for the person 2. Note that sitting behavior is an example of the "sitting posture" described in the claims.
[0028] For example, as a method for the output determination unit 23 to determine the reliability of the state type detected by the state detection unit 21, it may be determined whether the state type detected by the state detection unit 21 is a false detection (whether the state type can be trusted). In this case, for example, when the confidence level of the state type output when the state detection unit 21 detects the state type is less than the confidence level threshold, it is determined that the state type detected by the state detection unit 21 is a false detection (the state type cannot be trusted), and when it is equal to or higher than the confidence level threshold, it is determined that the state type detected by the state detection unit 21 is not a false detection (the state type can be trusted). Note that the confidence level threshold is an example of the "determination criterion" described in the claims.
[0029] For example, when the state type detected by the state detection unit 21 is the sitting behavior, the output determination unit 23 may set a higher confidence level threshold when the person 2 is shown in the directly below image areas 3D and 4D than when the person 2 is shown in an area other than the directly below image areas 3D and 4D. FIG. 6 is a diagram showing an example of the setting of the confidence level threshold. When the state type detected by the state detection unit 21 is the sitting behavior, the confidence level threshold is set to 95% when the person 2 is shown in the directly below image areas 3D and 4D, and is set to 80% when the person 2 is shown in an area other than the directly below image areas 3D and 4D.
[0030] From this, when the person 2 is shown in an area other than the directly below image areas 3D and 4D, if the confidence level is not 95% or higher, it is determined as a false detection, while when the person 2 is shown in an area other than the directly below image areas 3D and 4D, if the confidence level is 80% or higher, it is determined that there is no false detection. On the other hand, when the state type detected by the state detection unit 21 is a state type other than the sitting behavior or the angry behavior, the confidence level threshold is set to 80% regardless of whether the image position of the person 2 is within the directly below image areas 3D and 4D. Therefore, when the state type detected by the state detection unit 21 is a state type other than the sitting behavior or the angry behavior, the confidence level threshold does not change depending on whether the image position of the person 2 is within the directly below image areas 3D and 4D.
[0031] In addition, when Person 2 appears in the directly below image areas 3D and 4D, it is difficult to detect the movement of Person 2's arm. Even if the arm is waved (such as when punching), the movement is shown as small, making it difficult to detect angry behaviors such as fights and destruction. Also, it is difficult to detect the vertical position of the person's arm, and although the arm is being raised, it appears in the image as if it is being lowered, making it difficult to detect actions such as hold-ups and confrontations. Therefore, when the state type detected by the state detection unit 21 is an angry behavior, a hold-up, or a confrontation action, and Person 2 appears in the directly below image areas 3D and 4D, the output determination unit 23 moderately determines the reliability of the state type detected by the state detection unit 21 as compared to the case where Person 2 appears in an area other than the directly below image areas 3D and 4D. Moderate determination means determining the reliability of the detection result of the state detection unit 21 to be high for a state type that is difficult to detect in a predetermined image area. By determining the reliability to be high, it is possible to suppress misjudging that a certain state type is not the case when the state type of Person 2 is difficult to detect.
[0032] For example, the output determination unit 23 may set the confidence threshold when the state type detected by the state detection unit 21 is an angry behavior to be lower when Person 2 appears in the directly below image areas 3D and 4D than when Person 2 appears in an area other than the directly below image areas 3D and 4D. For example, when the state type detected by the state detection unit 21 is an angry behavior, the confidence threshold is set to 70% when Person 2 appears in the directly below image areas 3D and 4D, and set to 80% when Person 2 appears in an area other than the directly below image areas 3D and 4D.
[0033] The output determination unit 23 determines the output from the output unit 15 for the abnormal states among the state types detected by the state detection unit 21 according to the reliability determined by the output determination unit 23. For example, when the detected state type is an abnormal state and the confidence of the state type is equal to or higher than the confidence threshold, the output determination unit 23 may output the state type from the output unit 15, and may prohibit the output of the state type when the confidence is less than the confidence threshold.
[0034] For example, the output determination unit 23 may also determine the device to which the output unit 15 outputs the state type according to the determined reliability. For example, the output unit 15 may output the state type to an external device of the image processing apparatus 10 by wired communication or wireless communication. The external device may be, for example, a center terminal or a local monitoring terminal used by a monitor who monitors the monitoring space 1, or may be a customer mobile terminal.
[0035] The output determination unit 23 may switch which of these terminal devices to output the state type according to the reliability determined by the output determination unit 23. For example, when the confidence level of the state type is equal to or higher than the confidence level threshold, the state type may be output to the center terminal or the local monitoring terminal, and when the confidence level of the state type is less than the confidence level threshold, the state type may be output to the customer mobile terminal.
[0036] (Modification example) (1) In addition to the directly below image regions 3D and 4D, there are also image regions where the detection accuracy of the state type of the person 2 decreases. For example, even in an image region in which a distant region away from the camera 11 is captured (which may be referred to as a "distant image region" in the following description), there is a possibility that the detection accuracy of the state type by the state detection unit 21 decreases.
[0037] FIG. 7(a) is a schematic diagram of the distant image region 3F in the captured image 3 of the single-direction type camera, and FIG. 7(b) is a schematic diagram of the distant image region 4F in the captured image 4 of the omnidirectional camera. The distant image region 3F in the captured image 3 of the single-direction type camera is located at the lower end of the captured image 3. On the other hand, the distant image region 4F in the captured image 4 of the omnidirectional camera is located at the image edge (periphery) of the captured image 4.
[0038] When Person 2 appears in the distant image regions 3F and 4F, the image of Person 2 becomes smaller, making it difficult to detect the fine movements of the joint points and thus difficult to detect the angry behaviors (arguments / destruction) that require such detection. Also, when Person 2 appears in the distant image regions 3F and 4F, the line-of-sight direction from the camera 11 to Person 2 becomes closer to the horizontal direction. For this reason, it becomes difficult to detect the actions of Person 2 when Person 2 moves in the line-of-sight direction of the camera 11. For example, when Person 2 falls in the line-of-sight direction of the camera 11, it becomes difficult to detect the falling action (e.g., "fall" or "collapse").
[0039] Therefore, when the state type detected by the state detection unit 21 is an angry behavior or a falling action, the output determination unit 23 gently determines the reliability of the state type detected by the state detection unit 21 when Person 2 appears in the distant image regions 3F and 4F to be lower than when Person 2 appears in a region other than the distant image regions 3F and 4F.
[0040] FIG. 8 is a diagram showing another example of setting the confidence threshold. For example, the output determination unit 23 may set the confidence threshold when the state type detected by the state detection unit 21 is an angry behavior or a falling action to be lower when Person 2 appears in the distant image regions 3F and 4F than when Person 2 appears in a region other than the distant image regions 3F and 4F. For example, the confidence threshold when the state type detected by the state detection unit 21 is an angry behavior or a falling action is set to 80% when Person 2 appears in a region other than the distant image regions 3F and 4F, and set to 70% when Person 2 appears in the distant image regions 3F and 4F.
[0041] On the other hand, when the state type detected by the state detection unit 21 is a state type other than an angry behavior or a falling action, the confidence threshold is set to 80% regardless of whether the image position of Person 2 is within the distant image regions 3F and 4F. For this reason, when the state type detected by the state detection unit 21 is a state type other than an angry behavior or a falling action, the confidence threshold does not change depending on whether the image position of Person 2 is within the distant image regions 3F and 4F.
[0042] (2) The output determination unit 23 may set the above confidence threshold value used to determine the reliability of the state type detected by the state detection unit 21 based on the situation where the camera 11 is installed. FIGS. 9(a) and 9(b) are diagrams showing setting examples of the correction amount of the confidence threshold value according to the installation conditions of the camera 11. When the camera 11 is a unidirectional camera, for example, as the installation condition of the camera 11, according to the depression angle θd of the optical axis OA of the camera 11, when the image position of the person 2 is within the directly below image areas 3D and 4D, the confidence threshold value may be corrected. This is because as the depression angle θd increases, it becomes easier to misdetect the sitting behavior and more difficult to detect the angry behavior.
[0043] When the state type detected by the state detection unit 21 is the sitting behavior, when the person 2 is shown in the area of the directly below image area 3D, if the depression angle θd is 30 degrees, 40 degrees, or 45 degrees, for example, 2%, 3%, and 5% are respectively added to the confidence threshold value. When the state type detected by the state detection unit 21 is the angry behavior, for example, 2%, 3%, and 5% are respectively subtracted from the confidence threshold value.
[0044] When the camera 11 is a unidirectional camera, for example, as the installation condition of the camera 11, according to the angle of view of the camera 11 (for example, the vertical angle of view θf), when the image position of the person 2 is within the directly below image area 3D, the confidence threshold value may be corrected. This is because as the angle of view increases, it becomes easier to misdetect the sitting behavior and more difficult to detect the angry behavior.
[0045] When the state type detected by the state detection unit 21 is the sitting behavior, when the person 2 is shown in the area of the directly below image area 3D, if the vertical angle of view θf is 70 degrees, 80 degrees, or 90 degrees, for example, 2%, 3%, and 5% are respectively added to the confidence threshold value. When the state type detected by the state detection unit 21 is the angry behavior, for example, 2%, 3%, and 5% are respectively subtracted from the confidence threshold value.
[0046] Thus, for example, when the depression angle θd is 45 degrees and the vertical field angle θf is 90 degrees, the correction amount of the confidence threshold for the sitting behavior is +10%. When the confidence threshold when person 2 is shown in an area other than the directly below image area 3D is 80%, the confidence threshold when person 2 is shown in the directly below image area 3D is set to 80 + 10 = 90%.
[0047] In addition to or instead of correcting the confidence threshold, the size of the directly below image area 3D may be corrected according to the installation conditions of the camera 11. For example, the larger the depression angle θd and the field angle (for example, the vertical field angle θf), the larger the directly below image area 3D may be set.
[0048] Also, as the depression angle θd or the vertical field angle θf increases, the image area capturing the vicinity directly below the camera expands. Therefore, as a method of making the reliability determination criteria different, the image area for determining reliability may be expanded as the depression angle θd or the vertical field angle θf increases.
[0049] Also, the higher the installation height of the camera, the more difficult it is to detect angry behavior or falling behavior in the distant image area. Therefore, the confidence threshold may be subtracted as the installation height increases.
[0050] (3) When the state detection unit 21 cannot continuously detect the same state type for a predetermined duration or more, the output determination unit 23 may determine that the state type detected by the state detection unit 21 is a false detection, and when the same state type is continuously detected for a predetermined duration or more, the output determination unit 23 may determine that the state type detected by the state detection unit 21 is not a false detection. The output determination unit 23 may set a predetermined duration according to the image position of person 2.
[0051] In this case, when person 2 is shown in a specific image area (for example, the directly below image area 3D, 4D or the distant image area 3F, 4F), if the reliability of the state type detected by the state detection unit 21 is strictly determined compared to when shown in an area other than the specific image area, the predetermined duration may be set longer, and if it is determined leniently, it may be set shorter.
[0052] For example, when the state type detected by the state detection unit 21 is sitting behavior, the predetermined duration when the person 2 appears in the directly below image areas 3D and 4D may be set longer than the predetermined duration when the person 2 appears in an area other than the directly below image areas 3D and 4D. Also, when the state type detected by the state detection unit 21 is angry behavior, the predetermined duration when the person 2 appears in the directly below image areas 3D and 4D may be set shorter than the predetermined duration when the person 2 appears in an area other than the directly below image areas 3D and 4D. Also, when the state type detected by the state detection unit 21 is angry behavior or falling behavior, the predetermined duration when the person 2 appears in the distant image areas 3F and 4F may be set shorter than the predetermined duration when the person 2 appears in an area other than the distant image areas 3F and 4F.
[0053] Similarly, the output determination unit 23 may output a false detection of the state type detected by the state detection unit 21 according to whether the ratio of the time during which the state detection unit 21 detects the same state type within a period of a predetermined length is less than a predetermined ratio threshold. The output determination unit 23 may set the predetermined ratio threshold according to the image position of the person 2. In this case, when the person 2 appears in a specific image area (for example, the directly below image areas 3D and 4D or the distant image areas 3F and 4F), if the reliability of the state type detected by the state detection unit 21 is strictly determined compared to when the person 2 appears in an area other than the specific image area, the predetermined ratio threshold may be set larger, and if it is determined leniently, it may be set smaller. The predetermined duration and the predetermined ratio threshold are examples of the "determination criteria" described in the claims.
[0054] (4) As a method for the output determination unit 23 to determine the reliability of the state type detected by the state detection unit 21, the confidence level of the state type detected by the state detection unit 21 may be corrected. The output determination unit 23 may output the confidence level corrected by the output determination unit 23, or may switch the output destination device of the state type detected by the state detection unit 21 according to the confidence level corrected by the output determination unit 23. Also, a confidence threshold for determining whether the state type detected by the state detection unit 21 is a false detection may be a common value for all state types, and the corrected confidence level and the threshold may be compared to determine whether it is a false detection.
[0055] In this case, the output determination unit 23 may correct the confidence level output by the state detection unit 21 according to the image position of person 2. For example, when person 2 is shown in a specific image area (for example, the immediate lower image areas 3D, 4D or the distant image areas 3F, 4F), if the reliability of the state type detected by the state detection unit 21 is strictly determined compared to the case of being shown in an area other than the specific image area, the correction may be made so that the confidence level output by the state detection unit 21 becomes smaller, and when it is determined leniently, the correction may be made so that it becomes larger. The correction value for the output determination unit 23 to correct the reliability of the state type detected by the state detection unit 21 is an example of the "determination criterion" described in the claims.
[0056] (5) The output determination unit 23 may determine the reliability using the score obtained by the state detection unit 21. At this time, the state detection unit 21 outputs the obtained score to the output determination unit 23. For example, when person 2 in the sitting state exists in the area directly below the camera, the value of the score is made lower than that in other image areas and compared with the threshold for determining the reliability, or when person 2 in the sitting state exists in the area directly below the camera, the threshold for determining the reliability is made higher than that in other image areas and compared with the score output from the state detection unit 21. The correction value for the output determination unit 23 to correct the value obtained by the state detection unit 21 for state detection is an example of the "determination criterion" described in the claims.
[0057] (6) The state detection unit 21 may correct the confidence level to be output to the output determination unit 23 for the state type detected by the state detection unit 21 according to the image position of the person 2, and determine the reliability of the state type detected by the state detection unit 21. For example, when the person 2 is shown in a specific image area (for example, the immediate lower image areas 3D, 4D or the distant image areas 3F, 4F), if the reliability of the state type detected by the state detection unit 21 is strictly determined compared to the case where the person 2 is shown in an area other than the specific image area, the confidence level may be corrected to be smaller, and if it is determined leniently, it may be corrected to be larger. Since the output determination unit 23 determines whether to output to the output 15 using the corrected confidence level, the correction value corrected by the state detection unit 21 is an example of the "determination criterion" described in the claims.
[0058] (7) When the state detection unit 21 detects the state type from a plurality of frames in the input image, the state type detected by the number of frames equal to or more than the threshold value among a predetermined number of consecutive frames may be detected as the state type of the person 2, and the reliability of the state type detected by the state detection unit 21 may be determined. In the following description, the threshold value of the number of frames required to detect the state type among a predetermined number of consecutive frames may be referred to as the "required number of frames threshold value". For example, when the person 2 is shown in the immediate lower image areas 3D, 4D, if the sitting behavior is detected in a number of frames equal to or more than the required number of frames threshold value (for example, 12 frames or more) among a predetermined number of consecutive frames (for example, 15 frames), the sitting behavior may be output to the output determination unit 23 as the state type of the person 2. On the other hand, when the person 2 is shown in an area other than the immediate lower image areas 3D, 4D, no threshold value is provided for the required number of frames for detecting the sitting behavior. Since the output determination unit 23 determines whether to output to the output 15 using the result of determining the reliability of the state type according to the image position, the required number of frames threshold value provided by the state detection unit 21 according to the image position is an example of the "determination criterion" described in the claims.
[0059] (8) The state detection unit 21 may determine the reliability of the posture of the person 2 detected by the posture detection AI according to the image position of the person 2. That is, the state detection unit 21 may be a part of the output determination unit 23. The state detection unit 21 may detect the state type of the person and determine the reliability of the detected state type based on the image position of the person 2. As a method for determining the reliability of the detected posture, the state detection unit 21 may determine whether the posture detected by the posture detection AI is correct. When the detected posture is correct, the state detection unit 21 may output the action detected based on the time-series change information of the posture detected by the posture detection AI to the output determination unit 23 as the state type of the person 2. In an embodiment where the state type output to the output determination unit 23 is a posture, the posture detected by the posture detection AI may be output to the output determination unit 23 as the state type of the person 2.
[0060] When the detected posture is incorrect, detecting the action of the person 2 using the posture detected by the posture detection AI and outputting the posture detected by the posture detection AI as the state type of the person 2 are prohibited. At this time, the state detection unit 21 may determine whether the posture of the person 2 detected by the posture detection AI is correct according to the angle and movement amount of the joint points (skeleton) recognized from the image of the person 2.
[0061] For example, the state detection unit 21 may determine that the sitting posture detected by the posture detection AI is correct when the angle between the line connecting the thigh (waist) and the knee and the line connecting the knee and the foot is less than or equal to a threshold value. The state detection unit 21 may set the threshold value when the person 2 is shown in the immediate lower image regions 3D and 4D to be smaller than the threshold value when the person 2 is shown in a region other than the immediate lower image regions 3D and 4D.
[0062] For example, the state detection unit 21 may determine that the angry posture detected by the posture detection AI is correct when the amount of movement of the line connecting the shoulder and the arm is equal to or greater than a threshold value, or when the number of times the line connecting the shoulder and the arm moves away from and then returns to the vicinity of the person 2 is equal to or greater than a threshold value. The state detection unit 21 may set a threshold value for the case where the person 2 appears in the directly below image areas 3D and 4D or the distant image areas 3F and 4F to be smaller than the threshold value for the case where the person 2 appears in other areas.
[0063] For example, the state detection unit 21 may determine the confidence level of the posture detected by the posture detection AI as a method for determining the reliability of the posture detected by the posture detection AI. When the posture detected by the posture detection AI is a sitting posture, the state detection unit 21 may correct the confidence level so that it is smaller when the person 2 appears in the directly below image areas 3D and 4D than when the person 2 appears in areas other than the directly below image areas 3D and 4D. When the posture detected by the posture detection AI is an angry posture, the state detection unit 21 may correct the confidence level so that it is larger when the person 2 appears in the directly below image areas 3D and 4D or the distant image areas 3F and 4F than when the person 2 appears in other areas.
[0064] (Effect of the Embodiment) (1) The image processing apparatus 10 includes an image acquisition unit 20 that acquires an input image captured by the camera 11, a state detection unit 21 that detects the state type of an object appearing in the input image, an image position detection unit 22 that detects an image position that is the position of the object on the input image, and an output determination unit 23 that determines the reliability of the detected state type based on the state type detected by the state detection unit 21 and the image position detected by the image position detection unit 22. This can suppress a decrease in the detection accuracy of the state type of the object in a specific image area.
[0065] (2) The output determination unit 23 may change the reliability determination criterion between the case where the object appears in a predetermined image area among a plurality of image areas of the input image and the case where the object appears in other image areas. This can suppress a decrease in the detection accuracy of the state type of the object in a predetermined image area.
[0066] (3) Criteria may be defined for each of a plurality of image areas for each state type. This enables determination of reliability according to the difference in state type when an object appears in a specific image area, and can suppress a decrease in the detection accuracy of the state type of the object. (4) Criteria may be defined based on the installation conditions of the camera. This can suppress a decrease in the detection accuracy of the state type of the object according to the installation conditions of the camera when the object appears in a specific image area.
[0067] (5) The predetermined image area may be a directly below image area in which the vicinity of the area directly below the camera is shown. When an object appears in the directly below image area, criteria may be defined so as to more strictly determine reliability than when the object appears in other image areas. This can suppress a decrease in the detection accuracy of the state type of the object in the directly below image area.
[0068] (6) The output determination unit 23 may determine reliability based on the image position only when the detected state type is a predetermined state type among a plurality of state types. This can suppress a decrease in the detection accuracy of the state when an object appears in a specific image area in a specific state type.
[0069] (7) The predetermined state type may be a sitting state in which the posture of the person who is the object is sitting. The output determination unit 23 may more strictly determine reliability when the object appears in a predetermined image area among a plurality of image areas of the input image than when the object appears in other image areas. This can suppress a decrease in the detection accuracy of the state type when the object appears in a specific image area when the state type detected by the state detection unit 21 is the sitting state.
[0070] (8) When the detected state type is the first state type, the output determination unit 23 determines the reliability more strictly when the object appears in a predetermined image area among a plurality of image areas of the input image than when the object appears in other image areas. When the detected state type is the second state type, the output determination unit 23 may determine the reliability more leniently when the object appears in a predetermined image area than when the object appears in other image areas. Thereby, it is possible to suppress false detection of the first state type when the object appears in a specific image area, and to suppress failure of detection of the second state type.
[0071] (9) The state detection unit 21 may include a posture detection unit that inputs the input image to a learning model generated by machine learning that uses image data as input and outputs the posture of the object appearing in the image of the image data, and an action detection unit that detects the type of action of the object based on the change in the detected posture and outputs the detected type of action as the state type. Thereby, when estimating the posture of the object using the learning model, it is possible to suppress a decrease in the detection accuracy of the state type of the object due to the image position of the object.
[0072] (10) The learning model may estimate the posture of the object based on the joint positions of the object. Thereby, when estimating the posture of the object based on the joint positions of the object, it is possible to suppress a decrease in the detection accuracy of the state type of the object due to the image position of the object. (11) The state detection unit 21 may detect the state type of the object based on the image position detected by the image position detection unit 22, or may calculate the confidence level of the state type of the object based on the image position. Thereby, it is possible to improve the detection accuracy of the state type and the calculation accuracy of the confidence level in the state detection unit 21.
[0073] The image processing apparatus according to an embodiment of the present invention can contribute to solving social problems such as a decrease in the working population.
Explanation of Reference Numerals
[0074] 1…Monitoring space, 2…Person (object to be detected), 10…Image processing device, 11…Camera, 12…Input unit, 13…Memory unit, 14…Control unit, 15…Output unit, 20…Image acquisition unit, 21…State detection unit, 22…Image position detection unit, 23…Judgment unit, 23…Output judgment unit
Claims
1. An image acquisition unit that acquires an input image captured by a camera; A state detection unit that detects a predetermined state caused by a moving object appearing in the input image; An image position detection unit that detects an image position that is the position of the moving object on the input image; A determination unit that determines the reliability of the detection result of the predetermined state by the state detection unit based on the image position detected by the image position detection unit; An image processing apparatus comprising the above.
2. The determination unit varies the reliability determination criteria when the predetermined state is detected between the case where the moving object is located in a predetermined image region among a plurality of image regions of the input image and the case where the moving object is located in other image regions. The image processing apparatus according to claim 1, characterized in that.
3. The moving object is a person, The predetermined image region is a partial image region of the input image in which the vicinity directly below the camera is shown, The predetermined state is a posture or action that detects with emphasis on the lower body of the person, The determination unit determines that the reliability is lower when the moving object is located in the predetermined image region than when the moving object is located in the other image regions. The image processing apparatus according to claim 2, characterized in that.
4. The state detection unit detects a plurality of state types caused by the moving object, The position on the input image of the predetermined image region is set for each state type. The image processing apparatus according to claim 2, characterized in that.
5. The state detection unit detects a plurality of state types caused by the moving object, When the determination unit detects a first state type or a second state type in the predetermined image region, when the first state type is detected, the determination unit determines that the reliability is lower than that of the other image regions, and when the second state type is detected, the determination unit determines that the reliability is higher than that of the other image regions. The image processing apparatus according to claim 2, characterized in that.
6. The state detection unit detects a plurality of state types caused by the moving object, The determination unit determines the reliability based on the image position when a predetermined state type among the plurality of state types is detected. The image processing apparatus according to claim 1, characterized in that.
7. The determination unit varies the reliability determination criteria when the predetermined state is detected based on the image position and the installation state of the camera. The image processing apparatus according to claim 1, characterized in that.
8. The moving object is a person, The image processing apparatus according to claim 1, wherein the state detection unit detects the posture or behavior of a person as the state based on the joint information of the person.
9. The image processing apparatus according to claim 1, wherein the state detection unit inputs the input image into a learning model generated by machine learning that uses image data as an input and outputs a predetermined state of the moving object shown in the image of the image data as output data, and outputs the predetermined state of the moving object.
10. An image acquisition process for acquiring an input image captured by a camera, A state detection process for detecting a predetermined state of a moving object shown in the input image, An image position detection process for detecting an image position that is the position of the moving object on the input image, A determination process for determining the reliability of the detection result of the predetermined state by the state detection process based on the image position detected in the image position detection process, A computer program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
Action analysis system and action analysis method
JP2023030965A