Image processing device and computer program

The image processing device enhances detection accuracy by identifying and accounting for occluded portions, correcting for errors in state detection due to partial concealment, thereby maintaining reliable identification of the target's state.

JP2025118211APending Publication Date: 2025-08-13SECOM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024013403
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-31
Publication Date
2025-08-13

AI Technical Summary

Technical Problem

Existing image detection systems suffer from decreased accuracy when a part of the detection target is occluded, leading to erroneous detections, especially for actions like kneeling, bending, or crouching, and may fail to detect the state of the target altogether if occlusion prohibits determination processing.

Method used

An image processing device with an occluded portion detection unit that identifies concealed parts of a moving object and a determination unit to assess the reliability of the detection result based on these concealed portions, using learning models and rule-based methods to correct for occlusions.

Benefits of technology

The device effectively suppresses the decrease in detection accuracy by determining the reliability of the detected state, reducing false positives and ensuring accurate identification of the target's state even when parts are hidden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025118211000001_ABST
    Figure 2025118211000001_ABST
Patent Text Reader

Abstract

To minimize a decrease in detection accuracy due to the fact that a portion of a detection target is hidden when detecting the state of the detection target appearing in an image.SOLUTION: An image processing device 10 is provided, comprising: an image acquisition unit 20 for acquiring an input image captured by a camera 11; a state detection unit 21 for detecting a predetermined state of a moving object appearing in the input image; a concealed portion detection unit 22 for detecting a concealed portion of the moving object; and a determination unit 23 configured to determine reliability of a detection result of the predetermined state by the state detection unit 21 on the basis of the concealed portion detected by the concealed portion detection unit 22.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image processing device and a computer program. [Background technology]

[0002] Conventionally, technologies for detecting the state (posture, behavior, etc.) of a detection target captured in an image have been proposed. For example, Patent Document 1 below discloses a security system that identifies suspicious behavior of a person captured on a surveillance camera. In this security system, if an occlusion occurs in a person area in a camera image where a person exists, the system prohibits suspicious behavior determination processing for the occluded person area. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-086471 Summary of the Invention [Problem to be solved by the invention]

[0004] If a part of the detection target is occluded by something in an image showing the detection target, the accuracy of detecting the state of the detection target may decrease. For example, even if the person being detected is not in a sitting position (e.g., a sitting position such as bending, squatting, or kneeling), the state of the detection target may be erroneously detected as sitting. Furthermore, for example, if a certain part is occluded, the accuracy of detecting a specific state may decrease significantly. Furthermore, if the determination process is prohibited when occlusion occurs, as in Patent Document 1, it may become impossible to detect the state of the detection target at all. The present invention aims to suppress a decrease in detection accuracy caused by occlusion of a part of a detection target when detecting the state of the detection target shown in an image. [Means for solving the problem]

[0005] An image processing device according to one embodiment of the present invention includes an image acquisition unit that acquires an input image captured by a camera, a mode detection unit that detects a predetermined mode of a moving object captured in the input image, a concealed portion detection unit that detects concealed portions of the moving object, and a determination unit that determines the reliability of the detection result of the predetermined mode by the mode detection unit based on the concealed portions detected by the concealed portion detection unit. [Effects of the Invention]

[0006] According to the present invention, when detecting the state of a detection target appearing in an image, it is possible to suppress a decrease in detection accuracy due to the part of the detection target being occluded. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a schematic configuration diagram of an example of an image processing apparatus according to an embodiment; [Figure 2] FIG. 1 is a block diagram illustrating an example of a functional configuration of an image processing apparatus according to an embodiment. [Figure 3] FIG. 10 is an explanatory diagram of an example of a method for detecting an occluded portion. [Figure 4] FIG. 10 is a diagram illustrating an example of setting part information indicating essential parts for each type of condition. [Figure 5] 10A and 10B are diagrams illustrating examples of setting body part information indicating the importance of each body part of a person for each aspect type. [Figure 6] 1A is a schematic diagram of an example of the optical axis direction and angle of view of a unidirectional camera, and FIG. 1B is a schematic diagram of the direct image area in an image captured by the unidirectional camera. [Figure 7] 1A is a schematic diagram of an example of the optical axis direction and angle of view of an omnidirectional camera, and FIG. 1B is a schematic diagram of the direct image area in an image captured by the omnidirectional camera. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Note that the embodiments of the present invention shown below are merely examples of devices and methods for embodying the technical concept of the present invention, and the technical concept of the present invention does not limit the structure, arrangement, etc. of the components to those described below. The technical concept of the present invention can be modified in various ways within the technical scope defined by the claims.

[0009] (First embodiment) 1 is a schematic configuration diagram of an example of an image processing device according to an embodiment. The image processing device 10 detects the type of appearance of a detection target in a monitored space 1. As an example, assume that a person 2 and a structure 5 exist in the monitored space 1. In this specification, a case where the type of appearance of the person 2 as a detection target existing in the monitored space 1 is detected is exemplified.

[0010] The image processing device 10 detects the posture and behavior of the person 2 as the type of appearance of the person 2. The type of appearance to be detected includes detection based on the posture and behavior of the person 2 alone, and detection based on the postures and behavior of multiple people. The image processing device 10 detects abnormal posture and abnormal behavior of the person 2 as an abnormal appearance of the person 2 that notifies a monitor or the like of the monitored space 1 that an abnormal situation has occurred.

[0011] Examples of postures and behaviors to be detected as abnormal include "threatening (a person raising their arms, etc.)," "fighting (a person hitting another person with their fist, etc.)," "destruction (a person swinging their arms, etc.)," "falling (a person who has collapsed from a standing position and changed to a lying position)," "holding up (a person who continues to raise both hands due to being threatened, etc.)," "shoving (a person who has extended their arms toward another person (e.g., a person pointing a weapon, etc. at another person))," "crouching (a person who continues to crouch with their knees or waist bent)," "dogeza (a person who bows to another person)," "crouching (a person who is prostrating and bowing to another person)," "crouching (a person who is lying on their stomach)," and "dogeza (a person who is lying down and sleeping)." Note that "crouching" and "dogeza" are examples of a "first posture type classified as a sitting position" as defined in the claims, and "threatening" and "holding up" are examples of a "second posture type classified as a standing position."

[0012] Furthermore, a type of state other than the abnormal state may be detected. For example, postures such as "standing," "standing upright," and "sitting," or actions such as "walking" and "running" may be detected. The image processing device 10 may detect the state type of a detection target other than the person 2. For example, the detection target may be a machine having a manipulator or a non-humanoid robot. For example, the image processing device 10 may detect the orientation of the manipulator or its posture, such as its bending or stretching state, as the state type, or may detect a movement, which is a change in posture, as the state type.

[0013] The image processing device 10 includes a camera 11, an input unit 12, a storage unit 13, a control unit 14, and an output unit 15. Of these, the storage unit 13 and the control unit 14 may be realized by a so-called computer, and the input unit 12 and the output unit 15 may be realized as peripheral devices of the computer. The camera 11 is placed in the monitored space 1 to generate images of objects present in the monitored space 1. The camera 11 may be a unidirectional camera with a horizontal angle of view and a vertical angle of view of about 90 degrees, for example. The camera 11 may also be an omnidirectional camera with a shooting area (monitoring area) in all directions (360 degrees). Note that the image processing device 10 may not be equipped with the camera 11 and may instead acquire images from an external shooting device.

[0014] The input unit 12 includes a user interface such as a keyboard, a mouse, etc. that is operated by a user to input data, etc. The input unit 12 is connected to the control unit 14, converts user operations into operation signals, and outputs the signals to the control unit 14. The input unit 12 may also include a DVD (Digital Versatile Disc) drive and a USB (Universal Serial Bus) interface. The input unit 12 inputs data to the control unit 14 as a file, and outputs data from the control unit 14 as a file. The storage unit 13 is a memory device such as a ROM (Read Only Memory) or a RAM (Random Access Memory), and stores various programs and various data. The storage unit 13 is connected to the control unit 14, and inputs and outputs this information to and from the control unit 14.

[0015] The control unit 14 is composed of arithmetic devices such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), and an MCU (Micro Control Unit). The control unit 14 is connected to the storage unit 13, and operates as various processing units by reading and executing computer programs from the storage unit 13, and stores and reads various data in the storage unit 13. The functions of the image processing device 10 described below are realized by the control unit 14 executing computer programs stored in the storage unit 13.

[0016] The output unit 15 is connected to the control unit 14 and outputs the detection result of the mode type by the control unit 14. The output unit 15 may include a display device such as a liquid crystal display or a CRT (Cathode Ray Tube) display. The output unit 15 may also include an audio signal output device such as a speaker or a buzzer. The output unit 15 may also include a network interface or the like that transmits and receives data between the image processing device 10 and an external device via wired or wireless communication.

[0017] The control unit 14 acquires a captured image generated by the camera 11 by capturing an image of the person 2 present in the monitored space 1, and detects the type of appearance of the person 2 appearing in the captured image. In this case, if a part of a body part of person 2 is hidden from the view of camera 11 and this part does not appear in the captured image generated by camera 11, there is a risk that the accuracy of detecting the state of person 2 may decrease depending on the part hidden from the view of camera 11 (hereinafter simply referred to as a "hidden part"). Note that a hidden part is not limited to a case where the entire part is hidden, but may also include a case where only a part of the part is hidden. Furthermore, in this specification, "hidden" is explained to mean not only intentionally hiding a part of person 2, but also accidentally hiding a part of person 2.

[0018] For example, actions such as "kneeling," "bending," and "crouching" tend to have a greater drop in detection accuracy compared to other actions (such as "holding up"), even if the same body part is hidden. In particular, actions such as "kneeling," "bending," and "crouching" tend to have a more significant drop in detection accuracy when specific body parts such as the knees or ankles are hidden.

[0019] Therefore, the image processing device 10 of the embodiment detects the appearance type of the person 2 appearing in the captured image 3, detects the hidden parts of the person 2, and determines the reliability of the detected appearance type based on the detected appearance type and the hidden parts. 2 is a block diagram showing an example of the functional configuration of the image processing device 10 according to the embodiment. The image processing device 10 includes the camera 11 and output unit 15, an image acquisition unit 20, a state detection unit 21, an occluded part detection unit 22, and an output determination unit 23.

[0020] The image acquisition unit 20 acquires an image generated by the camera 11 by capturing an image of an object present in the monitored space 1 as an input image. The manner detection unit 21 detects the manner type of the person 2 appearing in the input image acquired by the image acquisition unit 20. For example, the manner detection unit 21 may detect the type of behavior of the person 2 as the manner type. In this case, the manner detection unit 21 first detects the posture of the person 2 appearing in the input image.

[0021] For example, the posture detection unit 21 may estimate the posture of person 2 by inputting the input image into a learning model (hereinafter, sometimes referred to as "posture detection AI (Artificial Intelligence)") generated by machine learning, which uses image data as input and postures of people appearing in the image data as output data. For example, the posture detection AI may identify the positions of joint points as joint information of person 2 appearing in the input image, and estimate the posture of person 2 based on feature quantities of the positions of the joint points. Note that the posture may also be estimated based on skeletal information (position of the skeleton) of person 2 as joint information.

[0022] Furthermore, for example, the manner detection unit 21 may detect the posture of the person 2 by means other than a learning model (for example, a rule base, etc.). For example, the manner detection unit 21 may extract a silhouette image of the person 2 by a background subtraction method, and detect the posture by obtaining a score from the similarity between the extracted silhouette image and silhouette images of each posture registered in advance, or may extract the positions of the person's joint points from the silhouette of the person extracted by the background subtraction method, and detect the posture by obtaining a score from the similarity between the positions of the joint points of each posture registered in advance and the positions of the extracted joint points of the person.

[0023] Next, the manner detection unit 21 determines the behavior of the person 2 based on the time-series change information of the detected posture. For example, the manner detection unit 21 may determine the behavior of the person 2 using a rule base based on the time-series change information of the detected posture. For example, if the posture of the person 2 changes in the order of standing upright, collapsing, and lying down, the type of behavior of the person 2 may be determined to be a falling behavior. The manner detection unit 21 may output a "certainty" that is the degree of accuracy (correctness probability) of the detected posture. For example, the type of the detected behavior and its certainty may be output, such as "the probability that it is a falling behavior is 80%."

[0024] Furthermore, for example, the behavior detection unit 21 may determine the behavior of each frame of the input image based on time-series change information of the posture between the target frame and the previous frame, and detect the type of behavior of person 2 by majority vote of the postures detected in each frame for the same person appearing in multiple frame images. Furthermore, the manner detection unit 21 may detect the type of behavior of the person 2 using a learning model generated by machine learning that uses time-series information of moving images and joint points as input and person behavior as output data.

[0025] The manner detection unit 21 may output not only a single manner type but also a plurality of manner types and their certainty levels (for example, "standing upright" 90%, "walking" 80%, etc.). Furthermore, for example, the manner detection unit 21 may detect, as the manner type, the type of posture of the person 2. In this case, the manner detection unit 21 may input an input image to a posture detection AI to estimate the posture of the person 2, or may detect the posture of the person 2 by means other than the posture detection AI (for example, a rule base, etc.).

[0026] The occluded portion detection unit 22 detects occluded portions of the person 2 (head, neck, wrist, elbow, shoulder, etc.) that are hidden from the view of the camera 11. For example, the occluded portion detection unit 22 may detect occluded portions by fixed obstacles, occluded portions by moving objects (e.g., other people), or occluded portions of the person 2 that are hidden by the person 2's own body (so-called occluded portions due to self-occlusion).

[0027] For example, when detecting a part obscured by a fixed obstacle, the obscured part detection unit 22 may store in advance position information of the fixed obstacle such as the structure 5 in the monitored space 1, and determine which part of the person 2 is obscured based on the positional relationship between the position of each part of the person 2 and the position of the obstacle. For example, the image processing device 10 may store the position information of the fixed obstacle such as the structure 5 in the storage unit 13 of FIG. 1. The occluded portion detection unit 22 may acquire position information of each part of the person 2 from the manner detection unit 21. The manner detection unit 21 may detect the position of each part of the person 2 when estimating the posture of the person 2, and output the position information to the occluded portion detection unit 22. For example, the manner detection unit 21 may output position information of the joints of the person 2 to the occluded portion detection unit 22.

[0028] For example, when detecting a part obscured by a moving object, the obscured part detection unit 22 may detect the moving object from the input image acquired by the image acquisition unit 20, and determine which part of the person 2 is obscured based on the positional relationship between the position of each part of the person 2 and the position of the obstacle.

[0029] Furthermore, for example, when detecting a part obscured by self-occlusion, the obscured part detection unit 22 may detect the part obscured based on a learning model that has learned the positions of the parts of the person 2 and whether or not the parts are obscured. 3 is an explanatory diagram of another example of a method for detecting occluded parts due to self-occlusion. In Fig. 3, circles (e.g., ja, jb) indicate joints of person 2 detected by manner detection unit 21, and lines (e.g., c1, c2) connecting the joints indicate connections between the joint parts.

[0030] For example, when the posture detection unit 21 calculates the score of the connection relationship between joint parts using existing posture estimation software such as OpenPose, the occlusion part detection unit 22 may determine that the joint part jb is an occlusion part when the connection c between the joint parts ja and jb is broken or the connection c is weak. Alternatively, for example, the occluded part detection unit 22 may track joint parts using history information of past joint parts and predict whether a part will be occluded. In this case, if three-dimensional part information is estimated, more accurate prediction is possible based on the anteroposterior relationship of the parts.

[0031] See Fig. 2. The output determination unit 23 determines the reliability of the state type detected by the state detection unit 21. As described above, when a part of the person 2 is hidden from the view of the camera 11, there is a risk that the accuracy of detecting the appearance of the person 2 may decrease depending on the hidden part.

[0032] Therefore, the output determination unit 23 determines the reliability of the detection result of the state type by the state detection unit 21 based on the occluded portion detected by the occluded portion detection unit 22. As a method for determining the reliability of the state type detected by the state detection unit 21, for example, it may be determined whether or not the state type detected by the state detection unit 21 is an erroneous detection (whether or not the state type is reliable).

[0033] In this case, the output determination unit 23 may determine whether the mode type detected by the mode detection unit 21 is a false detection (the mode type is unreliable) depending on whether a predetermined part (sometimes referred to as a "required part" in the following description) for each mode type is a hidden part.

[0034] For example, the storage unit 13 may store body part information indicating required body parts for each type of state. Fig. 4 is a diagram showing an example of setting body part information indicating required body parts for each type of state. In the body part information exemplified in Fig. 4, for example, "wrist" and "elbow" are set as required body parts for the state type "hold up", and "ankle", "knee", and "waist" are set as required body parts for the state type "dogeza".

[0035] The output determination unit 23 may determine whether the mode type detected by the mode detection unit 21 is a false detection based on a comparison of the essential parts and hidden parts corresponding to the mode type detected by the mode detection unit 21 among the part information stored in the memory unit 13. For example, the output determination unit 23 may determine whether the mode type detected by the mode detection unit 21 is a false detection depending on whether the hidden part detected by the hidden part detection unit 22 is included in the essential parts.

[0036] For example, if an obscured part is included in the essential parts, the mode type detected by the mode detection unit 21 may be determined to be a false detection, and if not, it may be determined to be not a false detection. For example, if the mode type detected by the mode detection unit 21 is "hold-up" and the obscured part detected by the obscured part detection unit 22 includes at least one of "wrist" and "elbow," it may be determined to be a false detection, and if the obscured part is neither "wrist" nor "elbow," it may be determined to be not a false detection.

[0037] Furthermore, for example, the output determination unit 23 may determine whether the type of state detected by the state detection unit 21 is an erroneous detection based on the proportion of the hidden parts detected by the hidden part detection unit 22 among the parts designated as essential parts. For example, if all of the parts designated as essential parts are concealed parts, the type of the part detected by the aspect detection unit 21 may be determined to be an erroneous detection, and if only some of the parts designated as essential parts are concealed parts, it may be determined not to be an erroneous detection. That is, if the proportion of concealed parts among the parts designated as essential parts is 100%, it may be determined to be an erroneous detection, and if it is less than 100%, it may be determined not to be an erroneous detection.

[0038] The rate at which it is determined that a detection is false may be a value less than 100%. For example, assuming that the rate at which it is determined that a detection is false is set to 50%, when the type of condition detected by the condition detection unit 21 is "kneeling," if any two of the essential parts "ankle," "knee," and "waist" are concealed (i.e., when the concealed parts are 66%), it is determined that it is false, and if there is one or less concealed part (i.e., when the concealed parts are 33% or less), it is determined that it is not false. The rate at which it is determined that a detection is false may be a fixed value, or different values may be stored according to the type of condition.

[0039] See Fig. 2. The output determination unit 23 determines the output from the output unit 15 for an abnormal state among the state types detected by the state detection unit 21, depending on the determination result of the reliability of the state type detected by the state detection unit 21. For example, when the output determination unit 23 determines that the mode type detected by the mode detection unit 21 is a mode type previously stored in the storage unit 13 as an abnormal mode and that the mode type is not an erroneous detection, the output determination unit 23 outputs the mode type from the output unit 15. On the other hand, even if the mode type detected by the mode detection unit 21 is a mode type previously stored in the storage unit 13 as an abnormal mode, when the output determination unit 23 determines that the mode type is an erroneous detection, the output of the mode type is prohibited.

[0040] For example, the output determination unit 23 may determine an output destination device to which the output unit 15 outputs the mode type based on the determination result of the reliability of the mode type detected by the mode detection unit 21. For example, the output unit 15 may output the mode type determined not to be a false detection to an external device of the image processing device 10 via wired or wireless communication. The external device may be, for example, a center terminal or local monitoring terminal used by a monitor monitoring the monitored space 1, or a customer mobile terminal. For example, the output determination unit 23 may determine the mode of information to be output from the output unit 15 based on the determination result of the reliability of the mode type detected by the mode detection unit 21. For example, when the output unit 15 determines that the reliability is low (false detection or low certainty) due to a high degree of obscuration, the output unit 15 may add auxiliary information indicating that the reliability is low and output the information. For example, to visually identify the low reliability, the output may display a bounding box (a frame) of a color corresponding to the reliability surrounding the person.

[0041] The output determination unit 23 may switch to which of these terminal devices the mode type is to be output, depending on the determination result of the reliability of the mode type detected by the mode detection unit 21. For example, if the mode detection unit 21 determines that the mode type detected is not an erroneous detection, the output determination unit 23 may output the mode type to the center terminal or the local monitoring terminal, and if the mode detection unit 21 determines that the detection is an erroneous detection, the output determination unit 23 may output the mode type to the customer mobile terminal.

[0042] (Second embodiment) The storage unit 13 of the second embodiment stores body part information indicating the importance of each body part of the person 2 for each aspect type. The output determination unit 23 of the second embodiment determines the reliability of the aspect type detected by the aspect detection unit 21 based on an evaluation value calculated based on the body part information corresponding to the aspect type detected by the aspect detection unit 21 and the concealed part detected by the concealed part detection unit 22, among the body part information stored in the storage unit 13. In the following description, the evaluation value calculated based on the body part information corresponding to the aspect type detected by the aspect detection unit 21 and the concealed part detected by the concealed part detection unit 22 may be simply referred to as an "evaluation value."

[0043] As a method for determining the reliability of the mode type detected by the mode detection unit 21, for example, it may be determined whether or not the mode type detected by the mode detection unit 21 is a false detection. For example, the output determination unit 23 may determine that the mode type detected by the mode detection unit 21 is a false detection (the mode type is unreliable) when the evaluation value is equal to or greater than a predetermined determination threshold, and may determine that the mode type detected by the mode detection unit 21 is not a false detection (the mode type is reliable) when the evaluation value is less than the determination threshold.

[0044] Fig. 5 is a diagram showing an example of setting body part information indicating the importance of each body part of person 2 for each state type. In the body part information shown in Fig. 5, for example, in the case of the second state type "hold up" classified as a standing position, the importance of each of the body parts "head," "neck," "right wrist," "left wrist," "right elbow," "left elbow," "right shoulder," "left shoulder," "right hip," "left hip," "right knee," "left knee," "right ankle," and "left ankle" is set to "0.5."

[0045] For example, in the case of the first posture type "dogeza" classified as a sitting position, the importance of each part of the body, "head," "neck," "right wrist," "left wrist," "right elbow," "left elbow," "right shoulder," "left shoulder," "right hip," "left hip," "right knee," "left knee," "right ankle," and "left ankle," is set to "0.5," "0.5," "0.5," "0.5," "0.5," "0.5," "0.5," "0.7," "0.7," "0.9," "0.9," "0.8," and "0.8," respectively.

[0046] The output determination unit 23 may calculate the evaluation value by, for example, adding up the importance levels set for the concealed parts detected by the concealed part detection unit 22. For example, if the manner type detected by the manner detection unit 21 is "dogeza" and the concealed parts detected by the concealed part detection unit 22 are "right wrist," "right elbow," "left hip," and "left knee," the output determination unit 23 may calculate the evaluation value as 0.5+0.5+0.7+0.9=2.6, which is the sum of the importance levels "0.5," "0.5," "0.7," and "0.9" set for these concealed parts, respectively.

[0047] Note that for the first posture type classified as a sitting position, the detection accuracy tends to decrease significantly when parts of the lower body are occluded. For this reason, in the body part information illustrated in Fig. 5, the importance of the lower body parts "right hip," "left hip," "right knee," "left knee," "right ankle," and "left ankle" for the first posture type "kneeling" classified as a sitting position is set to a higher value than the importance of the lower body parts for the second posture type "holding up" classified as a standing position. This makes it possible to reduce erroneous detection of the first posture type when parts of the lower body are occluded.

[0048] The importance of the lower body part of the first mode type may be set to be greater than the importance of the upper body part of the first mode type. The importance of the upper body part of the first mode type may be set to be the same as or approximately the same as the importance of the upper body part of the second mode type.

[0049] (Variation) Modifications of the embodiments will be described below, which are applicable to both the first and second embodiments. (1) The output determination unit 23 may determine the reliability of the state type detected by the state detection unit 21 based on the image position, which is the position of the person 2 on the input image, and the body part information set according to at least one of the installation conditions of the camera 11.

[0050] The detection accuracy of the state type by the state detection unit 21 may be reduced depending on the image position of the person 2. For example, in a direct-below image area in which the vicinity of the area directly below the camera 11 is captured, the detection accuracy of the state type by the state detection unit 21 may be reduced. Figure 6(a) is a schematic diagram of an example of the optical axis OA direction and angle of view θf of the camera when the camera 11 is a unidirectional camera, and Figure 6(b) is a schematic diagram of the direct image area 3D in the captured image 3 of the unidirectional camera 11.

[0051] When the camera 11 is a unidirectional camera, the vertical angle of view θf of the camera 11 is approximately 70 to 90 degrees, and the camera 11 is installed above the monitored space 1, with the direction of the optical axis OA set to have a depression angle θd. In the captured image 3 of such a unidirectional camera 11, the nadir image area 3D is located at the bottom edge of the captured image 3.

[0052] Fig. 7(a) is a schematic diagram of an example of the direction of the optical axis OA of the camera and the shooting range when the camera 11 is an omnidirectional camera, and Fig. 7(b) is a schematic diagram of the direct-down image area 4D in the image 4 captured by the omnidirectional camera 11. The hatched area RF in Fig. 7(a) indicates the shooting range of the omnidirectional camera 11. The omnidirectional camera 11 has a shooting range in all directions (360 degrees), is installed above the monitored space 1, and the direction of the optical axis OA is directed directly downward. The image 4 captured by such an omnidirectional camera 11 is a circular image as a whole, and the direct-under image area 4D is located at the center of the captured image 4.

[0053] If Person 2 appears in the 3D or 4D direct image area, the lower half of Person 2's body will be hidden, making it easier for Person 2's posture to be mistakenly detected. For example, when posture detection AI is used to detect a person's posture from the position of each joint point, the posture detection AI will infer invisible joint points and detect the posture including these inferred joint points. If Person 2 appears in the 3D or 4D direct image area, it will often infer that the joint points of the lower half of the body are bent, resulting in a false detection of Person 2 kneeling even though they are not.

[0054] Therefore, when person 2 is captured in the 3D or 4D direct image area, the part information may be set so that the lower body parts are essential parts, or the part information may be modified so that the lower body parts are given greater importance than when person 2 is captured in an area other than the 3D or 4D direct image area. For example, the output determination unit 23 may modify the body part information depending on whether the person 2 appears in the direct-behind image areas 3D and 4D. Alternatively, the body part information modified depending on whether the person 2 appears in the direct-behind image areas 3D and 4D may be set in advance and stored in the storage unit 13.

[0055] Furthermore, when the camera 11 is a unidirectional camera, the body part information may be corrected according to the depression angle θd of the optical axis OA of the camera 11, for example, as an installation condition of the camera 11. This is because the greater the depression angle θd, the more likely the lower body parts are to be concealed. For example, the part information may be modified so that the greater the depression angle θd, the more essential the lower body parts become, or the part information may be modified so that the importance of the lower body parts becomes greater compared to when the depression angle θd is small.

[0056] (2) The output determination unit 23 may correct the certainty of the mode type detected by the mode detection unit 21 as a method of determining the reliability of the mode type detected by the mode detection unit 21. The output determination unit 23 may output the certainty corrected by the output determination unit 23, or may switch the output destination device of the mode type detected by the mode detection unit 21 in accordance with the certainty corrected by the output determination unit 23.

[0057] The output determination section 23 may correct the confidence level output by the manner detection section 21 in accordance with the occluded portion detected by the occluded portion detection section 22 . For example, the output determination unit 23 may correct the certainty factor output by the manner detection unit 21 depending on whether or not an essential part is included in the hidden parts detected by the hidden part detection unit 22. For example, if an essential part is included in the hidden parts, the certainty factor may be corrected to be smaller.

[0058] For example, the output determination unit 23 may correct the certainty factor output by the manner detection unit 21 based on the evaluation value. For example, the higher the evaluation value, the smaller the certainty factor output by the manner detection unit 21 may be corrected.

[0059] (3) If the state detection unit 21 is unable to detect the same state type continuously for a predetermined duration or longer, the output determination unit 23 may determine that the state type detected by the state detection unit 21 is a false detection, and if the state detection unit 21 is able to detect the same state type continuously for a predetermined duration or longer, the output determination unit 23 may determine that the state type detected by the state detection unit 21 is not a false detection. The output determination unit 23 may set the predetermined duration in accordance with the occluded portion detected by the occluded portion detection unit 22 .

[0060] For example, the output determination unit 23 may set the predetermined duration depending on whether or not an essential part is included in the hidden parts detected by the hidden part detection unit 22. For example, if an essential part is included in the hidden parts, the predetermined duration may be corrected to be longer. For example, the output determination unit 23 may set the predetermined duration based on the evaluation value. For example, the larger the evaluation value, the longer the predetermined duration may be set.

[0061] Similarly, the output determination unit 23 may determine whether the mode type detected by the mode detection unit 21 is a false detection, depending on whether the proportion of time during which the mode detection unit 21 detected the same mode type in a predetermined period of time is less than a predetermined proportion threshold. The output determination unit 23 may set the predetermined proportion threshold depending on the concealed part detected by the concealed part detection unit 22. For example, if the concealed portion includes an essential portion, the predetermined percentage threshold may be set to a larger value.Furthermore, for example, the predetermined percentage threshold may be set to a larger value as the evaluation value increases.

[0062] (6) When detecting a manner type from multiple frames in an input image, the manner detection unit 21 may detect a manner type detected in a number of frames equal to or greater than a threshold among a predetermined number of consecutive frames as the manner type of person 2, and determine the reliability of the manner type detected by the manner detection unit 21. In the following description, the threshold value for the number of frames required to detect a manner type among a predetermined number of consecutive frames may be referred to as the “required frame number threshold value.”

[0063] For example, if the concealed part includes an essential part or the evaluation value exceeds a predetermined threshold, when the same condition type is detected in a required frame number threshold or more (e.g., 12 frames or more) out of a predetermined number of consecutive frames (e.g., 15 frames), it may be detected as the condition type of person 2. On the other hand, if the concealed part does not include an essential part or the evaluation value is equal to or less than the predetermined threshold, no threshold is set for the required number of frames for detecting the condition type.

[0064] (7) The manner detection unit 21 may determine the reliability of the posture of the person 2 detected by the posture detection AI, depending on the occluded parts detected by the occluded part detection unit 22. In other words, the manner detection unit 21 may be part of the output determination unit 23, and the manner detection unit 21 may detect the type of the person's posture and determine the reliability of the detected type of posture based on the occluded parts. As a method of determining the reliability of the detected posture, the posture detection unit 21 may determine whether the posture detected by the posture detection AI is correct depending on the occluded part. In an embodiment in which the posture type output to the output determination unit 23 is posture, if the posture detection unit 21 determines that the occluded part is not included in the essential parts and the detected posture is correct, the posture detection unit 21 may output the posture detected by the posture detection AI to the output determination unit 23 as the posture type of the person 2.

[0065] If it is determined that the detected posture is incorrect because the hidden body part is included in the essential body parts, the posture detection AI is prohibited from outputting the detected posture as the state type of person 2. At this time, the posture detection unit 21 may determine whether the posture of the person 2 detected by the posture detection AI is correct or not, depending on the angles and movement amounts of joint points (skeleton) and the like recognized from the image of the person 2.

[0066] For example, the manner detection unit 21 may determine that the threatening posture detected by the posture detection AI is correct if the number of times the arm is swung down is equal to or greater than a threshold. In this case, the manner detection unit 21 may set a higher threshold when the concealed body part includes an essential body part than when the concealed body part does not include an essential body part. Furthermore, the threshold may be set higher as the evaluation value increases. For example, the posture detection unit 21 may determine that the sitting posture detected by the posture detection AI is correct when the angle between the line connecting the waist and knee and the line connecting the knee and ankle is equal to or smaller than a threshold. The threshold when the hidden part includes an essential part may be set smaller than the threshold when the hidden part does not include an essential part. Furthermore, the threshold may be set smaller as the evaluation value increases.

[0067] Furthermore, for example, the posture detection unit 21 may determine that the fighting posture detected by the posture detection AI is correct if the amount of movement of the line connecting the shoulder and arm is equal to or greater than a threshold, or if the number of times that the line connecting the shoulder and elbow and / or the elbow and wrist move away from and return to the periphery of person 2 is equal to or greater than a threshold. In this case, the posture detection unit 21 may set the threshold value when the hidden body part includes an essential body part to be higher than the threshold value when the hidden body part does not include an essential body part. Furthermore, the threshold value for the amount of movement or the threshold value for the number of times that the line moves away from and returns to the periphery may be set higher as the evaluation value increases.

[0068] Furthermore, for example, the manner detection unit 21 may determine the reliability of the posture detected by the posture detection AI by changing the certainty factor of the posture detected by the posture detection AI. The manner detection unit 21 may set the certainty factor when an essential part is included in the occluded part to be smaller than the certainty factor when an essential part is not included in the occluded part. Furthermore, the certainty factor may be set to be smaller as the evaluation value increases.

[0069] (Effects of the embodiment) (1) The image processing device 10 includes an image acquisition unit 20 that acquires an input image captured by the camera 11, a state detection unit 21 that detects a predetermined state of a moving object captured in the input image, a concealed portion detection unit 22 that detects concealed portions of the moving object, and an output determination unit 23 that determines the reliability of the detection result of the predetermined state by the state detection unit 21 based on the concealed portions detected by the concealed portion detection unit 22. This makes it possible to suppress a decrease in detection accuracy due to the part of the detection target being occluded when detecting the state of the detection target shown in the image.

[0070] (2) The state detection unit 21 may detect a state type that indicates the type of state of the moving object. The image processing device 10 may further include a storage unit 13 that stores part information that indicates essential parts for each state type. The output determination unit 23 may determine the reliability of the state type detection result by the state detection unit 21 based on a comparison between the part information corresponding to the state type detected by the state detection unit 21 and the hidden parts. This makes it possible to determine the reliability according to the difference in the state type when detecting the state of the detection object captured in the image, and prevents a decrease in the detection accuracy of the state type of the moving object.

[0071] (3) The output determination unit 23 may determine the reliability of the detection result of the mode type by the mode detection unit 21 based on the proportion of the parts determined to be hidden parts among the parts specified as essential parts. This makes it possible to determine reliability according to the degree to which essential parts are hidden, and to prevent a decrease in the detection accuracy of the state type of the moving body.

[0072] (4) The manner detection unit 21 may detect manner types that indicate the types of manners of the moving object. The image processing device 10 may further include a storage unit 13 that stores part information that indicates the importance of each part of the moving object for each manner type. The output determination unit 23 may determine the reliability of the manner type detection result by the manner detection unit 21 based on an evaluation value calculated based on the part information corresponding to the manner type detected by the manner detection unit 21 and the concealed part. This makes it possible to determine the reliability of the state type detection according to the importance of each part, and to prevent a decrease in the detection accuracy of the state type of the moving body.

[0073] (5) The manner detection unit 21 may determine a manner type that indicates a type of manner, which is a posture or behavior of a person who is a moving body. The manner type may include a first manner type classified as a sitting position and a second manner type classified as a standing position. The storage unit 13 may store a higher value of importance for a part of the lower body for the first manner type than a value of importance for a part of the lower body for the second manner type. This makes it possible to determine the reliability according to the degree to which the lower body parts of a person are concealed, and makes it possible to prevent a decrease in the detection accuracy of the person's appearance type.

[0074] (6) The output determination unit 23 may determine the reliability of the detection result of the condition type by the condition detection unit 21 based on the part information that is corrected according to at least one of the installation conditions of the camera 11 and the image position, which is the position of the moving object on the input image. This makes it possible to prevent a decrease in the detection accuracy of the state type of the moving object depending on the photographing conditions of the camera 11.

[0075] An image processing apparatus according to an embodiment of the present invention can contribute to solving social issues such as a decline in the working population. [Explanation of symbols]

[0076] 1...Monitored space, 2...Person (detection target), 10...Image processing device, 11...Camera, 12...Input unit, 13...Memory unit, 14...Control unit, 15...Output unit, 20...Image acquisition unit, 21...Aspect detection unit, 22...Occluded part detection unit, 23...Output determination unit

Claims

1. an image acquisition unit that acquires an input image captured by a camera; a state detection unit that detects a predetermined state of a moving object captured in the input image; an occluded portion detection unit that detects an occluded portion of the moving object; a determination unit that determines reliability of a detection result of the predetermined mode by the mode detection unit based on the occluded portion detected by the occluded portion detection unit; An image processing device comprising:

2. the state detection unit detects a state type that indicates a type of state of the moving object; the image processing device further includes a storage unit that stores part information indicating essential parts for each of the condition types; the determination unit determines reliability of a result of detection of the state type by the state detection unit based on a comparison between the part information corresponding to the state type detected by the state detection unit and the concealed part.

2. The image processing device according to claim 1, wherein:

3. The image processing device according to claim 2, characterized in that the determination unit determines the reliability of the detection result of the aspect type by the aspect detection unit based on the proportion of the areas determined to be the hidden areas among the areas designated as the essential areas.

4. the state detection unit detects a state type that indicates a type of state of the moving object; the image processing device further includes a storage unit configured to store part information indicating the importance of each part of the moving object for each of the aspect types; the determination unit determines reliability of a detection result of the state type detected by the state detection unit based on the part information corresponding to the state type detected by the state detection unit and the concealed part.

2. The image processing device according to claim 1, wherein:

5. the manner detection unit determines a manner type that indicates a type of manner, which is a posture or behavior, of the person who is the moving object; The posture type includes a first posture type classified into a sitting position and a second posture type classified into a standing position, the storage unit stores the importance level related to the lower body part of the first condition type as a value greater than the importance level related to the lower body part of the second condition type; 5. The image processing device according to claim 4.

6. the manner detection unit determines a manner type that indicates a type of manner, which is a posture or behavior, of the person who is the moving object; the storage unit stores the importance of a lower body part of the posture type classified as a sitting position as a value greater than the importance of an upper body part of the posture type; 5. The image processing device according to claim 4.

7. The image processing device described in any one of claims 2 to 6, characterized in that the determination unit determines the reliability of the detection result of the condition type by the condition detection unit based on the part information that is set according to at least one of the installation conditions of the camera and the image position, which is the position of the moving object on the input image.

8. an image acquisition process for acquiring an input image captured by the camera; A behavior detection process for detecting a predetermined behavior of a moving object captured in the input image; an occluded portion detection process for detecting an occluded portion of the moving object; a determination process for determining reliability of a detection result of the predetermined mode by the mode detection process based on the occluded portion detected by the occluded portion detection process; A computer program characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Crime prevention system, crime prevention method, and crime prevention program

    JP2023086471A