Image processing device and computer program

The image processing apparatus enhances detection accuracy by integrating congestion assessment to adjust detection reliability, addressing the issue of decreased accuracy in crowded environments.

JP2025105027APending Publication Date: 2025-07-10SECOM CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2023223285
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The detection accuracy of a detection target's state in an image decreases due to the degree of congestion around the target, leading to potential misclassification of actions such as angry behaviors in crowded environments.

Method used

An image processing apparatus that includes an image acquisition unit, a state detection unit to identify the target's state, a congestion detection unit to assess the level of congestion, and an output determination unit to determine the reliability of the detection based on the congestion level, adjusting confidence thresholds and detection criteria accordingly.

Benefits of technology

The apparatus effectively suppresses the decrease in detection accuracy by ensuring reliable classification of states, particularly in congested conditions, by refining detection criteria based on congestion levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025105027000001_ABST
    Figure 2025105027000001_ABST
Patent Text Reader

Abstract

To minimize a decrease in detection accuracy due to the state of congestion around a target object when detecting the state of a detection target appearing in an image.SOLUTION: An image processing device 10 is provided, comprising: an image acquisition unit 20 for acquiring an input image captured by a camera; a state detection unit 21 for detecting a predetermined state of a moving object appearing in the input image; a congestion level detection unit 22 for detecting the level of congestion around the moving object; and a determination unit 23 configured to determine reliability of a detection result of the predetermined state by the state detection unit 21 on the basis of the level of congestion detected by the congestion level detection unit 22.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus and a computer program.

Background Art

[0002] Conventionally, techniques for detecting the state of a detection target shown in an image have been proposed. For example, Patent Document 1 below discloses a system for detecting states such as the posture and behavior of a person from image data obtained by photographing a person who is the detection target.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] However, there is a risk that the detection accuracy of the state of the detection target may decrease depending on the degree of congestion around the detection target shown in the image. For example, even if the person who is the detection target is not performing an angry action (actions when angry, such as destructive actions or fights), if the area around the person who is the detection target is congested, there is a risk of misdetecting the state of the person who is the detection target as an angry action. An object of the present invention is to suppress a decrease in detection accuracy due to the degree of congestion around an object when detecting the state of a detection target shown in an image.

Means for Solving the Problems

[0005] An image processing apparatus according to one aspect of the present invention includes an image acquisition unit that acquires an input image captured by a camera, a state detection unit that detects a predetermined state by a moving object shown in the input image, a congestion detection unit that detects the degree of congestion around the moving object, and a determination unit that determines the reliability of a detection result of the predetermined state by the state detection unit based on the degree of congestion detected by the congestion detection unit. [Effect of the Invention]

[0006] According to the present invention, when detecting the state of a detection target shown in an image, it is possible to suppress a decrease in detection accuracy due to the degree of congestion around the target object. [Brief Description of the Drawings]

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0008] Hereinafter, embodiments of the present invention will be described with reference to the drawings. The embodiments of the present invention shown below illustrate devices and methods for embodying the technical idea of the present invention, and the technical idea of the present invention does not specify the structure, arrangement, etc. of the components as the following. The technical idea of the present invention can be variously modified within the technical scope defined by the claims described in the claims.

[0009] FIG. 1 is a schematic configuration diagram of an example of the image processing apparatus according to the embodiment. The image processing apparatus 10 detects the type of state of a detection target in the monitoring space 1. In this specification, the case of detecting the type of state of a person 2 as a detection target existing in the monitoring space 1 is exemplified.

[0010] The image processing device 10 detects the posture and behavior of the person 2 as the mode type of the person 2. The mode types to be detected include those detected by the posture and behavior of the person 2 alone and those detected by the postures and behaviors of multiple persons. The image processing device 10 detects abnormal postures and abnormal behaviors of the person 2 as an abnormal mode, which is a mode of the person 2 notifying the monitor or the like in the monitoring space 1 that an abnormal situation has occurred. The postures and behaviors to be detected as abnormal modes are, for example, "threatening (a person waving their arms, etc.)", "fighting (a person hitting someone else with a fist, etc.)", "destroying (a person swinging their arm downwards, etc.)", "falling (a person who has changed from a standing posture to a collapsed and lying-on-side posture)", "holding up (a person who keeps their hands raised due to being threatened, etc.)", "pointing (a person stretching their arm towards someone else (for example, a person pointing a weapon at someone else))", "bending (a person who keeps bending their knees and waist and crouching)", "bowing (a person who bows down to show respect to someone else)", "crawling (a person lying on their stomach and spreading their body)", "lying on side (a person lying on their side and continuing to sleep)", etc. Furthermore, mode types other than the abnormal mode may be detected. For example, postures such as "standing position", "standing upright", "sitting position", etc. and behaviors such as "walking", "running", etc. may be detected. Note that the image processing device 10 may detect the mode types of detection targets other than the person 2. For example, it may detect the mode types of a machine having a manipulator or a non-humanoid robot as the detection target. For example, the image processing device 10 may detect postures such as the direction and bending state of the manipulator as the mode type, and may also detect operations that are changes in these postures as the mode type.

[0011] The image processing device 10 includes a camera 11, an input unit 12, a storage unit 13, a control unit 14, and an output unit 15. Among these, the storage unit 13 and the control unit 14 may be realized by a so-called computer, and the input unit 12 and the output unit 15 may be realized as peripheral devices of the computer. The camera 11 is arranged in the monitoring space 1 to generate an image of an object existing in the monitoring space 1. The camera 11 may be, for example, a single-direction type camera with a horizontal viewing angle and a vertical viewing angle of about 90 degrees. The camera 11 may also be an omnidirectional camera with a 360-degree omnidirectional shooting area (monitoring area). Note that the image processing device 10 may not include the camera 11 and may acquire video from an external imaging device.

[0012] The input unit 12 includes user interfaces such as a keyboard and a mouse that are operated by a user and used for data input and the like. The input unit 12 is connected to the control unit 14, converts the user's operation into an operation signal, and outputs it to the control unit 14. Further, the input unit 12 may include a DVD (Digital Versatile Disc) drive and a USB (Universal Serial Bus) interface. The input unit 12 inputs data to the control unit 14 as a file and outputs data from the control unit 14 as a file. The storage unit 13 is a memory device such as a ROM (Read Only Memory) and a RAM (Random Access Memory), and stores various programs and various data. The storage unit 13 is connected to the control unit 14 and inputs and outputs this information to and from the control unit 14.

[0013] The control unit 14 is composed of an arithmetic device such as a CPU (Central Processing Unit), a DSP (Digital Signal Processor), and an MCU (Micro Control Unit). The control unit 14 is connected to the storage unit 13, reads out a computer program from the storage unit 13 and executes it to operate as various processing units, stores various data in the storage unit 13, and reads it out. The functions of the image processing device 10 described below are realized by the control unit 14 executing a computer program stored in the storage unit 13.

[0014] The output unit 15 is connected to the control unit 14 and outputs the detection result of the state type by the control unit 14. The output unit 15 may include a display device such as a liquid crystal display or a CRT (Cathode Ray Tube) display. The output unit 15 may be provided with an audio signal output device such as a speaker or a buzzer. The output unit 15 may be provided with a network interface or the like that transmits and receives data between the image processing apparatus 10 and an external device by wired communication or wireless communication.

[0015] The control unit 14 acquires a captured image generated by photographing a person 2 existing in the monitoring space 1 by the camera 11, and detects the state type of the person 2 shown in the captured image. At this time, as shown in FIG. 1, when a large number of other persons 2a to 2d exist around the person 2 and the surroundings of the person 2 are crowded, the detection accuracy of the state of the person 2 may decrease depending on the degree of congestion around the person 2.

[0016] For example, among the state types detected by the control unit 14, there is a type that is detected on the condition that there are other persons near the person 2, that is, a plurality of persons exist nearby. Such state types include violent acts (such as destruction and quarrels), nuisance acts (kowtowing), and threatening acts (pushing), which are represented as dangerous acts. When kowtowing is being done in a crowded place, the person doing the kowtowing may be stepped on by other persons, or other persons may get caught on the person doing the kowtowing and fall, so it is regarded as a dangerous act. In the following description, a state type that is detected on the condition that there are other persons near the person 2, such as an angry act, kowtowing, or pushing, may be referred to as a "predetermined state type".

[0017] When the degree of congestion around the person 2 is high, there is a possibility that a predetermined state type may be erroneously detected as the state type of the person 2 even if the true state type of the person 2 is not the predetermined state type. Therefore, the image processing apparatus 10 according to the embodiment detects the state type of the person 2 shown in the captured image 3, detects the degree of congestion around the person 2, and determines the reliability of the detected state type based on the detected state type and the degree of congestion. FIG. 2 is a block diagram showing an example of the functional configuration of the image processing apparatus 10 according to the embodiment. The image processing apparatus 10 includes the above-described camera 11 and output unit 15, an image acquisition unit 20, a state detection unit 21, a congestion level detection unit 22, and an output determination unit 23.

[0018] The image acquisition unit 20 acquires, as an input image, an image generated by the camera 11 capturing an object existing in the monitoring space 1. The state detection unit 21 detects the state type of the person 2 shown in the input image acquired by the image acquisition unit 20. For example, the state detection unit 21 may detect the type of action of the person 2 as the state type. In this case, the state detection unit 21 first detects the posture of the person 2 shown in the input image.

[0019] For example, the state detection unit 21 may input the input image into a learning model (hereinafter sometimes referred to as "posture detection AI (Artificial Intelligence)") generated by machine learning that takes image data as input and outputs the posture of a person shown in the image data, and estimate the posture of the person 2. For example, the posture detection AI may identify the positions of the joint points as the joint information of the person 2 shown in the input image, and estimate the posture of the person 2 based on the feature amounts of the positions of the joint points. Note that the posture may be estimated based on the skeleton information (positions of the skeleton) of the person 2 as the joint information.

[0020] Also, for example, the state detection unit 21 may detect the posture of the person 2 by means other than the learning model (for example, rule-based, etc.). For example, the state detection unit 21 may extract the silhouette image of the person 2 by the background difference method, obtain a score from the similarity between the silhouette image of each registered posture and the extracted silhouette image, and detect the posture, or extract the positions of the joint points of the person from the silhouette of the person extracted by the background difference method, and obtain a score from the similarity between the positions of the joint points of each registered posture and the positions of the extracted joint points of the person, and detect the posture. Next, the state detection unit 21 determines the action of person 2 based on the time-series change information of the detected posture. For example, the state detection unit 21 may determine the action of person 2 based on a rule-based approach based on the time-series change information of the detected posture. For example, when the posture of person 2 changes in the order of upright, falling down, and lying horizontally, it may be determined that the type of action of person 2 is a falling action. The state detection unit 21 may output the "confidence level", which is the correct probability of the detected posture. For example, the type of detected action and its confidence level may be output, such as the probability of a falling action being 80%.

[0021] Also, for example, the state detection unit 21 determines the action for each frame of the input image based on the time-series change information of the posture between the target frame and the previous frame, and detects the type of action of person 2 by a majority vote of the postures detected in each frame for the same person existing in the images of multiple frames. Further, the state detection unit 21 may detect the type of action of person 2 using a learning model generated by machine learning that takes video or time-series information of joint points as input and outputs the action of a person as output data.

[0022] The state detection unit 21 may output not only the type of a single state but also the types of multiple states and their confidence levels (e.g., "upright standing position" 90%, "walking" 80%, etc.). Also, for example, the state detection unit 21 may detect the type of the posture of person 2 as the state type. In this case, the state detection unit 21 may input the input image to the posture detection AI to estimate the posture of person 2, or may detect the posture of person 2 by means other than the posture detection AI (e.g., rule-based, etc.).

[0023] The congestion detection unit 22 detects the congestion level of people around person 2 based on the input image. For example, the congestion detection unit 22 may scan the input image acquired by the image acquisition unit 20 with a density estimator that has learned feature amounts for each density using density images obtained by photographing spaces where people exist at various densities in advance, so as to estimate the congestion level of people in the area around person 2. Also, for example, a plurality of levels may be defined as the congestion level output by the congestion detection unit 22. Also, for example, the congestion detection unit 22 may detect the population density itself (e.g., 2.5 people / m 2 etc.) as the congestion level in the area around person 2. Also, the definition of the congestion levels may be by a method other than the number of people per 1m 2 ; for example, it may be defined that the higher the range of the image area where people are overlapping and shown, the higher the congestion level. Also, a congestion map may be created on the input image, and based on the image position where person 2 exists and the congestion map, the congestion level at the image position where the person 2 exists may be detected. Note that the congestion detection unit 22 may be a part of the state detection unit 21, and the state detection unit 21 may detect the state of person 2 and also detect the congestion level of people in the area around the person 2.

[0024] In the following description, a case will be exemplified where the congestion detection unit 22 detects whether the area around person 2 is in any of the three congestion level areas, the "low congestion area", the "medium congestion area", or the "high congestion area", as the congestion level of people around person 2. For example, the "low congestion area" may be an area of 0.0 people / m 2 or more and 2.0 people / m 2 or less, the "medium congestion area" may be an area higher than 2.0 people / m 2 and 4.0 people / m 2 or less, and the "high congestion area" may be an area higher than 4.0 people / m 2 or more.

[0025] The output determination unit 23 determines the reliability of the detection result of the state type detected by the state detection unit 21 based on the state type detected by the state detection unit 21 and the congestion level detected by the congestion detection unit 22. As described above, the mode type detected on the condition that another person exists near person 2 may be misdetected when the degree of congestion around person 2 is high.

[0026] Therefore, when the mode type detected by the mode detection unit 21 is a predetermined mode type, the output determination unit 23 determines the reliability of the mode type detected by the mode detection unit 21 based on the degree of congestion detected by the congestion detection unit 22. For example, the higher the degree of congestion, the more strictly the reliability is determined. To determine more strictly means to determine that the reliability of the detection result of the mode detection unit 21 is low for a mode type that is likely to be misdetected when person 2 is in a highly congested area. By determining the reliability to be low, it is possible to suppress the determination that the mode type is the mode type that person 2 is likely to detect even though it is not the mode type that person 2 is likely to detect.

[0027] For example, as a method for the output determination unit 23 to determine the reliability of the mode type detected by the mode detection unit 21, it may be determined whether the mode type detected by the mode detection unit 21 is a misdetection (whether the mode type can be trusted). In this case, for example, the output determination unit 23 determines that the mode type detected by the mode detection unit 21 is a misdetection (the mode type cannot be trusted) when the confidence level of the mode type output when the mode detection unit 21 detects the mode type is less than the confidence level threshold, and determines that the mode type detected by the mode detection unit 21 is not a misdetection (the mode type can be trusted) when the confidence level is equal to or higher than the confidence level threshold. Note that the confidence level threshold is an example of the "determination criterion" described in the claims.

[0028] For example, the output determination unit 23 may set the confidence level threshold to be larger as the degree of congestion is higher when the mode type detected by the mode detection unit 21 is an angry behavior, a bowing gesture, or a pushing gesture. Figure 3 is a diagram showing an example of setting the confidence level threshold. When the mode type detected by the mode detection unit 21 is an angry behavior (quarreling, destruction), the confidence level thresholds are set to 80%, 90%, and 95% when the area around person 2 is a low-congestion area, a medium-congestion area, and a high-congestion area, respectively.

[0029] Accordingly, when the area around Person 2 is a low-congestion area, if the confidence level is 80% or more, it is determined that there is no false detection. On the other hand, when it is a medium-congestion area or a high-congestion area, if the confidence level is less than 90% or 95% respectively, it is determined as a false detection. Thus, the higher the congestion level, the more strictly the reliability is judged. That is, the reliability of the detection result of a predetermined state by the state detection unit is judged to be low. On the contrary, when the state type detected by the state detection unit 21 is a state type other than the predetermined state types (angry behavior, bowing, thrusting), the confidence threshold is set to 80% regardless of the congestion level. For this reason, when the state type detected by the state detection unit 21 is a state type other than the predetermined state types, the confidence threshold does not change depending on the congestion level.

[0030] Also, the confidence thresholds when the state types detected by the state detection unit 21 are bowing or thrusting are set higher as the congestion level is higher. Thus, the higher the congestion level, the more strictly the reliability is judged. However, in the case of bowing, special conditions such as Person 2 sitting hunched or crouching are required, so the risk of false detection is smaller compared to angry behavior. In the case of thrusting as well, conditions such as holding an object like a weapon and pointing it at the other person are required, so the risk of false detection is smaller compared to angry behavior. Therefore, if the confidence thresholds for bowing and thrusting are set to be the same as those for angry behavior, the determination of reliability may become overly strict and there is a risk of failing to detect bowing and thrusting.

[0031] For this reason, the confidence thresholds in the medium-congestion area and the high-congestion area when the state types detected by the state detection unit 21 are bowing or thrusting may be set to 85% and 88% which are smaller than those for angry behavior. Accordingly, when the area around Person 2 is a low-congestion area, if the confidence level is 80% or more, it is determined that there is no false detection. On the other hand, when it is a medium-congestion area or a high-congestion area, if the confidence level is less than 85% or 88% respectively, it is determined as a false detection. By setting the confidence threshold in this way, when the state type detected by the state detection unit 21 is an angry behavior, the reliability of the state type detected by the state detection unit 21 is judged more strictly than when the state type detected by the state detection unit 21 is a squat or a push. Note that violent behavior is an example of "the first state type representing the first dangerous behavior" described in the claims, and squatting or pushing is an example of "the second state type representing the second dangerous state" described in the claims.

[0032] The output determination unit 23 determines the output from the output unit 15 for the abnormal states among the state types detected by the state detection unit 21 according to the reliability determined by the output determination unit 23. For example, when the detected state type is an abnormal state and the confidence level of the state type is equal to or higher than the confidence threshold, the output determination unit 23 may output the state type from the output unit 15, and when the confidence level is less than the confidence threshold, the output of the state type may be prohibited.

[0033] Also, for example, the output determination unit 23 may determine the device to which the state type is output from the output unit 15 according to the determined reliability. For example, the output unit 15 may output the state type to an external device of the image processing device 10 by wired communication or wireless communication. The external device may be, for example, a center terminal or a local monitoring terminal used by a monitor who monitors the monitoring space 1, or may be a customer mobile terminal.

[0034] The output determination unit 23 may switch which of these terminal devices the state type is output to according to the reliability determined by the output determination unit 23. For example, when the confidence level of the state type is equal to or higher than the confidence threshold, the state type may be output to the center terminal or the local monitoring terminal, and when the confidence level of the state type is less than the confidence threshold, the state type may be output to the customer mobile terminal.

[0035] (Modification example) (1) Depending on the image position of Person 2, there is a risk that the detection accuracy of the state type by the state detection unit 21 will decrease. For example, in the directly below image area where the area near the directly below area of the camera 11 is captured, there is a risk that the detection accuracy of the state type by the state detection unit 21 will decrease.

[0036] FIG. 4(a) is a schematic diagram of an example of the optical axis OA direction and the imaging angle θf of the camera when the camera 11 is a single-direction type camera, and FIG. 4(b) is a schematic diagram of the directly below image area 3D in the captured image 3 of the single-direction type camera 11. When the camera 11 is a single-direction type camera, the vertical imaging angle θf of the camera 11 is about 70 degrees to 90 degrees, and it is installed above the monitoring space 1, and the optical axis OA direction is set to have a depression angle θd. In the captured image 3 of such a single-direction type camera 11, the directly below image area 3D is located at the lower end of the captured image 3.

[0037] FIG. 5(a) is a schematic diagram of an example of the optical axis OA direction and the imaging range of the camera when the camera 11 is an omnidirectional camera, and FIG. 5(b) is a schematic diagram of the directly below image area 4D in the captured image 4 of the omnidirectional camera 11. The hatched area RF in FIG. 5(a) indicates the imaging range of the omnidirectional camera 11. The omnidirectional camera 11 takes the entire 360-degree range as the imaging area, is installed above the monitoring space 1, and the optical axis OA direction is directed straight down. In the captured image 4 of such an omnidirectional camera 11, the image is circular as a whole, and the directly below image area 4D is located at the center of the captured image 4.

[0038] When Person 2 is captured in the directly below image areas 3D and 4D, the lower body is hidden by Person 2's own body, and the posture of Person 2 is likely to be misdetected. For example, when using a posture detection AI to detect a person's posture from the positions of each joint point, the posture detection AI estimates the joint points that cannot be seen and detects the posture including the estimated joint points. When Person 2 is captured in the directly below image areas 3D and 4D, it is often the case that the joint points of the lower body are estimated to be bent, and although the person is not in a prostrate position, it is misdetected as being in a prostrate position.

[0039] Therefore, when the state type detected by the state detection unit 21 is a predetermined state type and the person 2 appears in an area other than the immediate image areas 3D and 4D, the output determination unit 23 determines the reliability of the state type detected by the state detection unit 21 based on the congestion level detected by the congestion detection unit 22. When the person 2 appears in the immediate image areas 3D and 4D, the output determination unit 23 may determine the reliability of the state type detected by the detection unit 21 more strictly based on the congestion level detected by the congestion detection unit 22 than when the person 2 appears in an area other than the immediate image areas 3D and 4D. For example, the confidence thresholds when the person 2 appears in an area other than the immediate image areas 3D and 4D are 80%, 90%, and 95% for the low congestion area, medium congestion area, and high congestion area respectively, while the confidence thresholds when the person 2 appears in the immediate image areas 3D and 4D are 82%, 92%, and 97% for the low congestion area, medium congestion area, and high congestion area respectively.

[0040] (2) When the state detection unit 21 cannot continuously detect the same state type for a predetermined duration or more, the output determination unit 23 determines that the state type detected by the state detection unit 21 is a false detection. When the state detection unit 21 can continuously detect the same state type for a predetermined duration or more, the output determination unit 23 may determine that the state type detected by the state detection unit 21 is not a false detection. The output determination unit 23 may set a predetermined duration according to the congestion level around the person 2. For example, the higher the congestion level, the longer the predetermined duration may be set. For example, when the area around the person 2 is a low congestion area, medium congestion area, or high congestion area, the predetermined duration may be set to 2 seconds, 5 seconds, and 10 seconds respectively.

[0041] Similarly, the output determination unit 23 may output a false detection of the state type detected by the state detection unit 21 according to whether the ratio of the time during which the state detection unit 21 detects the same state type in a period of a predetermined length is less than a predetermined ratio threshold. The output determination unit 23 may set a predetermined ratio threshold according to the congestion level around the person 2. For example, the higher the congestion level, the larger the predetermined ratio threshold may be set. The predetermined duration and the predetermined ratio threshold are examples of the "determination criteria" described in the claims.

[0042] (3) As a method for the output determination unit 23 to determine the reliability of the state type detected by the state detection unit 21, the confidence level of the state type detected by the state detection unit 21 may be corrected. The output determination unit 23 may output the confidence level corrected by the output determination unit 23, or may switch the output destination device of the state type detected by the state detection unit 21 according to the confidence level corrected by the output determination unit 23. Also, a confidence level threshold for determining whether or not the state type detected by the state detection unit 21 is a false detection may be a common value for all state types, and it may be determined whether or not it is a false detection by comparing the corrected confidence level with the threshold.

[0043] In this case, the output determination unit 23 may correct the confidence level output by the state detection unit 21 according to the degree of congestion around person 2. For example, the higher the degree of congestion, the confidence level output by the state detection unit 21 may be corrected to be smaller. The correction value by which the output determination unit 23 corrects the reliability of the state type detected by the state detection unit 21 is an example of the "determination criterion" described in the claims.

[0044] (4) The output determination unit 23 may determine the reliability using the score obtained by the state detection unit 21. At this time, the state detection unit 21 outputs the obtained score to the output determination unit 23. For example, the higher the degree of congestion, the value of the score is made lower and compared with the threshold for determining the reliability, or the higher the degree of congestion, the threshold for determining the reliability is made higher and compared with the output score. The correction value by which the output determination unit 23 corrects the value obtained by the state detection unit 21 for state detection is an example of the "determination criterion" described in the claims.

[0045] Also, the state detection unit 21 may correct the obtained feature amount or score value according to the degree of congestion. The state detection unit 21 detecting the state using the degree of congestion is an example of the "determination criterion" described in the claims.

[0046] (5) The state detection unit 21 may correct the confidence level to be output to the output determination unit 23 for the state type detected by the state detection unit 21 according to the degree of congestion around person 2, and determine the reliability of the state type detected by the state detection unit 21. For example, the correction may be made such that the higher the degree of congestion, the smaller the confidence level output by the state detection unit 21. Since the output determination unit 23 determines whether to output using the corrected confidence level, the correction value corrected by the state detection unit 21 is an example of the "determination criterion" described in the claims.

[0047] (6) When the state detection unit 21 detects a state type from a plurality of frames in the input image, the state type detected based on the number of frames equal to or greater than a threshold value among a predetermined number of consecutive frames may be detected as the state type of person 2, and the reliability of the state type detected by the state detection unit 21 may be determined. In the following description, the threshold value of the number of frames required to detect the state type among a predetermined number of consecutive frames may be referred to as the "required number of frames threshold value".

[0048] For example, the required number of frames threshold value when the state detection unit 21 detects a predetermined state type may be set according to the degree of congestion around person 2. For example, the higher the degree of congestion, the larger the required number of frames threshold value may be set. FIG. 6 is a diagram showing an example of setting the required number of frames threshold value. The setting in FIG. 6 shows an example when detecting a state type from 15 consecutive frames (that is, when the "predetermined number" is 15). The required number of frames threshold value when detecting an angry behavior (quarreling, destruction) is set to 10 frames, 12 frames, and 14 frames when the area around person 2 is a low congestion area, a medium congestion area, and a high congestion area, respectively.

[0049] Thereby, when the area around person 2 is a low congestion area, if the state detection unit 21 detects an angry behavior in 10 or more frames out of 15 consecutive frames, the state detection unit 21 outputs the angry behavior to the output determination unit 23 as the action type of person 2. On the other hand, in the case of a medium congestion area and a high congestion area, the angry behavior is not output as the action type of person 2 unless it is 12 or more frames and 14 or more frames, respectively. Thereby, since the required number of frames threshold value is set so that it becomes more difficult to detect an angry behavior as the degree of congestion increases, false detection of an angry behavior in a situation with a high degree of congestion can be suppressed.

[0050] On the other hand, when detecting a behavior type other than the predetermined behavior types (i.e., angry behavior, bowing, and pushing), regardless of the degree of congestion, the required number of frames threshold is set to 10 frames. Therefore, when detecting a behavior type other than the predetermined behavior types, the required number of frames threshold does not change according to the degree of congestion.

[0051] Also, the required number of frames threshold for detecting bowing and pushing is set to be larger as the degree of congestion is higher. As a result, the required number of frames threshold is set so that it becomes more difficult to detect bowing and pushing as the degree of congestion increases, and thus false detection of bowing and pushing in a highly congested situation can be suppressed. On the other hand, if the required number of frames threshold for detecting bowing and pushing is set to the same size as the required number of frames threshold for angry behavior, the criteria for detecting bowing and pushing will become overly strict and there is a risk of failing to detect bowing and pushing. Therefore, the required number of frames threshold in the medium congestion area and high congestion area for bowing and pushing is set to 11 frames and 12 frames, which are smaller than those for angry behavior.

[0052] As a result, when the area around Person 2 is in the low congestion area, if the behavior detection unit 21 detects bowing or pushing in 10 or more frames out of 15 consecutive frames, it outputs these behaviors to the output determination unit 23 as the behavior type of Person 2. In contrast, when it is in the medium congestion area or high congestion area, it does not output bowing or pushing as the behavior type of Person 2 unless it is 11 frames or more and 12 frames or more, respectively. In order to determine whether the output determination unit 23 outputs using the result of determining the reliability of the behavior type according to the degree of congestion, the required number of frames threshold set by the behavior detection unit 21 according to the degree of congestion is an example of the "criteria" described in the claims.

[0053] (7) The state detection unit 21 may determine the reliability of the posture of person 2 detected by the posture detection AI according to the degree of congestion around person 2. That is, the state detection unit 21 may be a part of the output determination unit 23, and the state detection unit 21 may detect the state type of the person and determine the reliability of the detected state type based on the degree of congestion. As a method for determining the reliability of the detected posture, the state detection unit 21 may determine whether the posture detected by the posture detection AI is correct. When the detected posture is correct, the state detection unit 21 may output to the output determination unit 23 the action detected based on the time-series change information of the posture detected by the posture detection AI as the state type of person 2. In an embodiment where the state type output to the output determination unit 23 is a posture, the posture detected by the posture detection AI may be output to the output determination unit 23 as the state type of person 2.

[0054] When the detected posture is incorrect, detecting the action of person 2 using the posture detected by the posture detection AI and outputting the posture detected by the posture detection AI as the state type of person 2 are prohibited. At this time, the state detection unit 21 may determine whether the posture of person 2 detected by the posture detection AI is correct according to the angle and movement amount of the joints (skeleton) recognized from the image of person 2.

[0055] For example, when the number of times of swinging down the arm is equal to or greater than a threshold value, the state detection unit 21 may determine that the threatening posture detected by the posture detection AI is correct. For example, the state detection unit 21 may set a larger threshold value for the number of times of swinging down the arm as the degree of congestion increases. For example, the threshold values in the low congestion area, medium congestion area, and high congestion area may be set to 2 times, 4 times, and 8 times, respectively. Also, when the angle between the line connecting the thigh (waist) and the knee and the line connecting the knee and the foot is equal to or less than a threshold value, and when the state detection unit 21 determines that the crouched (seiza) posture detected by the posture detection AI is correct, the threshold value may be set smaller as the degree of congestion increases. By making the angle smaller, the posture of bending the knees and sitting on the floor can be accurately detected.

[0056] For example, the state detection unit 21 may determine the confidence level of the posture detected by the posture detection AI as a method for determining the reliability of the posture detected by the posture detection AI. For example, the state detection unit 21 may set the requirements for obtaining the same height of confidence level more strictly as the degree of congestion increases. For example, when the posture of person 2 detected by the posture detection AI is a threatening posture, the state detection unit 21 may determine the confidence level of the posture detected by the posture detection AI according to the number of times of swinging down the arm.

[0057] And the number of times for the confidence level to reach 100% in the low congestion area, medium congestion area, and high congestion area may be set to 2 times, 4 times, and 8 times respectively, for example. When the calculation method of the confidence level is determined in this way, for example, when the number of times of swinging down the arm is 2 times, the confidence level will be 100% in the low congestion area, 50% in the medium congestion area, and 25% in the high congestion area.

[0058] (Effect of the embodiment) (1) The image processing apparatus 10 includes an image acquisition unit 20 that acquires an input image captured by the camera 11, a state detection unit 21 that detects the state type of an object shown in the input image, a congestion detection unit 22 that detects the degree of congestion around the object, and an output determination unit 23 that determines the reliability of the detected state type based on the state type detected by the state detection unit 21 and the degree of congestion detected by the congestion detection unit 22. Thereby, it is possible to suppress a decrease in the detection accuracy of the state type of the object when the periphery of the object is congested.

[0059] (2) The output determination unit 23 may determine the reliability based on the degree of congestion only when the detected state type is a predetermined state type among a plurality of state types. Thereby, it is possible to suppress a decrease in the detection accuracy of a predetermined state type when the periphery of the object is congested. (3) The output determination unit 23 may determine the reliability of the detected state type based on a reliability determination criterion defined in advance according to the degree of congestion. Thereby, it becomes possible to determine the reliability according to the difference in the degree of congestion around the object, and it is possible to suppress a decrease in the detection accuracy of the state type due to the degree of congestion.

[0060] (4) When the detected state type is a predetermined state type, the determination criterion may be defined so that the higher the degree of congestion, the stricter the determination of reliability. Thereby, it is possible to suppress a decrease in the detection accuracy of the state type of the object when the degree of congestion is high. (5) The image processing apparatus 10 may detect the state type of a person as the object. The predetermined state type may be a state type detected on the condition that there is another person near the person who is the object. Thereby, it is possible to suppress a decrease in the detection accuracy of the state type detected on the condition that there is another person near the object due to the degree of congestion.

[0061] (6) The determination criterion may be defined such that the higher the degree of congestion, the stricter the determination of reliability in both the case where the detected state type is the first state type and the case where it is the second state type, and the determination of reliability is stricter in the case of the first state type than in the case of the second state type. Thereby, when there are a plurality of state types in which the higher the degree of congestion, the stricter the determination of reliability, it is possible to provide a degree of urgency of the determination criterion between these plurality of state types.

[0062] (7) The state detection unit 21 may include a posture detection unit that inputs the input image to a learning model generated by machine learning that uses the image data as input and outputs the posture of the object shown in the image of the image data as output data to estimate the posture of the object, and an action detection unit that detects the type of action of the object based on the detected change in posture and outputs the detected type of action as the state type. Thereby, when estimating the posture of the object using the learning model, it is possible to suppress a decrease in the detection accuracy of the state type of the object due to the degree of congestion around the object.

[0063] (8) The state detection unit 21 may detect the state type of the object based on the joint positions of the object. Thereby, when estimating the posture of the object based on the joint positions of the object, it is possible to suppress a decrease in the detection accuracy of the state type of the object due to the degree of congestion around the object. (9) The state detection unit 21 may detect the state type of the object based on the degree of congestion detected by the degree-of-congestion detection unit 22, or may calculate the confidence level of the state type of the object based on the degree of congestion. Thereby, it is possible to improve the detection accuracy of the state type and the calculation accuracy of the confidence level in the state detection unit 21.

[0064] The image processing apparatus according to an embodiment of the present invention can contribute to solving social problems such as a decrease in the working population.

Explanation of Reference Numerals

[0065] 1... Monitoring space, 2... Person (detection target), 10... Image processing apparatus, 11... Camera, 12... Input unit, 13... Storage unit, 14... Control unit, 15... Output unit, 20... Image acquisition unit, 21... State detection unit, 22... Degree-of-congestion detection unit, 23... Determination unit, 23... Output determination unit

Claims

1. An image acquisition unit that acquires an input image captured by a camera; A state detection unit that detects a predetermined state caused by a moving object appearing in the input image; A congestion degree detection unit that detects the congestion degree around the moving object; A determination unit that determines the reliability of the detection result of the predetermined state by the state detection unit based on the congestion degree detected by the congestion degree detection unit; An image processing apparatus, characterized by comprising the above.

2. The image processing apparatus according to claim 1, wherein the determination unit determines that the reliability is lower when the predetermined state is detected when the congestion degree is high than when it is low.

3. The state detection unit detects a plurality of state types caused by the moving object, The image processing apparatus according to claim 1 or 2, wherein the determination unit determines the reliability based on the congestion degree when a predetermined state type among the plurality of state types is detected.

4. The moving object is a person, The image processing apparatus according to claim 3, wherein the predetermined state type is a state type representing a dangerous act detected on the condition that a plurality of the persons are present nearby.

5. The image processing apparatus according to claim 3, wherein the determination unit varies the reliability determination criteria depending on whether the detected state type is a first state type representing a first dangerous act or a second state type representing a second dangerous act.

6. The moving object is a person, The image processing apparatus according to claim 1, wherein the state detection unit detects the posture or action of the person as the state based on the joint information of one or more of the persons.

7. The image processing apparatus according to claim 1, wherein the state detection unit inputs the input image into a learning model generated by machine learning that takes image data as input and outputs a predetermined state caused by the moving object appearing in the image data as output data, and outputs the predetermined state caused by the moving object.

8. An image acquisition process for acquiring an input image captured by a camera; A state detection process for detecting a predetermined state caused by a moving object appearing in the input image; A congestion degree detection process for detecting the congestion degree around the moving object; A determination process for determining the reliability of the detection result of the predetermined state by the state detection process based on the congestion degree detected by the congestion degree detection process; A computer program, characterized by causing a computer to execute the above.

Citation Information

Patent Citations

  • Action analysis system and action analysis method

    JP2023030965A