Image analysis system and image analysis method

The video analysis system reduces processing load and enhances recognition accuracy by identifying candidates for monitoring based on motion analysis and selectively applying action recognition models, addressing the challenge of high processing loads in existing systems.

JP7710369B2Active Publication Date: 2025-07-18HITACHI LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021213802
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-28
Publication Date
2025-07-18
Estimated Expiration
2041-12-28

AI Technical Summary

Technical Problem

Existing video monitoring systems face high processing loads due to the need to perform action recognition on all detected individuals, which increases with the number of surveillance cameras, making it difficult to enhance the analysis server capacity.

Method used

A video analysis system that detects candidates for monitoring by calculating motion amounts and using pre-set conditions to select appropriate action recognition models for these candidates, reducing the need for full-scale action recognition on all individuals.

Benefits of technology

This approach reduces the processing load and improves recognition accuracy by selectively applying action recognition only to candidates, thereby optimizing system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007710369000001
    Figure 0007710369000001
  • Figure 0007710369000002
    Figure 0007710369000002
  • Figure 0007710369000003
    Figure 0007710369000003
Patent Text Reader

Abstract

To reduce a processing load on a video monitoring system.SOLUTION: A video analysis system detects an event in a region being monitored, by using a video of the region being monitored. A storage device stores a preset motion amount condition for defining the range of a motion amount. A computation deice calculates the motion amount of a person detected on the basis of the video, determines whether or not the person is a candidate to be monitored on the basis of the calculated motion amount and the motion amount condition, selects an action recognition model on the basis of the motion amount and the motion amount condition, and detects an event concerning the person determined as a candidate to be monitored, by using the selected action recognition model.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a video analysis system and a video analysis method, and more particularly, to a video analysis system and a video analysis method for detecting a person from a video captured in a monitoring area and detecting a monitoring target based on the detection result.

Background Art

[0002] In public facilities such as stations and airports, and in concert halls and amusement facilities, etc., it is necessary to quickly detect and deal with terrorist acts and dangerous acts, etc. in order to ensure the safety of users. Since the number of monitors and security guards deployed on-site is limited, the demand for video monitoring by surveillance cameras is increasing. However, with the reduction in cost, miniaturization, power saving, and high resolution of videos of surveillance cameras, while the trend of increasing the number of surveillance cameras continues, it becomes necessary to monitor a huge amount of videos with limited personnel.

[0003] Therefore, in order to achieve effective video monitoring and labor saving for administrators, automation of video monitoring is required. In particular, the technology for recognizing human behavior is one of the important technologies as an automatic monitoring technology for ensuring safety and security within a facility. For example, by detecting the early fall or huddled movement of a person being photographed, the facility manager can quickly protect the person in need of first aid who has occurred within the facility. In addition, by detecting running or violent acts at an early stage, it becomes possible to contribute to maintaining security within the facility.

[0004] On the one hand, with the increase in the number of surveillance cameras, it is necessary to enhance the analysis server for video analysis. Therefore, when enhancement is difficult, a reduction in the analysis processing volume is required. On the other hand, for example, in the information processing apparatus described in Patent Document 1, for the purpose of reducing the analysis processing volume, detection means for analyzing the input image and detecting a person included in the image, person determination means for determining whether the detected person is a pre-registered person, action determination means for determining whether a person determined by the person determination means not to be the pre-registered person has performed a first action seeking assistance, and output means for outputting a notification regarding the first person determined by the action determination means to have performed the first action to an external device are provided, and the detection means is characterized in that it executes further image analysis related to the first person determined by the action determination means to have performed the first action. An information processing apparatus is disclosed.

Prior Art Documents

Patent Documents

[0005]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0006] In the above-mentioned conventional technology, for a person detected in an image and having performed a predetermined action, the analysis processing volume is reduced by performing further image analysis. However, since the action recognition in the action determination unit needs to be performed for all the photographed persons, the analysis processing volume of the action recognition itself is not reduced. With the increase in the number of photographed persons, the analysis processing volume of the action recognition also increases accordingly. Therefore, when it is difficult to enhance the analysis server, a reduction in the processing volume of the action recognition is required.

Means for Solving the Problems

[0007] As one aspect of the present invention, a video analysis system uses a video captured of a monitoring area to detect an event in the monitoring area. The video analysis system includes one or more computing devices and one or more storage devices that store preset motion amount conditions that define a range of motion amounts. The one or more computing devices calculate the motion amount of a person detected based on the video, determine whether the person is a candidate for monitoring based on the calculated motion amount and the motion amount conditions, select an action recognition model based on the motion amount and the motion amount conditions, and use the selected action recognition model to detect an event for the person determined to be a candidate for monitoring.

[0008] As one aspect of the present invention, a video analysis method uses a video captured of a monitoring area to detect an event in the monitoring area. The video analysis method includes the video analysis system calculating the motion amount of a person detected based on the video, determining whether the person is a candidate for monitoring based on the calculated motion amount and preset motion amount conditions that define a range of motion amounts, selecting an action recognition model based on the motion amount and the motion amount conditions, and using the selected action recognition model to detect an event for the person determined to be a candidate for monitoring.

Advantages of the Invention

[0009] According to one aspect of the present invention, a reduction in the processing load of a video monitoring system is realized.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Embodiments for Carrying Out the Invention

[0011] Hereinafter, embodiments of the video surveillance system according to the present invention will be described. In one embodiment of this specification, for the purpose of reducing the processing load of the video analysis system, the action recognition process is multi-staged, and according to the result of the pre-stage process, an appropriate recognition model is selected to perform the action recognition process.

[0012] According to one embodiment of this specification, for a person detected within the screen, as a pre-stage process for determining whether to perform an action recognition process, an operation amount estimation that is less computationally intensive and faster than the action recognition process is performed, and it is determined whether to perform the action recognition process according to the operation amount and a predetermined condition set in advance. As a result, the number of times of the action recognition process can be reduced, and the processing load of the video surveillance system can be reduced. Also, by selecting an action recognition model according to the operation amount, it is possible to improve the recognition accuracy.

[0013] Note that the "event" in the present embodiment is a situation preset as a detection target in a certain monitoring area. In particular, in the present embodiment, the important attention actions of a person are the detection targets. For example, actions such as falling, squatting, violence, and running are included. Also, the "candidate for person to be monitored" refers to a person who has reached a predetermined condition set in advance regarding the operation amount as a result of the operation amount estimation, and the "person to be monitored" refers to a person who has been determined to have performed a specific action as a result of the action recognition.

[0014] In one embodiment of the present specification, by performing action recognition processing only on the candidate persons to be monitored, it is possible to reduce the processing load of the system compared to a conventional video monitoring system that performs action recognition processing on all persons. Also, the "user" refers to a person who can access the video monitoring system or operate the system settings, such as the administrator, monitor, or on-site security guard of the space where the monitoring camera is installed. Hereinafter, the embodiment will be described with reference to the drawings.

[0015] FIG. 1 is an explanatory diagram of a video monitoring system according to an embodiment of the present specification. As shown in FIG. 1, the video monitoring system 1 includes a photographing system 2, a video analysis system 3, and a monitoring center system 4. The photographing system 2 includes one or more camera units 13 installed in the monitoring target area 12.

[0016] Also, in the video analysis system 3, by analyzing the input video from the imaging device, a person photographed in the monitoring area is detected, the amount of movement of the person is calculated, and it is determined whether to execute action recognition processing by comparing the amount of movement with a predetermined condition preset regarding the amount of movement. At the time of execution determination, an appropriate action recognition model is selected from the amount of movement to execute the action recognition processing.

[0017] In the monitoring center system 4, the analysis result from the video analysis system 3 is received, an effective display for users such as the monitor 14 and the security guard is performed, and a predetermined condition preset regarding the amount of movement is transmitted to the video analysis system 3.

[0018] Hereinafter, the photographing system 2, the video analysis system 3, and the monitoring center system 4 will be described.

[0019] FIG. 2 is a diagram showing the overall configuration of a video surveillance system according to an embodiment of the present specification. The imaging system 2 includes one or more camera units 21 installed in a monitoring target area, and the captured video is sequentially input to the video input unit 31 of the video analysis system 3. The camera unit 21 is a monitoring camera arranged so as to be able to image the entire area to be monitored. The type of the monitoring camera is not limited to its form, such as a fixed camera, a PTZ camera capable of pan-tilt-zoom (PTZ) operation, or a mobile camera such as a drone-mounted camera or a wearable camera.

[0020] In the case of a PTZ camera or a mobile camera, it is important to prevent the movement of the camera from affecting the subsequent calculation of the amount of movement and action recognition. For example, preprocessing such as performing camera calibration in advance and estimating the position of a person in the world coordinate system is executed. In addition, it is not necessary to perform video analysis on the entire area of the image obtained by the monitoring camera, and limiting it to a partial monitoring area by mask processing also contributes to reducing the amount of analysis processing.

[0021] Further, the camera unit 21 and the video input unit 31 are connected by wired communication means or wireless communication means, and frame images are continuously transmitted from the camera unit 21 to the video input unit 31. When using a time-series data analysis model that assumes the input of a plurality of frame images for calculating the amount of movement and / or action recognition, it is desirable that the frame rate of the continuous transmission of the frame images is equal to or higher than the required values for calculating the amount of movement and action recognition. However, if the decrease in recognition accuracy caused by the frame rate dropping below the required value can be tolerated, the frame rate may be lower than the required value.

[0022] In this case, in the calculation of the amount of motion and action recognition, processes for suppressing a decrease in accuracy, such as interpolation by interpolation or extrapolation of time-series data, may be performed. Also, the camera unit 21 and the video analysis system 3 do not necessarily have a one-to-one correspondence, and one video analysis system may receive video data from a plurality of camera units. The video analysis system processes the video data from each of the plurality of camera units. Even in the case of executing such a multi-process, the frame rate from each camera unit required for each process conforms to the above-described constraints.

[0023] Note that the camera unit 21 may incorporate some or all of the functions of the video analysis system described later. For example, the camera unit 21 includes one or more arithmetic units and one or more storage units, performs edge processing for calculating the amount of motion, and transmits only information on candidates for persons to be monitored to the video analysis system, thereby reducing the processing load on the video analysis system.

[0024] The video analysis system 3 includes a video input unit 31, a determination unit 32, a model selection unit 33, an action recognition unit 34, an output control unit 35, and a storage unit 36. The video input unit 31 receives input of video from the camera unit 21 and transmits the video data to the determination unit 32. Note that the video to be analyzed may be a video stored separately in a recorder instead of the video directly input from the camera unit 21, and the storage location of the video does not matter. The determination unit 32 performs person detection on the persons in the screen, and performs a motion amount estimation that is faster and has a smaller amount of calculation than the action recognition process as a preprocessing for determining whether to perform the action recognition process. It has a function of determining whether to perform the action recognition process on the person as a candidate for a person to be monitored according to the estimated amount of motion per unit time and a predetermined condition preset for the amount of motion.

[0025] In the model selection unit 33, an action recognition model is selected according to the amount of movement and a predetermined condition regarding the amount of movement preset in the storage unit 36. The action recognition unit 34 performs action recognition processing using the model selected in the model selection unit 33 and the image information around the candidate person to be monitored. When the result of the action recognition satisfies a predetermined condition, the action recognition unit 34 determines that the person is a person to be monitored. Note that in the present embodiment, the video analysis system 3 is not limited to an on-premises type system constructed on a server within the operation facility, and may be constructed on an external server of the facility such as by utilizing cloud services.

[0026] The storage unit 36 has importance information that can be set by a user such as a supervisor or a security guard regarding the action type and / or the occurrence location. The importance information indicates the importance of each of the action type and / or the occurrence location. The storage unit 36 calculates the monitoring importance for each person who caused the event based on the importance information. In the output control unit 35, highlighting processing on the display terminal is performed according to the monitoring importance.

[0027] The monitoring center system 4 includes a recording unit 41, a video display unit 42, and a management control unit 43. The recording unit 41 has a function of holding, as a database, information such as the image information, movement trajectory, events caused by the person, person attributes, occurrence area, and occurrence time of the candidate person to be monitored or the person to be monitored obtained by video analysis by the video analysis system 3.

[0028] The video display unit 42 displays the actions of the candidate person to be monitored or the person to be monitored at the current time and information regarding a part or all of the frames at the time of event occurrence. The management control unit 43 has a function of inputting setting information by the user in order to store the predetermined condition regarding the amount of movement used in the determination unit 32 and the importance information in the storage unit 36.

[0029] Figure 3 is a hardware configuration diagram of a video surveillance system according to an embodiment of the present specification. In Figure 3, a camera unit 51 is connected to a computer 52 via a network. The computer 52 can communicate with a computer 53 via the network. Further, a computer 54 can communicate with the computer 53 via the network. The computers 52, 53, and 54 each include one or more arithmetic units and one or more storage devices.

[0030] The camera unit 51 functions as, for example, the photographing system 2 or the camera unit 21, and the computer 52 functions as the video analysis system 3. The computer 53 functions as, for example, the monitoring center system 4, and the computer 54 functions as a user terminal used by a user (for example, a monitor).

[0031] One or more camera units 51 are installed in the monitoring area and appropriately transmit video data to the computer 52. The computer 52 includes a CPU (Central Processing Unit) 521 as an arithmetic unit, a RAM (Random Access Memory) 522 as a main storage device, an HDD (Hard Disk Drive) 523 as an auxiliary storage device, and a communication interface (IF) 524.

[0032] The computer 52 reads various programs from the HDD 523, expands them in the RAM 522, and executes them by the CPU 521 to realize functions 31 to 36 as the video analysis system 3. The computer 52 also communicates with the camera unit 51 and the computer 53 via a predetermined communication interface 524. Although not shown, input / output devices such as a keyboard and a display may also be connected to the computer 52 via a predetermined IF.

[0033] The computer 53 includes a CPU 531 as an arithmetic unit, a RAM 532 as a main memory device, an HDD 533 as an auxiliary storage device, and a communication interface 524. The computer 53 reads various programs from the HDD 533, expands them in the RAM 532, and executes them by the CPU 531 to realize functions 41 to 43 as the monitoring center system 4. Further, the computer 53 is connected to the computer 52 via a predetermined interface 534. Although not shown, input / output devices such as a keyboard and a display may also be connected to the computer 53 via a predetermined IF.

[0034] The computer 54 includes a CPU 541 as an arithmetic unit, a RAM 542 as a main memory device, an HDD 543 as an auxiliary storage device, and a communication interface 544. The computer 54 reads various programs from the HDD 543, expands them in the RAM 542, and executes them by the CPU 541 to realize the function as a user terminal. The computer 54 further includes an input / output device (I / O) 545. The input / output device 545 receives an input from the user and further presents the monitoring result to the user. Note that instead of the computer 54, an input / output device directly connected to the computer 53 may be used.

[0035] Note that part or all of the processing of the video analysis system 3 may be processed on the monitoring camera side. In that case, the monitoring camera has a configuration including part or all of the hardware of the camera unit 51 and the computer 52.

[0036] Next, with reference to FIG. 4, the details of the video analysis system 3 will be described. FIG. 4 is a block diagram of the video analysis system according to an embodiment of the present specification. Hereinafter, the video input unit 31, determination unit 32, model selection unit 33, action recognition unit 34, output control unit 35, and storage unit 36 that constitute the video analysis system 3 will be described.

[0037] The video input unit 31 sequentially receives videos from one or more camera units 21 and outputs the videos to the subsequent determination unit 32. The determination unit 32 quantifies the degree of the temporal change in the state of the target person, which is the amount of movement. For example, the amount of movement is calculated from a plurality of images (video frames) that are temporally continuous. When the action recognition unit 34 does not handle temporal information, the input to the action recognition unit 34 may be an image.

[0038] The determination unit 32 includes a person detection unit 321, a movement amount calculation unit 322, and a candidate determination unit 323. The person detection unit 321 detects a person from the still image of the current frame using the image or video (multiple frames) received from the video input unit 31. As means for person detection, there are means for determination by HOG (Histogram of Oriented Gradients), R-CNN (Regions with CNN), YOLO (You Only Look Once), etc., and means for determining an estimation region from a group of skeleton coordinates estimated for each person using skeleton estimation means. In the present embodiment, any of these means may be used.

[0039] When temporal information is handled in either the movement amount calculation unit 322 or the action recognition unit 34, or when continuous capture of the current position of a person to be monitored or a person to be closely monitored is performed as shown in FIG. 8 described later, person tracking is also performed together. The movement amount calculation unit 322 calculates the movement amount of the detected person. A typical means for calculating the movement amount is optical flow. With optical flow, the movement direction and its flow intensity can be calculated.

[0040] For example, by combining the output of the person detection unit 321 with an optical flow method such as TV-L1 or Farneback, a dense optical flow within the person rectangle can be calculated. The movement amount may be calculated by other methods. For example, the movement amount per unit time may be calculated from the amount of movement per unit time at one or more specific positions of the person.

[0041] In the candidate determination unit 323, it is determined whether to perform the action recognition process according to the amount of motion and the motion amount condition information 361 of the amount of motion preset in the storage unit 36. As the determination index, for example, the average flow intensity of all pixels within the person rectangle or the average intensity of pixels whose flow intensity is in the top predetermined percentage (for example, several tens of percent) can be used.

[0042] For simplicity, for example, in the subsequent model selection unit 33, two types of recognition models can be selected, and consider the case where the upper threshold and the lower threshold are set in the motion amount condition information 361 as the motion amount condition information. In this case, the user sets the upper and lower thresholds to optimal values through prior video analysis using the above index. When exceeding one-sided threshold continuously for a predetermined number of frames of 1 or more, the candidate determination unit 323 sets the person as a candidate person to be monitored.

[0043] In the above, an example using two types of recognition models is shown, so two types of thresholds, the upper threshold and the lower threshold, are set. However, when using three or more types of recognition models, a plurality of thresholds are set. Also, the threshold may be a vector instead of a scalar. For example, the flow intensity can also be captured by decomposing it in the horizontal or vertical direction of the image.

[0044] Note that person tracking may be performed to trace the movement trajectory of the candidate person to be monitored. In person tracking, it is sufficient that the rectangular image of a person and the person ID assigned to the person are associated in the front and rear frames. In addition to using a general person tracking method represented by template matching, the result of the optical flow used in the previous motion amount calculation unit may also be used.

[0045] In the model selection unit 33, an action recognition model is selected according to the action amount calculated by the action amount calculation unit 322 and the action amount condition information 361. Through the above adaptive model selection, an improvement in recognition accuracy is also achieved. For example, for a person moving at high speed, it is not necessary to use a recognition model that outputs action types in a stationary state, such as falling or crouching. Therefore, by selecting a model that can recognize only actions involving high-speed movements, such as running and throwing, the probability of false detection can be reduced.

[0046] In the action recognition unit 34, action recognition processing is performed using the model selected by the model selection unit 33 and the image information around the person to be monitored. Not only image information, but also information related to action amounts such as optical flow, and attribute information of the person specified from the image may be used. Also, feature amounts related to the skeleton of the person may be calculated by estimating the skeleton of the image, and an action recognition method based on the skeleton may be used.

[0047] When using image information, means such as using HoG features or CNN features can be mentioned. When using skeleton information, means such as training SVM (Support Vector Machine), decision trees, RNN (Recurrent Neural Network), or GCN (Graph Convolutional Network) with feature amounts representing postures, moving speeds, etc. can be mentioned.

[0048] Also, these recognition models may not only be used alone, but may also be used in combination. The details of the recognition model are not limited in this embodiment. A person to be monitored candidate who is determined to have performed a specific action is set as a person to be monitored, and information such as images and actions around the person is transmitted to the output control unit 35.

[0049] In addition to the operation amount condition information 361, the memory unit 36 has importance information that can be set by the user for the action type and / or the occurrence location as monitoring reference information 362, and calculates the monitoring importance for each person who caused the event based on the importance information. The output control unit 35 performs processing for emphasizing and displaying the person to be monitored on the display terminal according to the monitoring importance.

[0050] FIG. 5 is an explanatory diagram of the determination unit in an embodiment of the present specification. FIG. 5 shows images at two times taken by a certain monitoring camera. Image 61 is an image at time t, and image 62 is an image at time t + i (i is a natural number). When i = 1, it indicates that the two images are adjacent frames.

[0051] In image 61, three people are photographed in regions 611, 612, and 613. Similarly, in image 62, three people are photographed in regions 621, 622, and 623. Here, the people in region 611 and region 621, the people in region 612 and region 622, and the people in region 613 and region 623 are the same person. That is, for example, it indicates that the person in region 611 has moved to region 621 after i seconds. In addition, the region indicated by the dotted line indicates a person who is not set as a candidate for the person to be monitored, and the region indicated by the solid line indicates a person who is set as a candidate for the person to be monitored.

[0052] Here, as described above, consider the case where two types of recognition models can be selected for simplicity. At this time, a person who exceeds the upper threshold value of the operation amount and a person who is below the lower threshold value are set as candidates for the person to be monitored, and action recognition processing will be performed. In FIG. 5, since the person in region 611 did not reach any of the threshold values during the movement to region 621, the person is not set as a candidate for the person to be monitored.

[0053] On the other hand, since it is determined that the person in area 612 exceeded the upper threshold of the amount of movement by i seconds and the person in area 613 fell below the lower threshold of the amount of movement by i seconds, it indicates that the person is set as a candidate for person to be monitored. However, for the sake of robustness of the target determination, it is also preferable to set a person to be monitored when either threshold is continuously exceeded for a certain number of frames or when a predetermined percentage of the number of frames among a predetermined number of frames exceeds the threshold.

[0054] Next, with reference to the flowchart shown in FIG. 6, the processing flow of the video analysis system in one embodiment of this specification will be described. When video is input from the imaging system 2 to the video analysis system 3 in step S1, in step S2, person detection is performed by the person detection unit 321. If a person is detected within the monitoring area (S2: YES), in steps 3 to 9, processing is performed for each detected person. If no person is detected (S2: NO), the process returns to step 1 to read the next frame image.

[0055] In the processing after step 3, when the amount of movement calculation or behavior recognition deals with time-series information, the person detection unit 321 first performs person tracking in step S4. At this time, since it is necessary to hold information of a plurality of frames, it is desirable to manage the image information and the amount of movement information around the person by person ID in a memory, a storage, or the like. Next, in step S5, the amount of movement calculation unit 322 calculates the amount of movement of the person. In step S6, when it is determined by the candidate determination unit 323 that the calculated amount of movement has reached a preset amount of movement condition, in order to perform behavior recognition processing, in the subsequent step S7, the model selection unit 33 selects a recognition model.

[0056] Here, for those who do not reach the conditions, the processing after step S7 is not performed, and the processing is restarted from step S4 for the next person. In step S8, the action recognition unit 34 performs action recognition using the selected recognition model and the peripheral image of the person. The calculated amount of motion may be used as a feature amount. Finally, in step S10, as output control to the monitoring center system, the output control unit 35 transmits the person ID assigned by person tracking, the image coordinates of the person to be monitored, etc. to the monitoring center system 4. After step S10, the process returns to step S1 to read the next frame image.

[0057] Note that the processing of the flow shown in this figure does not necessarily have to be processed by a single process, and may be processed asynchronously using a plurality of processes for improving the calculation efficiency.

[0058] Next, referring to FIG. 7, an example of a setting screen for the amount-of-motion condition information in the present embodiment is shown. FIG. 7 is a GUI (Graphical User Interface) displayed by the management control unit 43 of the monitoring center system 4 in an embodiment of this specification, and can be controlled by the user. In area 71, a graph is shown with the horizontal axis representing the amount of motion and the vertical axis representing the probability density. By setting a threshold according to the probability density, a more appropriate determination becomes possible.

[0059] In this example, the scalar flow intensity is assumed as the amount of motion, but settings regarding the flow direction etc. may be further included. The function 711 represents the probability density function in a certain action. In area 71, five types of actions from action A to action E are prepared as learning data, and it is shown that each has a unimodal probability density function for simplicity. These probability density functions can be obtained, for example, by calculating in advance using the average flow intensity of all pixels within the person rectangle or the average intensity of pixels in the top several tens % in terms of flow intensity. Or, it may be calculated using the average value or the median value of a predetermined number of frames.

[0060] In addition, it is necessary to create an action recognition model in advance according to the function group. For example, it can be read that actions A and B have less movement amounts compared to the other three types of actions. Recognition model A for detecting these actions with small movement amounts and, similarly, recognition model B for detecting actions with large movement amounts are created in advance. Examples of actions with small movement amounts include "falling", "crouching", "looking around the surroundings", etc. On the other hand, examples of actions with large movement amounts include "running fast", "violence", "throwing", etc.

[0061] Here, as described above, two types of action recognition models are created, and it is assumed that the movement amount is given as a scalar. In area 71, threshold A is prepared as the lower threshold and threshold B is prepared as the upper threshold. The user can set each threshold while referring to the function illustrated in area 71. For example, when operating threshold A on the graph, the threshold can be changed by sliding the threshold left and right while pressing area 712.

[0062] Alternatively, as shown in Table 72, the coverage rate (cumulative density) for each action that varies according to the threshold setting may be displayed. The user can set the threshold while referring to Table 72. The "action type" in Table 72 is the action that the corresponding action recognition model can recognize. In the "coverage rate" of the same table, for example, in action recognition "model A", information is displayed that actions with the currently set "<threshold A", that is, actions with a movement amount smaller than threshold A, cover 99% of action A and 90% of action B. By creating and displaying Table 72 in conjunction with area 71, the user can set the threshold based on quantitative recognition.

[0063] In the case of the example shown in FIG. 7, action A and action B are learned by action recognition model A, action D and action E are learned by action recognition model B, and action C has not been learned by any of the models. By setting the threshold so that action C is excluded, it is expected that the overall video analysis process can be speeded up. In addition, the ratio of being erroneously estimated as the "action C" in the action recognition result can be reduced, and an effect on improving the accuracy can also be expected. When using three or more types of recognition models, three or more types of thresholds are similarly required. The above setting information is stored in the operation amount condition information 361.

[0064] Next, with reference to FIG. 8, an example of displaying the person to be monitored in FIG. 5 will be described. FIG. 8 is a GUI displayed on the video display unit 42 in an embodiment of this specification. Region 8 indicates the output screen, and region 8 may be displayed across the entire notification screen or a part of the notification screen. Among regions 8, the persons displayed on screen A (region 81), screen B (region 82), screen C (region 83), and screen D (region 84) are persons detected as having performed an action learned by any of the action recognition models as a result of action recognition.

[0065] In FIG. 8, four display screens are illustrated, but the number and arrangement method thereof are not limited. FIG. 8 shows an example of displaying, for all persons on each screen, an image at the time of action detection. The image may be a video of several seconds before and after detection. In that case, a form in which the video is automatically looped and played is also suitable for improving the efficiency of situation awareness. Each of regions 81 to 84 may display the entire frame image, or may display an image or video trimmed within a certain range so that only the surrounding situation of the detected person can be seen.

[0066] In order to facilitate the identification of other persons photographed in the same image, the person to be monitored may be surrounded by a rectangular frame as shown in the figure. When performing person tracking within the video surveillance system, if the current position of the person to be monitored can be captured, it may be switched to a form that can display the position of the person to be monitored at the current position and / or a real-time image or video. In that case, in the candidate determination unit 323, information such as the appearance, attributes, and actions of the person determined to be the person to be monitored may be separately stored in the storage unit 36 as the person to be monitored information. The information on the attributes of a person may include, for example, gender, age, race, etc., the information on appearance may include information such as clothing and hairstyle, and the information on actions may include the types of actions. Thereby, the user can easily obtain information about the person to be monitored.

[0067] The storage destination of the person to be monitored information is not limited, such as a cloud or a storage server prepared within the monitoring center system. When occlusion occurs to the person to be monitored or when the person to be monitored frames out of the monitoring area, it is also preferable to present information regarding the last captured location and time.

[0068] In area 85, detailed information on the above four areas 81 to 84 or detailed information on all detection results is displayed in text form. Specifically, it includes the displayed screen, monitoring importance, the current position of the person to be monitored, the type of event, its detection time, and the occurrence location, etc. Here, the monitoring importance refers to the weight preset for the event. For example, by making a preset in advance that a fall is more important than squatting down, sorting in the order of monitoring importance becomes possible.

[0069] It is also possible to set importance levels for the occurrence locations or to combine the scores of the occurrence locations and the detected events. These pieces of information are set by the management control unit 43 as the monitoring standard information 362. Also, on the display screen such as the area 81, it is possible to change the screen size in the order of monitoring importance levels, and the more important the event is, the larger it is displayed. In addition to the highlighting by size, the display screen of an image whose importance level exceeds a predetermined value may be highlighted by being placed at a prominent position or by coloring the thick frame or the frame. In this way, the image may be highlighted according to the importance level, and the size may be changed according to the importance level, or the highlighting mode of combination according to the importance level may be changed.

[0070] As described above, it is effective in improving the user's situation judgment and responsiveness, and enables the efficiency of video monitoring. As an example of setting the monitoring standard information, regarding the occurrence location, the importance level of the concourse is set to 3 points, and the importance level of the parking lot is set to 1 point. On the other hand, regarding the action type, the importance level of a fall is set to 2 points, and the importance level of violence is set to 3 points. At this time, the monitoring importance level can be calculated based on the sum or product of the occurrence location and the action type.

[0071] When calculating the monitoring importance level by product in the above point setting, the event A of "a fall occurred in the concourse" is 6 points, and the event B of "violence occurred in the parking lot" is 3 points. As a result, the monitoring importance level of event A is set higher than that of event B, and in the example of FIG. 8, the area 81 is displayed on a larger screen than the area 82. Due to the limitation of the screen size, when a plurality of events occur, the screen may be scrollable.

[0072] When the response by the staff to the event is completed, when it is determined that no response is required, or when it is obvious that it is a false detection, etc., it is preferable that the user can select and delete the row of the display screen or the area 85. FIG. 8 illustrates a state in which the event B of "violence occurred in the parking lot" is selected by pressing the check box in the second row from the left end of the area 85 or the area 82. When the response to this event is completed, etc., pressing the area 86 while this event is selected shows that it can be deleted from the display.

[0073] Since events are assumed to occur one after another over time, in order to improve the readability of the display screen, events for which no user processing has been performed within a predetermined time from the detection time may be automatically deleted from the display. However, for particularly important events, processing may be performed to reduce the user's oversight, such as issuing a warning after a certain period of time has elapsed since the detection of the event or before the event is deleted.

[0074] This screen can be viewed not only by reporting targets such as surveillance center monitors who are assumed to use large displays, but also by on-site staff and security guards who can view part or all of area 8 on-site by using a smartphone terminal, a tablet terminal, or an AR goggles, etc.

[0075] As described above, the video surveillance system 1 aims to reduce the processing load of the video analysis system by multi-staging the action recognition process and adaptively selecting a recognition model according to the result of the pre-stage process to perform the action recognition process.

[0076] According to the present invention, as a pre-stage process for determining whether to perform an action recognition process on a person detected in the screen, a faster operation amount estimation than the action recognition process is performed, and according to a predetermined condition regarding the operation amount and a preset operation amount, it is possible to determine whether to perform the action recognition process, thereby reducing the number of times of the action recognition process and reducing the processing load of the video surveillance system. In addition, by selecting an action recognition model according to the operation amount, it is also possible to improve the recognition accuracy.

[0077] Hereinafter, another embodiment of the video surveillance system 1 according to the present invention will be described. Note that descriptions of parts common to the above-described embodiment will be omitted, and mainly the specific processes in this embodiment will be described. In the above-described example, as shown in FIG. 7, the probability density function for each action regarding the operation amount was calculated in advance, and the user set a predetermined condition regarding the operation amount based on that.

[0078] However, it is assumed that the distribution of the above function changes according to the type of facility where the surveillance camera is installed. Taking actions with a relatively large amount of movement, such as running, throwing, and violence, as an example, it is assumed that the amount of movement in a situation where the density of people occupying the space is large and the amount of movement in a situation where the density is small, the latter amount of movement will be larger.

[0079] That is, if the probability density function obtained from the learning data captured in a situation where the density of people occupying the space is large is used in a situation where the density is small, the threshold will be set at a position farther from the actual distribution, so the effect of reducing the analysis processing amount may be reduced. Therefore, a process of optimizing the distribution and threshold for each facility where the surveillance camera is installed or for each surveillance camera is considered.

[0080] FIG. 9 is a block diagram of a video analysis system according to an embodiment of the present specification. Compared with the block diagram of the video analysis system in the embodiment shown in FIG. 1, a condition calculation unit 37 is added. The action type output by the action recognition unit 34 is sent to the condition calculation unit 37 together with the amount of movement, which is the output result of the amount of movement calculation unit 322.

[0081] In the condition calculation unit 37, using the newly calculated amount of movement and the result of the action type, the probability density stored in the storage unit is updated, and the threshold is automatically adjusted to satisfy the set coverage rate condition. Referring to the setting situation in FIG. 7 as an example, for actions D and E that can be recognized by the action recognition model B, it is assumed that the distribution related to action E has been changed. Changing the threshold B also affects the coverage rate of the distribution related to action D. Therefore, in this case, the condition calculation unit 37 may adjust the threshold B so that the average value or median value of the conventional coverage rates of actions D and E is the same as the value before the update. Or, a process such as outputting a message prompting the user for readjustment may be performed.

[0082] When a certain amount of data is accumulated at the site where the surveillance camera is in operation and the probability density function can be calculated for each action, there is no need to reuse the data used to calculate the prior distribution, and the data at the operation site, similar facilities, or with a similar viewing angle can be used. The data of similar facilities or with a similar viewing angle has been previously associated with the operation site by the user.

[0083] Note that the present invention is not limited to the above-described embodiments, and various modifications are included. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and are not necessarily limited to those having all the configurations described.

[0084] Also, a part of the configuration of one embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can also be added to the configuration of one embodiment. Further, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations is possible. Also, the above-described respective configurations, functions, processing units, processing means, etc. may be realized in hardware, for example, by designing a part or all of them with an integrated circuit.

[0085] Also, the above-described respective configurations, functions, etc. may be realized in software by a processor interpreting and executing a program for realizing each function. Information such as a program, table, file, etc. for realizing each function can be placed in a memory, a recording device such as a hard disk, SSD (Solid State Drive), or a recording medium such as an IC card, SD card, DVD.

Description of Reference Numerals

[0086] 1…Video monitoring system, 2…Shooting system, 21…Camera unit, 3…Video analysis system, 31…Video input unit, 32…Judgment unit, 321…Person detection unit, 322…Motion amount calculation unit, 323…Candidate judgment unit, 33…Model selection unit, 34…Action recognition unit, 35…Output control unit, 36…Memory unit, 361…Motion amount condition information, 362…Monitoring standard information, 37…Condition calculation unit, 4…Monitoring center system, 41…Recording unit, 42…Video display unit, 43…Management control unit

Claims

1. A video analysis system that detects an event in the monitoring area using a video obtained by photographing the monitoring area, comprising: one or more arithmetic units; one or more storage devices that store preset action amount conditions that define a range of action amounts, wherein the one or more arithmetic units calculate the action amount of a person detected based on the video, determine whether the person is a candidate for monitoring based on the calculated action amount and the action amount condition, select an action recognition model based on the action amount and the action amount condition, detect an event for the person determined to be a candidate for monitoring using the selected action recognition model, wherein the action amount condition is set based on a probability density function for each action calculated in advance, wherein the one or more arithmetic units present information on the probability density function and the action amount condition to the user, adjust the action amount condition according to an input from the user, update the probability density function stored in the one or more storage devices using the action amount and the recognition result of the action type obtained at the operation site, and present the updated probability density function to the user. A video analysis system.

2. The video analysis system according to Claim 1, wherein the one or more arithmetic units determine a candidate for monitoring whose event satisfies a preset condition as a person to be monitored, and store information on the person to be monitored in the one or more storage devices. A video analysis system.

3. The video analysis system according to Claim 1, wherein the one or more storage devices store importance information for the action type and / or the occurrence location, and the one or more arithmetic units calculate a monitoring importance based on the importance information for the person who caused the event. A video analysis system.

4. The video analysis system according to Claim 3, wherein the one or more arithmetic units perform a process of highlighting the person based on the monitoring importance. A video analysis system.

5. The video analysis system according to Claim 1, wherein the one or more arithmetic units display a frame image at the time of detection of the event, a video including the time of detection of the event, or the current position or video of the person of the event. A video analysis system.

6. The video analysis system according to Claim 1, wherein the one or more arithmetic units update the action amount condition based on the updated probability density function. A video analysis system.

7. The video analysis system according to claim 1, wherein the one or more computing devices are a plurality of computing devices, a part of the plurality of computing devices is implemented in a surveillance camera that captures the video, a part of the processing by the one or more computing devices is executed by the computing devices implemented in the surveillance camera, the video analysis system.

8. A video analysis method for detecting an event in a surveillance area using a video captured of the surveillance area, wherein a video analysis system calculates an amount of movement of a person detected based on the video, the video analysis system determines whether the person is a candidate for being monitored based on the calculated amount of movement and a preset amount-of-movement condition that defines a range of the amount of movement, the video analysis system selects an action recognition model based on the amount of movement and the amount-of-movement condition, the video analysis system uses the selected action recognition model to detect an event for the person determined to be a candidate for being monitored, the amount-of-movement condition is set based on a probability density function for the amount of movement for each action calculated in advance, the video analysis method is such that the video analysis system presents information on the probability density function and the amount-of-movement condition to the user, the video analysis system adjusts the amount-of-movement condition in response to an input from the user, the video analysis system updates the probability density function using the amount of movement and the recognition result of the action type obtained at the operation site, the video analysis system presents the updated probability density function to the user, the video analysis method.

Citation Information

Patent Citations

  • Method for determining target person of service in service system by robot and service system by robot using same method

    JP2008142876A

  • Apparatus for detecting fall-down state

    JP2008152717A

  • System, identifier unit, identification model generator, information processing method and program

    JP2016099716A

  • Behavior detection system, information processing device, and program

    JP2018029671A

  • Information processing device, system, information processing device control method, and program

    JP2019133437A