Image processing system, image processing method, and image processing program
The system enables operators to enhance video analysis accuracy by inputting category names and frames, facilitating learning and adaptation of modules, thus improving surveillance system performance.
Patent Information
- Application Number
- JP2024232411
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2013-06-28
- Filing Date
- 2024-12-27
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2034-06-25
AI Technical Summary
Existing video analysis systems lack the ability for active operator intervention during operation, preventing the improvement of analysis accuracy and requiring significant time and effort to adapt to diverse surveillance environments.
A video processing system and method that allows operators to input and associate category names and graphic frames with video sections, enabling learning and updating of image analysis modules during operation, with semi-automatic collection of learning videos and category association.
Improves the accuracy of video analysis in real-time surveillance systems by allowing operators to enhance and adapt video analysis modules based on actual surveillance footage, reducing manual effort and time.
Smart Images

Figure 0007764941000001 
Figure 0007764941000002 
Figure 0007764941000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for analyzing video from a surveillance camera. [Background technology]
[0002] In the above technical field, Patent Document 1 discloses a technology that uses real-time learning to eliminate the need for prior knowledge and prior learning for a behavior recognition system. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] WO2008 / 098188 publication Summary of the Invention [Problem to be solved by the invention]
[0004] However, the technology described in the above document uses machine learning to recognize behavior and characterize certain behaviors as normal or abnormal based on past observations of similar objects. This does not allow for active intervention or support from the system operator, making it impossible to train a classifier during operation. In other words, it is not possible to improve the analysis accuracy of the behavior analysis system during actual operation.
[0005] An object of the present invention is to provide a technique for solving the above-mentioned problems. [Means for solving the problem]
[0006] In order to achieve the above object, the video processing system according to the present invention comprises: A setting means for setting a category of an object detected from an image captured by a camera; a first input means for receiving, by an operator's operation, an input of a category name of an object included in the video that is different from a preset category; a second input means for receiving an input of a graphic frame that specifies an area corresponding to the different category by an operation of an operator on the image; a third input means for receiving an input for associating the different categories with a part of the section of the video by an operator's operation on a display showing the relationship between the categories of objects included in the video and the section of the video; a data storage means for storing names of the different categories and positions of the figures in the video in association with each other so that the names can be learned by an image analysis module; The video processing system includes:
[0007] In order to achieve the above object, a video processing method according to the present invention comprises: Set the category of objects detected from the camera footage, receiving, through an operator's operation, an input of a category name of an object included in the video that is different from a preset category; receiving input of a graphic frame specifying an area corresponding to each of the different categories through an operation by an operator on the image; receiving an input by an operator to associate the different categories with a section of the video by operating a display showing a relationship between categories of objects included in the video and sections of the video; and storing names of the different categories and positions of the figures in the video in association with each other for training an image analysis module. It is a video processing method.
[0008] In order to achieve the above object, a video processing program according to the present invention comprises: A process of setting a category for objects detected from the video captured by the camera; a process of receiving, through an operator's operation, an input of a category name of an object included in the video that is different from a preset category; a process of receiving input of a graphic frame specifying an area corresponding to each of the different categories by an operation of an operator on the image; receiving an input by an operator on a display showing a relationship between categories of objects included in the video and sections in the video, to associate the different categories with certain sections of the video; a process of associating and storing names of the different categories with positions of the figures in the video for learning by an image analysis module; It is a program that causes a computer to execute the above. [Effects of the Invention]
[0009] According to the present invention, it is possible to effectively and efficiently improve the accuracy of video analysis during the actual operation of a surveillance system. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a block diagram showing a configuration of a video processing device according to a first embodiment of the present invention. [Figure 2A] 1 is a block diagram showing a configuration of a video monitoring system according to a prerequisite technique of the present invention; [Figure 2B] 1 is a block diagram showing a configuration of a video monitoring system according to a prerequisite technique of the present invention; [Figure 3] 1 is a flowchart showing a processing flow of a video monitoring system according to a prerequisite technique of the present invention. [Figure 4] 10 is a flowchart showing the flow of a learning process of the video monitoring system according to the prerequisite technology of the present invention. [Figure 5] FIG. 10 is a block diagram showing the configuration of a video monitoring system according to a second embodiment of the present invention. [Figure 6A] FIG. 10 is a diagram showing the configuration of a category table used in the video monitoring system according to the second embodiment of the present invention. [Figure 6B] FIG. 10 is a diagram showing the configuration of a category table used in the video monitoring system according to the second embodiment of the present invention. [Figure 6C] 10 shows the contents of category information sent from a group of video monitoring operation terminals to a learning video extraction unit in a video monitoring system according to a second embodiment of the present invention. [Figure 7A] FIG. 10 is a diagram showing an example of a display image of a video monitoring system according to a second embodiment of the present invention. [Figure 7B] FIG. 10 is a diagram showing an example of an occurrence event table for each camera stored in the video monitoring system according to the second embodiment of the present invention. [Figure 8] FIG. 10 is a diagram showing an example of a display image of a video monitoring system according to a second embodiment of the present invention. [Figure 9A] 10 is a flowchart showing the flow of learning processing in a video monitoring system according to a second embodiment of the present invention. [Figure 9B] 10 is a flowchart showing the flow of learning processing in a video monitoring system according to a second embodiment of the present invention. [Figure 10] FIG. 10 is a block diagram showing the configuration of a video monitoring system according to a third embodiment of the present invention. [Figure 11] FIG. 11 is a block diagram showing the contents of an incentive table of a video monitoring system according to a third embodiment of the present invention. [Figure 12] 10 is a flowchart showing the flow of learning processing in a video monitoring system according to a third embodiment of the present invention. [Figure 13] FIG. 10 is a diagram showing an example of a display image of a video monitoring system according to a fourth embodiment of the present invention. [Figure 14] 10 is a flowchart showing the flow of a learning process of a video monitoring system according to a fourth embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the components described in the following embodiments are merely examples and are not intended to limit the technical scope of the present invention.
[0012] [First embodiment] A video processing device 101 according to a first embodiment of the present invention will be described with reference to Fig. 1. As shown in Fig. 1, the video processing device 101 includes a video data storage unit 121, a video analysis unit 111, a display control unit 123, and a learning data storage unit 140.
[0013] The video data storage unit 121 stores video data captured by the surveillance camera 102. The video analysis unit 111 analyzes the video data stored in the video data storage unit 121 to detect events belonging to specific category information and outputs the detection results. The display control unit 123 also displays a category information setting screen for setting category information for events included in the video, along with the video of the video data stored in the video data storage unit 121. The learning data storage unit 140 stores, as learning data, the category information set in response to an operation by the operator 180 on the category information setting screen and the video data to which the category information has been set. The video analysis unit 111 then performs learning processing using the learning data stored in the learning data storage unit 140.
[0014] According to this embodiment, it is possible to effectively and efficiently improve the accuracy of video analysis during the actual operation of the surveillance system.
[0015] [Second embodiment] The second embodiment of the present invention relates to a technology that collects learning videos for detection by a video detection engine by category, and uses the collected learning videos to create new modules and improve the accuracy of existing modules. Note that in the following explanation, the term "video" is used as a concept that does not only refer to moving images, but also includes still images.
[0016] (Prerequisite technology) First, the underlying technology of a video monitoring system according to a second embodiment of the present invention will be described with reference to Figures 2A to 4. Figures 2A and 2B are diagrams for explaining a video monitoring system 200 as the underlying technology of this embodiment.
[0017] As shown in FIG. 2A, the video monitoring system 200 includes a data center 201 and a group of monitoring cameras 202. The data center 201 includes a video analysis platform 210 and a video monitoring platform 220, and further includes a group of multiple video monitoring operation terminals 232. This figure shows a video monitoring room 230 equipped with multiple video monitoring operation terminals 232. In the video monitoring room 230, an operator 240 checks the video of the monitoring target while viewing the dual-screen monitors provided on each terminal of the group of video monitoring operation terminals 232. Here, an example is shown in which a 16-screen split display is performed on the left screen and a single-screen enlarged display is performed on the right screen. However, the present invention is not limited to this, and any display may be used. For example, the left and right may be reversed, and the number of divisions on each screen may be any number. Multiple large monitors 231 are installed on the front wall of the video monitoring room 230, and problematic video or still images are displayed on them. In this type of video monitoring room 230, for example, 200 operators 240 take turns monitoring footage from 16 cameras each, continuously monitoring all 2,000 cameras 24 hours a day, 360 days a year. While viewing the footage from the 16 surveillance cameras assigned to them, the operators 240 identify and report to a supervisor any incidents of interest, such as reckless driving, dangerous objects like guns or knives, theft, snatching, running away from home, assault, murder, drug trafficking, and problematic behavior like trespassing, as well as similar objects and behaviors (e.g., suspicious individuals or crowd movements). The supervisor then reviews the footage and, if necessary, contacts the police or hospitals to assist in rescuing victims and arresting criminals.
[0018] The video monitoring platform 220 is called a VMS (Video Management System), and stores video data acquired from the surveillance cameras 202 and distributes it to the video monitoring operation terminals 232. As a result, the video monitoring operation terminals 232 display the video data in real time according to predetermined allocation rules. In response to a request from an operator 240 operating the video monitoring operation terminals 232, the video monitoring platform 220 selects one of the surveillance cameras 202 and sends a PTZ (pan-tilt-zoom) operation instruction.
[0019] Video analysis platform 210 performs an analysis process on the video data stored in video monitoring platform 220, and if video data that meets the conditions is found, it transmits category information specifying the target video data to video monitoring platform 220. Video monitoring platform 220 generates an alert screen according to the category information received from video analysis platform 210, and notifies a predetermined terminal of video monitoring operation terminal group 232. In some cases, the problematic video is forcibly enlarged and displayed on large monitor 231.
[0020] Fig. 2B is a block diagram showing the detailed configuration of video surveillance system 200. As shown in Fig. 2B, video surveillance platform 220 collects video data 250 from surveillance cameras 202, adds the capture time, camera position, camera ID, and other information, and stores the data in video storage 221. In addition, camera selection operation unit 222 receives camera designation and PTZ (pan-tilt-zoom) operation instructions from operator 240 via video surveillance operation terminals 232, and operates the designated surveillance camera.
[0021] The video surveillance platform 220 includes a display control unit 223 that displays an alert on the video surveillance operation terminal group 232, and a video reading processing unit 224 that plays / edits past video stored in the video storage 221 in response to instructions from the video surveillance operation terminal group 232.
[0022] The video analysis platform 210 includes predefined video analysis modules 211. Each video analysis module is configured with an algorithm and / or parameters for detecting a different type of problematic video. These predefined video analysis modules 211 use predefined algorithms and parameters to detect video containing predefined events, and transmit predefined category information for the detected video data to the video surveillance platform 220.
[0023] 3 is a flowchart for explaining the flow of processing in the data center 201. In step S301, the video monitoring platform 220 receives the video data 250 from the group of monitoring cameras 202.
[0024] Next, in step S302, the video monitoring platform 220 stores the received video data 250 in the video storage 221, and transmits it to the video monitoring operation terminal group 232 and the video analysis platform 210.
[0025] Next, in step S303, the camera selection operation unit 222 receives camera selection information and camera operation information from the operator 240 via the video monitoring operation terminal group 232, and transmits an operation command to the selected monitoring camera.
[0026] On the other hand, in step S304, the video analysis platform 210 performs analysis processing of the video data received from the video monitoring platform 220 using the default video analysis module 211.
[0027] Then, in step S305, if the default video analysis module 211 detects video that satisfies the predetermined conditions, the process proceeds to step S307, where it transmits category information to the video surveillance platform 220. Even if it does not detect video that satisfies the conditions, the process proceeds to step S310, where it transmits category information indicating "no category information" to the video surveillance platform 220.
[0028] Furthermore, in step S308, the display control unit 223 of the video monitoring platform 220 generates an alert screen and transmits it to the video monitoring operation terminal group 232 together with the video from the target monitoring camera.
[0029] In step S309, an operation by the operator 240 on the alert screen (an operation to report to a supervisor or the police) is accepted.
[0030] FIG. 4 is a flowchart illustrating the process flow for generating the default video analysis module 211. This video analysis module generation process is performed before the data center 201 is built on-site. First, in step S401, a large amount of video containing the event to be detected is extracted by human eyes from a huge amount of past video. Alternatively, in step S402, an event similar to the event to be detected is intentionally generated in an environment similar to the actual operating environment, filmed, and sample video is extracted. In step S403, additional information is manually added to each of the large amount of extracted / collected video data to create training video. Furthermore, in step S404, a researcher / engineer selects the optimal algorithm for the target object, event, or action, and trains the training video data to generate the default video analysis module 211.
[0031] (Issues with prerequisite technology) When creating a default video analysis module using the above prerequisite technologies, it required a huge amount of time and effort to collect and assign correct answers. For example, 2,000 images were required for facial recognition, and 1 million images were required for specific event detection using deep learning, making implementation difficult. In other words, creating a video analysis module (classifier) from training video was done manually, and during that process, the operation of the video analysis module was also verified individually, and the environment had to be set up separately, which required a lot of time and effort.
[0032] In recent years, however, as the types of crimes and accidents have become more diverse, customer demand for additional processing capabilities for detectable events has increased. Furthermore, with default video analysis modules that only apply learning footage collected in environments other than the actual surveillance environment, the accuracy of detecting problematic footage can sometimes drop significantly depending on the actual surveillance environment. Furthermore, adapting the module to the actual surveillance environment requires a significant amount of work and time.
[0033] (Configuration of this embodiment) FIG. 5 is a block diagram showing the configuration of a video surveillance system 500 as an example of a surveillance information processing system according to this embodiment. Components similar to those in the underlying technology are assigned the same reference numerals, and descriptions thereof will be omitted. Unlike the underlying technology shown in FIG. 2B, a data center 501 as an example of a video processing device according to this embodiment includes a training database 540 and a video surveillance operation terminal 570. The training database 540 stores training video data 560 to which category information 561 selected by an operator 240 is added. The video surveillance operation terminal 570 is a terminal through which a supervisor 580 specifies category information. The video analysis platform 510 includes a newly created new video analysis module 511, a category information adding unit 515, and a new video analysis module generating unit 516. Meanwhile, the video surveillance platform 520 includes a new training video extracting unit 525.
[0034] Category information is information that indicates the classification of objects and actions that are to be detected in video. Examples of category information include "gun," "knife," "fight," "reckless driving," "double riding on a motorcycle," and "drug dealing." The video analysis platform 510 includes category information tables 517 and 518, which store various types of category information and their attributes in association with each other.
[0035] 6A and 6B are diagrams showing the contents of the default category information table 517 and the new category information table 518. The category information tables 517 and 518 store category information types, shapes, trajectory information, size thresholds, etc. in association with category information names. By referencing these, the video analysis modules 211 and 511 can determine the categories of events included in the video data.
[0036] Figure 6C shows the content of category information 531 sent from the video monitoring operation terminal group 232 to the learning video extraction unit 525. As shown in Figure 6C, category information 531 includes the following: "category," "camera ID," "video capture time," "event area," "operator information," and "category type." Here, "category" refers to the name of the object or behavior to be detected. "camera ID" is an identifier for identifying the surveillance camera. "video capture time" is information indicating the date and time when the video data to which the category information should be added was captured. It may also refer to a specific period (start to end of capture). "event area" is information indicating the "target shape" in the video, "video background subtraction," and "position of the background subtraction within the overall video." Event areas can be rectangular or various types, such as masked video, background subtraction video, and polygonal video. "Operator information" indicates information about the operator who added the category information, including the operator ID and name. "Category type" refers to the type of object / event to be detected. For example, the category information "gun" is a category type that accumulates the "shape" of the target gun as learning video data. The category information type for "runaway" is a category information type that accumulates learning video data with the trajectory from the start point to the end point of the background difference of a specified area as the "action." The category information "drug trafficking" is a category information type that accumulates learning video data with the trajectory of the background difference in the entire video as the "action."
[0037] 5, if it is possible to increase the number of categories of new events to be detected by adjusting or adding parameters in the already existing video analysis modules 211, 511, category adding unit 515 adjusts or adds parameters to those video analysis modules 211, 511. On the other hand, if it is determined that it is not possible to increase the number of categories of new events to be detected by adjusting or adding parameters to the already existing video analysis modules 211, 511, new video analysis module generating unit 516 generates new video analysis module 511. Furthermore, learning video data 560 accumulated in learning database 540 is allocated to the video analysis modules 211, 511 selected based on category information 561, and each video analysis module is caused to perform learning processing.
[0038] The display control unit 523 generates a category information selection screen 701 as shown in FIG. 7A and sends it to the video monitoring operation terminal group 232. The category information selection screen 701 includes an "Other" 711 as a category information option in addition to pre-prepared category information (e.g., "No Helmet," "Two People Riding a Motorcycle," "Speeding," "Gun," "Knife," "Drug Deal," "Fight," etc.). To prevent the operator 240 from being discouraged from selecting category information, it is preferable to display only a portion of category information candidates (e.g., five candidates) predicted in advance by referring to the camera event table 702 shown in FIG. 7B on the category information selection screen 701. The camera event table 702 stores, for each camera ID 721, category information on events that are likely to be included in video captured by that camera. Specifically, the category 722 of the event that occurred and its occurrence rate 723 are stored in descending order of occurrence rate. The category information related to the selected category information is stored in the training database 540 together with the video data from which the category information was selected.
[0039] For video data for which "Other" 711 has been selected, the learning video extraction unit 525 separately stores the video data in the learning database 540 so that the supervisor 580 can input specific category information as needed via the video monitoring operation terminal 570. When "Other" is selected, a category information input request is sent to the video monitoring operation terminal 570 for the supervisor 580 along with video identification information indicating the video data at that time. When the operator 240 performs a category information selection operation during video monitoring via the video monitoring operation terminal group 232, a set 531 of video identification information and category is stored in the learning database 540 via the video monitoring platform 520. Meanwhile, the display control unit 523 sends a new category information generation screen to the video monitoring operation terminal 570 to prompt the supervisor 580 to perform new category information generation processing.
[0040] Fig. 8 is a diagram showing a specific example of such a new category information setting screen 801. In this new category information setting screen 801, for example, in addition to a category information name input field 811 and a category information type selection field 813, an area designation graphic object 812 is provided for designating an area to focus on in order to detect video included in that category information. In addition to the new category information setting screen 801 shown in Fig. 8, information specifying the operator who selected the "Other" category information, the location and time at which the video was acquired, etc. may also be displayed. It may be determined by video analysis which category information video categorized as "Other" is similar to, and this may be presented to supervisor 580.
[0041] The new video analysis module generation unit 516 selects an existing algorithm that fits the new category information and creates a new video analysis module using a neural network or the like that conforms to that algorithm. It then trains the module using accumulated video data for training. For the training and application process, either batch processing or on-the-fly processing can be selected according to the category information.
[0042] The category information addition unit 515 performs batch processing or real-time processing, registers information about the added category information in the category information table 518, and the existing video analysis module 211 or the new video analysis module 511 specifies the category information to be referenced in the category information tables 517 and 518.
[0043] Based on the specified category information, the default video analysis module 211 and the new video analysis module 511 perform learning processing using their respective learning video data. This improves the video analysis accuracy of the default video analysis module 211, and the new video analysis module 511 is completed as a new video analysis module.
[0044] If there is no existing algorithm that fits the new category, the new video analysis module generation unit 516 may automatically generate a new algorithm (for example, one that still recognizes a person behind a person even if multiple people pass in front of them).
[0045] (Processing flow) 9A and 9B are flowcharts for explaining the flow of processing by video monitoring system 500. Steps S301 to S309 are the same as the processing of the base technology explained in Fig. 3, so explanations will be omitted here, and only steps S900 to S911 after step S309 will be explained.
[0046] In step S309, the learning video extraction unit 525 examines the video and if an alert is issued indicating that the video contains an event that should be detected, the process proceeds to step S900, and the display control unit 523 displays a category information selection screen 701 as shown in Figure 7A on the video monitoring operation terminal group 232.
[0047] In step S901, operator 240 selects category information. If the selected category information is specific category information, the process proceeds to step S902, where learning video extraction unit 525 generates category information and assigns it to the video data, thereby generating learning video data.
[0048] Next, in step S903, learning video extraction unit 525 stores the generated learning video data in learning database 540, and further in step S904, learning processing is performed by default video analysis module 211.
[0049] On the other hand, if the operator 240 selects the "Other" category information in step S901, the process proceeds to step S905, where the learning video extraction unit 525 generates category information with the category information name set to "Other" and the category information type set to "NULL", and assigns this to the learning video data.
[0050] Next, in step S906, the learning video extraction unit 525 stores the learning video data with the category information "Other" added in the learning database 540, and at the same time, the display control unit 523 sends the learning video data and the new category information setting screen 801 to the video monitoring operation terminal 570 of the supervisor 580.
[0051] In step S907, the category adding unit 515 receives an instruction from the supervisor 580 to set new category information and associates it with the accumulated learning video data.
[0052] In step S908, category information adding unit 515 determines whether or not there is a default video analysis module 211 that fits the set new category information. If there is a default video analysis module 211 that fits the set new category information, the process proceeds to step S909, where it is set as new category information for the target default video analysis module 211, and then in step S904, learning is performed using the learning video data with the new category information added.
[0053] On the other hand, if it is determined in step S908 that there is no pre-defined video analysis module 211 that fits the set new category information, the process proceeds to step S911, where a new video analysis module 511 is generated and trained on the training video.
[0054] FIG. 9B is a flowchart showing a detailed flow of the new video analysis module generation process in step S911. In step S921, an algorithm database (not shown) is referenced. In step S923, an algorithm (e.g., an algorithm for extracting feature vectors, an algorithm for extracting clusters consisting of collections of blobs (small image regions) and their boundaries) is selected according to the category information type specified by the supervisor, and the framework of the video analysis program module is generated. Next, in step S925, the video region to be analyzed by the video analysis module is set using the judgment region specified by the supervisor. Furthermore, in step S927, thresholds for the shape and size of the object to be detected, the direction and distance of the movement to be detected, and the like are determined using multiple learning video data. At this time, feature vectors of the object to be detected and its movement in the learning video data may be extracted, and the feature vectors may be set as the threshold. When a cluster consisting of a collection of blobs and its cluster boundary are extracted as feature vectors, the feature vectors may be set as the threshold.
[0055] As described above, according to this embodiment, an operator can easily accumulate learning videos and simultaneously associate them with categories during actual operation. This allows for semi-automatic collection of learning videos and association with categories, thereby reducing the man-hours and time required for localization to the environment and creation of new video analysis modules.
[0056] Furthermore, because the system can learn from video footage of the operating environment, it is possible to build a more accurate video analysis module. This technology can be applied to security fields such as video surveillance and security guards. It can also be used for customer-oriented analysis using video in stores and public areas.
[0057] [Third embodiment] Next, a video monitoring system according to a third embodiment of the present invention will be described. The video monitoring system according to this embodiment differs from the second embodiment in that it takes into consideration incentives for operators. As the other configurations and operations are the same as those of the second embodiment, the same components are assigned the same reference numerals and detailed explanations will be omitted here.
[0058] 10 is a block diagram showing the configuration of a video monitoring system 1000 according to this embodiment. Unlike the second embodiment, the video monitoring platform 1020 included in the video monitoring system 1000 includes an incentive table 1026. The incentive table 1026 collects statistics on the number of training videos stored in the training database 540 for each operator and provides incentives, thereby improving collection efficiency.
[0059] An example of the incentive table 1026 is shown in FIG. 11. The incentive table 1026 stores and manages the number of learning videos 1102, the number of new categories 1103, and points 1104, linked to an operator ID 1101. The number of learning videos 1102 indicates the number of learning video data for which the operator selected and assigned a category. The number of new categories 1103 indicates the number of categories for which the operator selected "other" as a category and which were ultimately created as new categories by a supervisor. The number of learning videos and the number of new categories can be evaluated as contributions to the operator's monitoring work, so points 1104 are calculated based on these values, and are also stored and updated linked to the operator ID. The importance of the videos discovered by the operator and assigned a category may be taken into consideration and weighted when calculating the points 1104.
[0060] By determining the hourly wage or salary of an operator according to the value of this point 1104, it is possible to motivate the operator in the monitoring work. In addition to points, a value representing the accuracy of category assignment may be used in the incentive table 1026 as a value for evaluating an operator. For example, test video data to which a correct category has been assigned may be shown to multiple operators, and each operator's detection speed, detection accuracy, and probability of correct category assignment may be verified, and the operator evaluation value may be calculated using these.
[0061] 12 is a flowchart illustrating the processing flow of the video monitoring system 1000. When a default video analysis module is trained or a new video analysis module is generated by category selection or category generation, the process proceeds to step S1211, where points are awarded. That is, for example, an incentive (points) according to the category of the event discovered by the operator is saved in association with the ID of the operator who performed the categorization.
[0062] According to this embodiment, it is possible to stimulate the monitoring motivation of the operator.
[0063] [Fourth embodiment] As an example of this method, we will explain a system that allows operators to add and modify category information, as shown in Figure 2B, which is a prerequisite technology for the monitoring information system.
[0064] The display control unit 523 generates a category information selection screen 1301 as shown in FIG. 13 and sends it to the video monitoring operation terminal group 232. The category information selection screen 1301 includes a video display screen 1303, a category information bar 1302 indicating the occurrence status of an alert, video control components 1304 (e.g., "play," "stop," "pause," "rewind," "fast forward," etc.), a progress bar 1307 indicating the video playback status, and a category information setting button 1305. In addition to the category information prepared in advance, the category information setting button 1305 includes an "other" 1306 as a category information option. The operator 240 checks the video using the video control components 1304. The operator 240 modifies or adds category information for the video data displayed on the display screen 1303 using the category information setting button 705. The category information set by the operator 240 is stored in the learning database 540 together with the video data.
[0065] Figure 14 is a flowchart illustrating the flow of processing by video monitoring system 500. Steps S301 to S309 are the same as the processing of the underlying technology described in Figure 3, so their explanation will be omitted here, and steps S900 to S910 after step S309 will be explained. Also, steps S900 and steps S902 to S910 are the same as the processing described in Figure 9A, so their explanation will be omitted here, and step S1401 will be explained. In step S309, the learning video extraction unit 525 examines the video and if an alert is issued indicating that the video contains an event that should be detected, the process proceeds to step S900, and the display control unit 523 displays a category information selection screen 1301 as shown in Figure 13 on the video monitoring operation terminal group 232.
[0066] On the category information selection screen 1301, the video and its category information are displayed in a display section 1303 and a category information bar 1302. In addition, by using a video control component 1304, the category information and the video can be checked by playing, rewinding, or fast-forwarding. The category information bar 1302 displays, in color, the category information generated for the video displayed on the category information selection screen 1301 and the correct category information corrected or added by the operator 240. For example, a section in which a "no helmet" alert occurred is displayed in blue, a section in which no alert occurred or a section 1310 in which there is no category information is displayed in black, and a section in which a "two people riding a motorcycle" alert occurred is displayed in red. In addition, the correct category information corrected or added by the operator 240 using a category information setting button 1305 is also displayed in a similar manner.
[0067] In step S1401, the operator 240 checks the video in which an alert has occurred and its category information using the category information selection screen 1301, and modifies or adds category information for sections where modification or addition of category information is required.
[0068] When correcting or adding category information, a category information setting button 1305 corresponding to a specific category is pressed while a portion of the video is being played. For example, consider correcting a section in which a "no helmet" alert has occurred to a "two-person motorcycle" alert. Here, the section belonging to the "no helmet" category is displayed in blue in the category information bar 1302. Using the video control component 1304 on the category information selection screen 1301, the video is played on the display unit 1303. When the corresponding video data is displayed on the display unit 1303, the "two-person motorcycle" button of the category information setting buttons 1305 is pressed to correct the category to "two-person motorcycle." At this time, the color of the category information bar 1302 corresponding to the section for which category information has been set by the operator 240 changes from "blue" to "red."
[0069] It is also possible to add multiple types of category information to the same section using video control component 1304. In this case, the category information bar 1302 for a section to which multiple categories of category information have been added is displayed in multiple layered colors. For example, if "no helmet" and "two people riding a motorcycle" are added, blue and red are displayed in layers, as in category information bar 1309.
[0070] If you wish to delete category information, simply use the video control component 1304 to play back the video for the section for which you want to delete category information, and press the "No category information" button 1308. This is the same as correcting the category information to "No category information." The category information bar 1310 for the section corrected to "No category information" is displayed in black, just like a section in which no alert has occurred.
[0071] Here, when correcting or adding category information for a certain continuous section of video, while the video of the relevant section is being played back, it is sufficient to keep pressing a specific button of the category information setting button 1305. Furthermore, if the section for which you want to correct or add category information is long, you may make the category information setting button 1305 a toggle type so that you press the button only at the beginning and end of the section for which you want to correct or add category information.
[0072] In step S1401, if the operator has corrected or added the correct category information, the video analysis module 211 can learn the alert that should be output.
[0073] If the corrected or added category information is specific category information, the process proceeds to step S902, where learning video extraction unit 525 adds the category information to the video data, and generates learning video data.
[0074] By having an operator check the category information from the video and the surveillance information system and set the correct category information, a surveillance information system that can detect new objects and actions can be constructed. Also, by having the operator correct the category information, the accuracy of the surveillance information system can be improved.
[0075] [Other embodiments] Although the present invention has been described above with reference to the embodiments, the present invention is not limited to the above embodiments. Various modifications that are understandable to those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention. Furthermore, systems or devices that combine the separate features included in each embodiment in any manner are also included in the scope of the present invention.
[0076] The present invention may also be applied to a system consisting of multiple devices or to a single device. Furthermore, the present invention is also applicable when an information processing program that realizes the functions of the embodiments is supplied directly or remotely to a system or device. Therefore, the scope of the present invention also includes a program installed on a computer to realize the functions of the present invention, a medium storing the program, and a WWW (World Wide Web) server from which the program can be downloaded. In particular, the scope of the present invention includes at least a non-transitory computer-readable medium storing a program that causes a computer to execute the processing steps included in the above-described embodiments.
[0077] This application claims priority based on Japanese Patent Application No. 2013-136953, filed June 28, 2013, the disclosure of which is incorporated herein in its entirety by reference.
Claims
1. A setting means for setting a category of an object detected from an image captured by a camera; a first input means for receiving, by an operator's operation, an input of a name of a category of an object included in the video that is different from a preset category; a second input means for receiving an input of a graphic frame that specifies an area corresponding to the different category by an operation of an operator on the image; a third input means for receiving an input by an operator to associate the different categories with a section of the video by operating a display showing a relationship between a category of an object included in the video and a section of the video; a data storage means for storing names of the different categories and positions of the figures in the video in association with each other so that the names can be learned by an image analysis module; A video processing system comprising:
2. 2. The image processing system according to claim 1, wherein the names of the different categories and the positions of the graphics in the image are stored in association with identification information of the operator.
3. 3. The image processing system according to claim 1, wherein an object existing in the area corresponding to the different category is different from an object whose category is set by the setting means.
4. Further, a detection means is provided for performing detection using a result of learning the names of the different categories and the positions of the figures in the video. The video processing system according to any one of claims 1 to 3.
5. The image processing system according to claim 1 , wherein the graphic is a rectangle.
6. The image processing system according to claim 1 , wherein the graphic is a polygon.
7. Set the category of objects detected from the camera footage, receiving, through an operator's operation, an input of a category name of an object included in the video that is different from a preset category; receiving input of a graphic frame specifying an area corresponding to each of the different categories through an operation by an operator on the image; receiving an input by an operator to associate the different categories with a section of the video by operating a display showing a relationship between categories of objects included in the video and sections of the video; and storing names of the different categories and positions of the figures in the video in association with each other for training an image analysis module. Image processing method.
8. A process of setting a category for objects detected from the video captured by the camera; a process of receiving, through an operator's operation, an input of a category name of an object included in the video that is different from a preset category; a process of receiving input of a graphic frame specifying an area corresponding to each of the different categories by an operation of an operator on the image; receiving an input by an operator on a display showing a relationship between categories of objects included in the video and sections in the video, to associate the different categories with certain sections of the video; a process of associating and storing names of the different categories with positions of the figures in the video for learning by an image analysis module; A program that causes a computer to execute the following.
Citation Information
Patent Citations
Surveillance system set-up.
EP2511887A1
Method for sorting defect, device therefor and method for generating data for instruction
JP2000057349A
Elevator control device, monitoring system, searching revitalization program, in / out management system
JP2011068469A
Advertising effect measurement system and advertising effect measurement device
JP2011070629A
Image recognition device and image recognition method
JP2011192178A