Video monitoring system, video recording device, video analysis method, and video analysis program

The video surveillance system enhances reliability of object identification by reconstructing unreliable meta-information using nearby camera data and machine learning, addressing issues from obstructions and reflections.

JP2025128048APending Publication Date: 2025-09-02MITSUBISHI ELECTRIC CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025025287
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-21
Filing Date
2025-02-19
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

The reliability of meta-information obtained from video surveillance systems is reduced due to environmental obstructions and reflections, leading to unreliable object identification.

Method used

A video surveillance system with a video recording device that includes a meta-information reanalysis unit to reconstruct unreliable object identification information using reliable data from nearby cameras or machine learning models, ensuring high reliability.

Benefits of technology

The system effectively reconstructs unreliable object identification information into highly reliable information, smoothing out temporary fluctuations and improving overall analysis accuracy without overburdening the cameras.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025128048000001_ABST
    Figure 2025128048000001_ABST
Patent Text Reader

Abstract

To provide a video monitoring system, a video recording device, a video analysis method, and a video analysis program with which it is possible to reconstruct object identification information with reduced reliability into object identification information with high reliability.SOLUTION: A video monitoring system (100) includes one or more cameras (sensors) (12) that analyze acquired video information and output meta information, and a video recording device (10) that records the video information and the meta information output from the cameras (sensors) (12). The cameras (sensors) (12) each have a video analysis unit (14) that outputs, as the meta information, object identification information indicating an object included in the video information and reliability information indicating the reliability of the object identification information. The video recording device (10) comprises: a video recording unit (15) that records the video information; a meta information recording unit (16) that records the meta information; and a meta information reconstruction unit (17) that reconstructs the object identification information with reliability equal to or less than a predetermined threshold on the basis of the object identification information with reliability higher than the predetermined threshold.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a video monitoring system, a video recording device, a video analysis method, and a video analysis program. [Background technology]

[0002] In recent years, cameras and various sensors have been installed throughout society for purposes such as preventing, detecting, and post-mortem investigations of crimes and accidents, behavioral research for marketing, and autonomous driving. However, visually verifying the vast amount of information acquired by these cameras and sensors is difficult, and technologies for automatically detecting important scenes are being utilized for surveillance. Specifically, video analysis processing is performed on the output of cameras or sensors, such as detecting the movement of the monitored subject or detecting and recognizing people or objects by facial detection. As a result of such video analysis processing, identification information of the detected and recognized person or object, information such as the position and size of the person or object in the video, and reliability information indicating the accuracy of the identification information are output. Such identification information, information such as the position and size of the person or object, and reliability information are referred to as meta-information. Today's video surveillance systems enable visual surveillance by analyzing this meta-information.

[0003] However, there is a problem in that the reliability of meta-information obtained by analyzing images captured by cameras or sensors is reduced due to the influence of the capturing environment, such as when the monitored object is blocked by an obstruction.

[0004] Patent Document 1 discloses a technology in which an object to be monitored is photographed using multiple cameras or sensors, and based on the object detection information recognized by one camera or sensor, the detection setting information of other cameras or sensors is sequentially reviewed, making it easier to identify objects using those other cameras or sensors. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] Patent No. 4866754 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the technique described in Patent Document 1 has a problem in that the reliability of meta information once determined cannot be improved by post-processing.

[0007] The present disclosure aims to provide a video surveillance system, a video recording device, a video analysis method, and a video analysis program that can reconstruct object identification information with reduced reliability into object identification information with high reliability. [Means for solving the problem]

[0008] The video surveillance system disclosed herein is a video surveillance system including one or more sensors that acquire video information and analyze the acquired video information to output meta information, and a video recording device that records the video information and the meta information output by the sensors, wherein the sensors have a video analysis unit that analyzes the acquired video information and outputs object identification information indicating an object contained in the video information and reliability information indicating the reliability of the object identification information as the meta information, and the video recording device is characterized by having a video recording unit that records the video information, a meta information recording unit that records the meta information, and a meta information re-analysis unit that reconstructs object identification information whose reliability indicated by the reliability information contained in the meta information is below a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold.

[0009] The video analysis method disclosed herein is a computer-executed video analysis method, and includes the steps of recording video information, recording meta information including object identification information indicating an object included in the video information and reliability information indicating the reliability of the object identification information, and reconstructing object identification information whose reliability indicated by the reliability information included in the meta information is below a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold.

[0010] The video analysis program disclosed herein causes a computer to perform the following steps: recording video information; recording meta information including object identification information indicating an object included in the video information and reliability information indicating the reliability of the object identification information; and reconstructing object identification information whose reliability indicated by the reliability information included in the meta information is below a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold. [Effects of the Invention]

[0011] According to the present disclosure, it is possible to provide a video surveillance system, a video recording device, a video analysis method, and a video analysis program that can reconstruct object identification information with reduced reliability into object identification information with high reliability. [Brief explanation of the drawings]

[0012] [Figure 1] 1 is a system configuration diagram showing a schematic configuration of a video monitoring system according to a first embodiment. [Figure 2] 1 is a hardware configuration diagram showing a video recording device according to a first embodiment. [Figure 3] FIG. 1 is an image diagram of installation of cameras (sensors) in a video monitoring system according to the first embodiment. [Figure 4] is a graph showing the reliability of object identification information relating to the monitored object in a time series in relation to the detection of the monitored object by the camera (sensors) in the case shown in FIG. [Figure 5] FIG. 10 is an explanatory diagram showing an example of a case where the detection result is re-evaluated according to the first embodiment. [Figure 6] 10 is a flowchart showing an example of a process for reevaluating meta information in the video recording device according to the first embodiment. [Figure 7] 10 is a system configuration diagram showing a schematic configuration of a video monitoring system according to a second embodiment. [Figure 8] 10 is a flowchart showing an example of a process for reevaluating meta information in the video recording device according to the second embodiment. [Figure 9] 10 is a system configuration diagram showing a schematic configuration of a video monitoring system according to a third embodiment. [Figure 10] 11 is an image diagram of the inference processing of the video monitoring system according to the third embodiment. [Figure 11] This is an overview of the processing in the learning phase for the inference model recording unit of the video surveillance system in embodiment 3. [Figure 12] 10 is a flowchart showing an example of a process for generating an inference model in a video recording device according to embodiment 3. [Figure 13] 10 is a flowchart showing an example of a process for reevaluating meta information in a video recording device according to a third embodiment. [Figure 14] 10 is a system configuration diagram showing a schematic configuration of a video monitoring system according to a fourth embodiment. [Figure 15] FIG. 10 is an image diagram of the installation of cameras (sensors) in a video monitoring system according to the fourth embodiment. [Figure 16] is a graph showing the reliability of object identification information relating to a mirror image over time in relation to the detection of a mirror image reflected in a mirror or the like by a camera (sensor) in the case shown in FIG. [Figure 17] FIG. 10 is an explanatory diagram showing an example of a case where the detection result is re-evaluated according to the fourth embodiment. [Figure 18] 10 is a flowchart showing an example of a process for correcting meta information in a video recording device according to a fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Hereinafter, a video monitoring system according to an embodiment will be described with reference to the drawings. The following embodiments are merely examples, and the embodiments can be appropriately combined and modified.

[0014] First Embodiment Fig. 1 is a system configuration diagram showing a schematic configuration of a video monitoring system 100 according to the first embodiment. In Fig. 1, the video monitoring system 100 according to the first embodiment is composed of cameras (sensors) 12 (12A, 12B, 12C, ...) for capturing images of a monitoring target, and a video recording device 10 for storing images and meta information output by the cameras (sensors) 12. The cameras (sensors) 12 may be connected to the video recording device 10 via a network such as the Internet.

[0015] The cameras (sensors) 12 are imaging devices that acquire video information, which is moving images. One or more cameras (sensors) 12A, 12B, 12C, etc. may be provided. Each of the cameras (sensors) 12 may be equipped with a motion sensor that detects a person using infrared light, and acquire video information when the motion sensor detects a person. Each of the cameras (sensors) 12 also includes a video analysis unit 14 that analyzes the captured video information using known image processing techniques such as edge detection. The analysis results, including object identification information indicating an object included in the video information and reliability information indicating the reliability of the identification information, are transmitted to the video recording device 10 as meta-information. In addition to the object identification information and reliability information, the video analysis unit 14 may also transmit three-dimensional information indicating the position and size of the object extracted from the video information to the video recording device 10 as meta-information. When the video information includes multiple objects, the video analysis unit 14 outputs object identification information for each object and reliability information indicating the reliability of the object identification information. Furthermore, when the video information includes multiple objects, the video analysis unit 14 outputs three-dimensional information indicating the position and size of each object.

[0016] The video recording device 10 records the video information received from the camera (sensors) 12 in the video recording unit 15, and the meta information in the meta information recording unit 16. Furthermore, the video recording device 10 includes a meta information reanalysis unit 17 that reanalyzes the recorded video information and meta information. A display device 18 configured with a liquid crystal display, organic electroluminescence, or the like is connected to the video recording device 10. The display device 18 displays the video recorded in the video recording unit 15, the meta information recorded in the meta information recording unit 16, or the analysis results of the meta information reanalysis unit 17. The display device 18 may also include a printer or the like that outputs information.

[0017] 2 is a hardware configuration diagram showing a video recording device 10 according to embodiment 1. The video recording device 10 includes a CPU (Central Processing Unit) 21, a main memory 22, an input / output interface (I / O interface) 23, and a storage unit 24, all of which are connected to one another via a system bus 25. The video recording device 10 may be made up of multiple computers.

[0018] The CPU 21 is an integrated circuit (IC) that performs arithmetic processing. In addition to the CPU 21, an arithmetic element such as a digital signal processor (DSP) or a graphics processing unit (GPU) may be used. The CPU 21 operates as a meta information reanalysis function that replaces a portion of meta information that has become unreliable with replaceable meta information in the vicinity of the portion by executing a video analysis program. As a result, the CPU 21 functions as a meta information reanalysis unit 17 by executing the video analysis program. The video analysis program is provided, for example, on a recording medium on which it is recorded.

[0019] The main memory 22 is configured by a volatile storage device such as a RAM (Random Access Memory) or a non-volatile storage device such as a ROM (Read Only Memory). The storage unit 24 is configured by a non-volatile storage device such as an HDD (Hard Disk Drive) or a flash memory, and functions as the video recording unit 15 and the meta information recording unit 16.

[0020] The I / O interface 23 is a port to which the camera (sensors) 12, which is an input device, and the display device 18, which is an output device, are connected. Specific examples of the I / O interface 23 include a USB (Universal Serial Bus) terminal, an IEEE 1394 terminal, a Thunderbolt terminal, or the like, and further includes a communication interface such as Ethernet (registered trademark).

[0021] 3 is an image diagram of the installation of cameras (sensors) 12 of video monitoring system 100 in the first embodiment. In the case shown in FIG. 3, consider a case where a shielding object 27 passes in front of camera (sensors) 12 while camera (sensors) 12 captures one monitoring target 26 within its captureable range (angle of view). In such a case, camera (sensors) 12 captures monitoring target 26 within its angle of view until a certain point, at which point the shielding object 27 blocks part of monitoring target 26, and eventually the entire monitoring target 26 is blocked. Eventually, as the shielding object 27 moves, part of monitoring target 26 is captured within the angle of view of camera (sensors) 12, and finally the entire monitoring target 26 is captured within the angle of view of camera (sensors) 12 again.

[0022] Figure 4 is a time-series graph of the reliability 30 of the object identification information for the monitored object 26, as detected by the camera (sensors) 12 in the case shown in Figure 3. The reliability 30 of the detection and recognition results for the monitored object 26 by the camera (sensors) 12 is initially high in section G1 because the entire monitored object 26 is captured within the field of view of the camera (sensors) 12. However, as the obstruction 27 moves, the reliability 30 decreases in section G2, where the monitored object 26 is partially obstructed, and in section G3, where the monitored object 26 is completely obstructed, the reliability 30 reaches its minimum value, indicating that detection is impossible. Eventually, in section G4, as the obstruction 27 moves and part of the monitored object 26 becomes visible, the monitored object 26 becomes detectable again, and the reliability 30 of the detection and recognition results also recovers. In section G5, where the monitored object 26 can be completely captured, the reliability 30 recovers to the same initial value as in section G1.

[0023] As shown in FIG. 4, in sections G2 to G4, the monitored object 26 is blocked from the field of view of the camera (sensors) 12 by an obstruction 27. This reduces the reliability 30 of the detection result, or in the worst case, may make it impossible to detect the monitored object 26, even though the monitored object 26 is actually within the field of view of the camera (sensors) 12. Furthermore, detection and recognition results with reduced reliability 30 become noise, which can worsen the overall recognition results for the monitored object 26. Generally, meta-information obtained by analyzing images captured by a camera must be determined within a predetermined time. This predetermined time is the time it takes for the camera to encode and compress the video, e.g., within one second. In such cases, if the reliability 30 is low, detection accuracy will be significantly reduced.

[0024] FIG. 5 is an explanatory diagram showing an example of reevaluation of detection results according to the first embodiment. Meta information for sections G2 to G4, where reliability has decreased, is generated from the detection results for sections G1 or G5, where reliability 30 is high. Specifically, by applying the object identification information for sections G1 or G5 and the reliability information indicating the reliability of the object identification information to sections G2 to G4, the meta information for sections G2 to G4 is reconstructed based on the meta information for sections G1 or G5, as in the reconstructed reliability 32. This eliminates noise that temporarily reduces reliability, making it possible to provide highly reliable detection information to the user. In the first embodiment, the description will be given assuming that sections G2 to G4 are reconstructed to have the same reliability value as sections G1 or G5. However, the reliability value for sections G2 to G4 may be set to a value slightly lower than that for sections G1 or G5, or only the object identification information for sections G1 or G5 may be applied to sections G2 to G4, with the reliability information remaining unchanged. By setting the reliability values ​​of sections G2 to G4 to values ​​lower than those of section G1 or section G5, it becomes easier to grasp the sections whose reliability has been reconstructed. Furthermore, the reliability of each of sections G2 to G4 does not need to be the same value, and the value may be gradually subtracted according to the time difference from section G1 or section G5.

[0025] Next, the process of re-evaluating meta information, which is a detection result, in the first embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the process of re-evaluating meta information in the video recording device 10 according to the first embodiment.

[0026] In step S101, a search is performed for areas where the reliability of the meta information, which is the detection result, is low, such as sections G2 to G4 in Fig. 4. Specifically, areas where the reliability value is equal to or less than a predetermined threshold are detected. The predetermined threshold is, for example, the minimum reliability value of the video information when a machine-learned model such as a CNN (convolutional neural network) can recognize the monitoring target from the video information.

[0027] In step S102, it is determined whether or not a location where the reliability of the meta information, which is the detection result, is low can be replaced with information generated from detection results in the vicinity of the location. In the case shown in Figures 3 and 4, there are detection results of the monitored object 26 with reliability higher than a predetermined threshold before and after the time when the reliability becomes low, so it is determined that replacement with information generated by linearly connecting the detection results of the two points before and after the time when the reliability becomes low is possible. If it is determined in step S102 that replacement with the reliability of the detection result in the vicinity of the location is possible, the procedure proceeds to step S103. If it is determined that replacement with the reliability of the detection result in the vicinity of the location is not possible, the process ends.

[0028] In step S103, the part of the meta-information where the reliability is low is replaced with information generated from the detection result of the vicinity of the part, and the process ends.

[0029] In the first embodiment, the meta information reanalysis unit 17 chronologically examines meta information and reconstructs object identification information whose reliability is below a predetermined threshold based on object identification information that exists before and after the object identification information in the time series and whose reliability is higher than the predetermined threshold. In the embodiment, temporally neighboring information when the reliability of the meta information becomes low is used as highly reliable information near the object identification information whose reliability is below the predetermined threshold. However, spatially neighboring information may also be used. For example, if the camera (sensors) 12 is monitoring the congestion situation within its angle of view and part of the angle of view is obstructed by an obstruction, reevaluation can be performed using the detection results of the surrounding area that is not obstructed in addition to the situation before and after in time. More specifically, when the reliability value of the object identification information of the monitoring target 26 falls below a predetermined threshold, the meta information reanalysis unit 17 reanalyzes the video information. If the object identification information of the monitoring target 26 obtained by the reanalysis matches the object identification information immediately before the reliability value fell below the predetermined threshold, the meta information reanalyzed by the reanalysis replaces the object identification information whose reliability value fell below the predetermined threshold. Furthermore, because the video recording device 10 aggregates the detection results of one or more cameras (sensors) 12, it is also possible to use the detection results of other cameras (sensors) 12 as nearby. Even when it is necessary to perform advanced analysis that is not possible with a single camera (sensor) 12 in terms of processing power, the video recording device 10 is able to perform reanalysis using the video data, which is primary information, because it also aggregates video data. Specifically, when the reliability of object identification information output by one camera (sensor) 12 falls below a predetermined threshold, the meta information reanalysis unit 17 replaces the object identification information whose reliability falls below the predetermined threshold with object identification information output by other cameras (sensors) 12 whose reliability is higher than the predetermined threshold.

[0030] As described above, according to the first embodiment, a video surveillance system, a video recording device, a video analysis method, and a video analysis program can be provided that can reconstruct unreliable object identification information into highly reliable object identification information. In the embodiment, unreliable analysis results can be replaced with data generated from nearby highly reliable analysis results, thereby smoothing out temporary fluctuations in the analysis results that may otherwise cause noise, and providing the user with more reliable analysis results. Furthermore, by performing re-evaluation using the video recording device 10 that aggregates multiple cameras (sensors) 12, re-evaluation can be performed using the analysis results of the other cameras (sensors) 12, thereby providing the user with more reliable analysis results. Furthermore, by having the video recording device 10 bear the processing load related to the re-evaluation, the reliability of the analysis results can be improved without increasing the burden on the cameras (sensors) 12, which have relatively low processing power, thereby reducing costs.

[0031] Second Embodiment Next, a video monitoring system 110 according to a second embodiment will be described. The video monitoring system 110 according to the second embodiment shown in FIG. 7 differs from the video monitoring system 100 according to the first embodiment in that the video recording device 20 further includes a processing load monitoring unit 19. However, since the other configurations are the same as those of the first embodiment, the same components as those of the first embodiment are assigned the same reference numerals, and detailed description thereof will be omitted. Furthermore, since the hardware configuration of the video monitoring system 110 according to the second embodiment is the same as that of the video monitoring system 100 according to the first embodiment, detailed description thereof will be omitted. However, in the second embodiment, the CPU 21 executes a video analysis program to implement a video analysis method having a meta information reanalysis function that replaces a portion of meta information that has become unreliable with meta information in the vicinity of the portion that can be replaced, and a processing load monitoring function that monitors the processing load of the video recording device 20. As a result, the CPU 21 functions as the meta information reanalysis unit 17 and the processing load monitoring unit 19 by executing the video analysis program.

[0032] 8 is a flowchart showing an example of the process of re-evaluating meta information in the video recording device 20 according to embodiment 2. In step S201, the processing load monitoring unit 19 checks the load status of the video recording device 20 and determines whether or not to re-evaluate the meta information.

[0033] Re-evaluating meta information requires computationally intensive processing, such as re-analyzing and aligning multiple camera images, but such computationally intensive processing may not be possible in real time depending on the performance of the CPU 21, the current computational load of the CPU 21, or the free space of the main memory 22. In step S201, the processing load monitoring unit 19 monitors the performance of the CPU 21, the current computational load of the CPU 21, and the free space of the main memory 22, and determines whether the requested re-evaluation of meta information is possible. The processing load monitoring unit 19 also detects time periods when the CPU 21 is not operating or when the load due to the operation of the CPU 21 is low, and determines whether the re-evaluation of meta information is possible during those time periods.

[0034] If it is determined in step S201 that the re-evaluation of the meta information is executable, the procedure proceeds to step S202. If it is determined that the re-evaluation of the meta information is not executable, the procedure waits until it is executable. As described above, the processing load monitoring unit 19 can determine whether the re-evaluation of the meta information can be performed during a time period when the CPU 21 is not operating or when the load due to the operation of the CPU 21 is low. Therefore, even if it is determined in step S201 that the re-evaluation of the meta information is not executable due to the performance of the CPU 21, the current calculation load of the CPU 21, a lack of free space in the main memory 22, or the like, it is possible to re-evaluate the detection results later. The time period when the CPU 21 is not operating or when the load due to the operation of the CPU 21 is low is detected, for example, by referring to a job management system provided in the OS (Operating System) running on the CPU 21.

[0035] Step S202 is the same as step S101 in Fig. 6, step S203 is the same as step S102 in Fig. 6, and step S204 is the same as step S103 in Fig. 6. Therefore, from step S202 onwards, the same procedure as the process of re-evaluating meta information in embodiment 1 is carried out, and the process ends.

[0036] As explained above, according to the second embodiment, the processing load monitoring unit 19 checks the current load status of the video recording device 20 and re-evaluates the meta information during periods when the processing load is low, so that even if the amount of calculation required for the re-evaluation is large and it cannot be executed in real time, it is possible to re-evaluate the meta information later. Furthermore, since the video recording device 20 is configured not to analyze the meta information during periods when the operating load is high, it has the effect of not reducing the response of the video recording device 20 to user operations.

[0037] Third Embodiment Next, a video monitoring system 120 according to a third embodiment will be described. The video monitoring system 120 according to the third embodiment shown in FIG. 9 differs from the video monitoring system 100 according to the first embodiment in that the video recording device 33 further includes a video analysis unit 34 and an inference model recording unit 35. However, since the other configurations are the same as those of the first embodiment, the same components as those of the first embodiment are assigned the same reference numerals, and detailed description thereof will be omitted. Furthermore, since the hardware configuration of the video monitoring system 120 according to the third embodiment is the same as that of the video monitoring system 100 according to the first embodiment, detailed description thereof will be omitted. However, in the third embodiment, when the CPU 21 executes a video analysis program, for meta information whose reliability has become low, the video analysis unit 34 acquires recorded video from the storage unit 24 at a time close to the relevant location, and uses the inference model recording unit 35 to perform video analysis using machine learning to generate new, highly reliable meta information, thereby acting as a meta information reanalysis function that replaces the low-reliability meta information with nearby, more reliable meta information. As a result, the CPU 21 functions as a meta information reanalysis unit 17 and a video analysis unit 34 by executing the video analysis program.

[0038] Fig. 10 is an image diagram of the inference processing of the video monitoring system 120 in embodiment 3. As with Fig. 3, the right diagram of Fig. 10 shows a time series of sections G1 to G5 when a blocking object 27 passes in front of the camera (sensors) 12 while the monitoring target 26 is captured within the captureable range (angle of view). Sections G1 to G5 are equivalent to the sections shown in Fig. 4, and are assumed to have the same reliability.

[0039] The left diagram of Fig. 10 assumes a case in which monitored object 28 passes in front of camera (sensors) 12 within the range (angle of view) that can be photographed, along the same route as monitored object 26, and time passes through sections G11 to G15 that are similar to sections G1 to G5, as in the case of the right diagram of Fig. 10. Because monitored object 28 is not affected by obstruction 27, the reliability that can be obtained from the camera (sensors) is always a high value. Because monitored object 26 and monitored object 28 follow the same route, it can be assumed that the reliability that can be obtained for sections G11 to G15 of monitored object 28 will become learning data that indicates the ideal value of reliability for sections G1 to G5. In the case shown in Figure 10, if information on sections G1, G2, G4, and G5 can be obtained, and if machine learning can be performed to determine the relationship between reliability and similar sections G11, G12, G14, and G15, it will be possible to infer that the low-reliability section G3 has reliability equivalent to section G13.

[0040] FIG. 11 is a processing overview diagram for the video analysis unit 34 and the inference model recording unit 35 in the learning phase. In the learning phase, the video analysis unit 34 stores video information 40 obtained from the video recording unit 15, section information 41 (sections G11 to G15), and highly reliable meta information (correct answer) 42 obtained from the meta information recording unit 16 as learning data 46 in the inference model recording unit 35 via a learning program 47. During inference, the video analysis unit 34 receives video information 43 and section information 44 (sections G1 to G5) as input and generates meta information 45 as output. The inference model recording unit 35 trains the learning data created based on a combination of the video information, section information, and meta information (correct answer). As a result, the video surveillance system 120 generates an inference model that receives the video information 40 and section information 41 as input data (input 1) and generates meta information 45 as output data (output).

[0041] Next, the process of generating an inference model in the third embodiment will be described with reference to Fig. 12. Fig. 12 is a flowchart showing an example of the process of generating an inference model in the video recording device 33 according to the third embodiment.

[0042] In step S301, highly reliable section information and meta information of the corresponding section are obtained from the meta information recording unit 16.

[0043] In step S302, video information is acquired from the video recording unit 15 based on the section information acquired in step S301.

[0044] In step S303, the acquired video information, section information, and meta information are used as learning data.

[0045] In step S304, an inference model is generated by learning using the training data and recorded in the inference model section, and the process ends.

[0046] Next, the process of re-evaluating meta information, which is a detection result, in the third embodiment will be described with reference to Fig. 13. Fig. 13 is a flowchart showing an example of the process of re-evaluating meta information in the video recording device 33 according to the third embodiment.

[0047] In step S401, a search is made for a location where the reliability of the meta information, which is the detection result, is low, such as in the section G2 to G4 in Fig. 4. Specifically, a location where the reliability value is equal to or less than a predetermined threshold is detected.

[0048] In step S402, the video analysis unit 34 acquires from the video recording unit 15 video information for a time period around a location where the reliability of the meta information, which is a detection result, is low.

[0049] In step S403, it is determined whether or not the location where the reliability of the meta information detected as a result of the detection is low can be replaced with more reliable meta information generated by a trained inference model from video information acquired by the video analysis unit 34 at a time close to the location in question. If it is determined in step S403 that the location can be replaced with meta information generated from video information at a time close to the location, the procedure proceeds to step S404; if it is determined that the location cannot be replaced with meta information generated from video information at a time close to the location, the processing is terminated.

[0050] In step S404, the part where the reliability of the meta information is low is replaced with meta information generated from video information in the vicinity of the part, and the process ends.

[0051] As described above, according to the third embodiment, it is possible to provide a video surveillance system, a video recording device, a video analysis method, and a video analysis program that can reconstruct unreliable object identification information into highly reliable object identification information. In the third embodiment, it is possible to replace unreliable analysis results with data generated from highly reliable analysis results. While the camera (sensors) 12 alone needs to produce analysis results in a short period of time, the video recording device 33 can perform analysis taking into account past and future video for the analysis point, thereby enabling more reliable analysis results to be obtained than with the camera (sensors) 12 alone. Furthermore, this process does not assume that the video recording device will perform continuous evaluation, but rather that the meta information from the camera (sensors) 12 will be reevaluated during periods when the reliability of the information is low. By following the second embodiment, it is also possible to achieve the effect of not reducing the response time to user operations.

[0052] Fourth Embodiment Next, a video monitoring system 130 according to a fourth embodiment will be described. Meta information obtained by analyzing images captured by cameras or sensors has the problem of false detection due to the influence of the imaging environment in which the monitored object is reflected, such as in a mirror or window. Detection of reflections in mirrors or windows tends to be unreliable because the time required depends on the angle of reflection and the distance from the camera or sensors, and the entire image is rarely captured.

[0053] The video monitoring system 130 according to the fourth embodiment shown in FIG. 14 differs from the video monitoring system 100 according to the first embodiment in that the video recording device 36 further includes a video analysis unit 37 and an inference model recording unit 38. However, since the other configurations are the same as those of the first embodiment, the same components as those of the first embodiment are assigned the same reference numerals, and detailed descriptions thereof will be omitted. Furthermore, since the hardware configuration of the video monitoring system 130 according to the fourth embodiment is the same as that of the video monitoring system 100 according to the first embodiment, detailed descriptions thereof will be omitted. However, in the fourth embodiment, the CPU 21 executes a video analysis program. For a short section of time in which unreliable meta information is detected, the video analysis unit 37 acquires recorded video from the memory unit 24 for a time period near that location, and uses the inference model recording unit 38 to perform machine learning video analysis to determine whether erroneous meta information has been detected due to reflections from objects such as mirrors or windows. If erroneous meta information is detected, the inference model recording unit 38 functions as a meta information reanalysis function that deletes the meta information. As a result, the CPU 21 functions as a meta information reanalysis unit 17, a video analysis unit 37, and an inference model recording unit 38 by executing the video analysis program.

[0054] 15 is an image diagram of the installation of cameras (sensors) 12 of a video monitoring system 130 in the fourth embodiment. In the case shown in FIG. 15, it is assumed that the monitored object 26 passes in front of a highly reflective reflector 50, such as a window, while being captured within the captureable range (angle of view). The camera (sensors) 12 accurately captures the monitored object 26 up to a certain point, but as the camera (sensors) 12 passes in front of the reflector 50, a reflection occurs, and the camera (sensors) 12 captures a mirror image 29 similar to the monitored object 26 reflected by the reflector 50, in addition to the monitored object 26. Eventually, as the monitored object 26 moves, the mirror image 29 moves out of the angle of view of the camera (sensors) 12, and eventually only the monitored object 26 is captured in the angle of view of the camera (sensors) 12 again.

[0055] FIG. 16 is a time-series graph of the reliability 39 of the object identification information related to the mirror image 29 detected by the camera (sensors) 12 in the case shown in FIG. 15. The reliability 39 of the detection and recognition result of the monitoring target 26 by the camera (sensors) 12 is initially undetectable in section G21 because the mirror image 29 is not captured within the field of view of the camera (sensors) 12. However, as the monitoring target 26 moves, a portion of the mirror image 29 is captured, increasing in section G22 and reaching a maximum value in section G23. Eventually, as the monitoring target 26 moves and the mirror image 29 moves out of view in section G24, the reliability 39 of the detection and recognition result decreases again. In section G25, when the shape of the mirror image 29 is completely captured, the reliability 39 becomes undetectable, the same as in section G21.

[0056] As shown in Figure 16, in sections G22 to G24, the mirror image 29 is captured in the field of view of the camera (sensors) 12, so that the reliability 39 of the object identification information relating to the mirror image 29 is erroneously detected separately from the monitored object 26, resulting in noise that deteriorates the overall recognition results for the monitored object 26.

[0057] 17 is an explanatory diagram showing an example of re-evaluation of the detection results according to embodiment 4. For the meta information of the erroneously detected section G22 to G24, the video information of the section G22 to G24 acquired from the video recording unit 15 is applied to the inference model recording unit 38 to detect the enantiomer. Specifically, meta-information indicating highly reflective windows, mirrors, etc. is registered in advance in meta-information recording unit 16, and by having this learn video information containing mirror images 29 when monitored object 26 passes by, when a similar pattern occurs in the imaging environment, the mirror images are detected by inference model recording unit 38. To distinguish highly reflective windows, mirrors, etc., the user may register meta-information indicating highly reflective windows, mirrors, etc. in advance, or the camera (sensors) 12 may recognize and register the information as attribute information.

[0058] The process of reevaluating meta information, which is a detection result, in the fourth embodiment will be described with reference to Fig. 18. Fig. 18 is a flowchart showing an example of the process of correcting meta information in the video recording device 36 according to the fourth embodiment.

[0059] In step S501, a search is performed for a section where low-reliability meta information, which is a detection result, is detected over a short section, such as sections G22 to G24 in FIG. 16. Specifically, the search is performed for a section where the reliability value is equal to or less than a predetermined threshold, and the time during which the reliability value is equal to or less than the predetermined threshold is equal to or less than a predetermined value. The threshold is, for example, the minimum reliability value of the video information when a machine-learned model such as a CNN can recognize the monitoring target from the video information. The predetermined time value is specifically determined based on the length of the monitoring target 26, such as a window or mirror, in the movement direction, and the movement speed of the monitoring target 26.

[0060] In step S502, the video analysis unit 37 acquires from the video recording unit 15 the video information corresponding to the search result in step S501.

[0061] In step S503, the inference model recording unit 38 is used to determine whether the acquired image information contains a mirror image reflected in a window or mirror. In step S503, if it is determined that the acquired image information contains an enantiomer, i.e., if it is determined that an enantiomer has been detected from the acquired image information, the procedure proceeds to step S404, and if it is determined that the acquired image information does not contain an enantiomer, the processing is terminated.

[0062] In step S504, the portion where it is determined that an enantiomer has been detected from the acquired image information is determined to be a false detection, and the meta information corresponding to that enantiomer is deleted, and the process ends.

[0063] As described above, according to the fourth embodiment, it is possible to provide a video surveillance system, a video recording device, a video analysis method, and a video analysis program that can detect and correct object identification information that has been erroneously detected due to a highly reflective window or mirror, and reconstruct it into accurate object identification information. In the fourth embodiment, it is possible to delete object identification information that has been erroneously detected due to a highly reflective window or mirror. While the camera (sensors) 12 alone needs to produce analysis results in a short period of time, the video recording device 33 can perform analysis taking into account past and future images of the analysis point, making it possible to obtain analysis results that are more reliable than those obtained by the camera (sensors) 12 alone. By following the second embodiment, it is also possible to achieve the effect of not reducing the response to user operations. [Explanation of symbols]

[0064] 10 Video recording device, 12 Camera (sensors), 14 Video analysis unit, 15 Video recording unit, 16 Meta information recording unit, 17 Meta information reanalysis unit, 18 Display device, 19 Processing load monitoring unit, 20 Video recording device, 21 CPU, 22 Main memory, 24 Memory unit, 26 Monitoring target, 27 Obstruction, 28 Monitoring target, 29 Enantiomer, 30 Reliability, 32 Reconstructed reliability, 33 Video recording device, 34 Video analysis unit, 35 Inference model recording unit, 36 Video recording device, 37 Video analysis unit, 38 Inference model recording unit, 39 Reliability, 40 Video information, 41 Section information, 42 Reliability (correct answer), 43 Video information, 44 Section information, 45 Reliability, 46 Training data, 47 Training program, 48 Reconstructed reliability, 50 Reflectors, 100, 110, 120, 130 Video surveillance systems.

Claims

1. A video surveillance system including one or more sensors that acquire video information, analyze the acquired video information, and output meta information, and a video recording device that records the video information and the meta information output by the sensors, the sensor has a video analysis unit that analyzes the acquired video information and outputs, as the meta information, object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; The video recording device is a video recording unit that records the video information; a meta information recording unit for recording the meta information; a meta information reanalysis unit that reconstructs object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold. A video surveillance system characterized by:

2. The meta information reanalysis unit examines the meta information in chronological order, and reconstructs object identification information whose reliability is equal to or less than the predetermined threshold based on object identification information that exists before and after the object identification information in the chronological order and whose reliability is higher than the predetermined threshold.

2. The video monitoring system according to claim 1.

3. The meta information reanalysis unit reanalyzes the video information when the reliability value of the object identification information of the object falls below the predetermined threshold, and if the object identification information of the object obtained by the reanalysis matches the object identification information immediately before the reliability value fell below the predetermined threshold, adopts the object identification information obtained by the reanalysis in place of the object identification information whose reliability value fell below the predetermined threshold.

2. The video monitoring system according to claim 1.

4. When the reliability of the object identification information output by one of the sensors falls below the predetermined threshold, the meta information reanalysis unit reconstructs the object identification information output by the other sensors whose reliability is higher than the predetermined threshold and whose reliability falls below the predetermined threshold.

2. The video monitoring system according to claim 1.

5. the video recording device further includes a processing load monitoring unit that monitors a processing load of the video recording device; The meta information re-analysis unit reconstructs object identification information whose reliability indicated by the reliability information is equal to or less than the predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold during a period when the processing load of the video recording device obtained from the processing load monitoring unit is low.

2. The video monitoring system according to claim 1.

6. 1. A computer-implemented method for video analysis, comprising: recording video information; recording meta information including object identification information indicating an object included in the video information and reliability information indicating reliability of the object identification information; reconstructing object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold; A video analysis method comprising:

7. recording video information; recording meta information including object identification information indicating an object included in the video information and reliability information indicating reliability of the object identification information; reconstructing object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold; A video analysis program that runs on a computer.

8. a video recording unit that records video information; a meta information recording unit that records meta information including object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; a meta information reanalysis unit that reconstructs object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold based on object identification information whose reliability is higher than the predetermined threshold. A video recording device characterized by:

9. A video surveillance system including one or more sensors that acquire video information, analyze the acquired video information, and output meta information, and a video recording device that records the video information and the meta information output by the sensors, the sensor has a video analysis unit that analyzes the acquired video information and generates, as the meta information, object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; The video recording device is a video recording unit that records the video information; a meta information recording unit for recording the meta information; an inference model recording unit that generates the meta information using an inference model trained on learning data generated based on the highly reliable video information; and a meta information reanalysis unit that reconstructs object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold, based on meta information generated by the inference model recording unit by analyzing video information at a time close to the video information corresponding to the meta information. A video surveillance system characterized by:

10. 1. A computer-implemented method for video analysis, comprising: recording video information; recording meta information including object identification information indicating an object included in the video information and reliability information indicating reliability of the object identification information; generating the meta information using an inference model trained on training data generated based on the highly reliable video information; reconstructing object identification information, the reliability of which is indicated by the reliability information included in the meta information being equal to or less than a predetermined threshold, based on meta information generated by analyzing video information at a time close to the video information corresponding to the meta information; A video analysis method comprising:

11. a video recording unit that records video information; a meta information recording unit that records meta information including object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; an inference model recording unit that generates the meta information using an inference model trained on learning data generated based on the highly reliable video information; and a meta information reanalysis unit that reconstructs object identification information whose reliability indicated by the reliability information included in the meta information is equal to or less than a predetermined threshold, based on meta information generated by the inference model recording unit by analyzing video information at a time close to the video information corresponding to the meta information. A video recording device characterized by:

12. A video surveillance system including one or more sensors that acquire video information, analyze the acquired video information, and output meta information, and a video recording device that records the video information and the meta information output by the sensors, the sensor has a video analysis unit that analyzes the acquired video information and outputs, as the meta information, object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; The video recording device is a video recording unit that records the video information; a meta information recording unit for recording the meta information; an inference model recording unit that detects the mirror image of the object from the image information using an inference model trained with learning data generated based on image information containing the mirror image of the object; and a meta information reanalysis unit that deletes meta information corresponding to the mirror image when the inference model recording unit detects the mirror image from the video information. A video surveillance system characterized by:

13. 1. A computer-implemented method for video analysis, comprising: recording video information; recording meta information including object identification information indicating an object included in the video information and reliability information indicating reliability of the object identification information; detecting the mirror image of the object from the image information using an inference model trained with training data generated based on image information containing the mirror image of the object; a step of deleting meta information corresponding to the mirror image when the mirror image is detected from the video information; A video analysis method comprising:

14. a video recording unit that records video information; a meta information recording unit that records meta information including object identification information that indicates an object included in the video information and reliability information that indicates reliability of the object identification information; an inference model recording unit that detects the mirror image of the object from the image information using an inference model trained with learning data generated based on image information containing the mirror image of the object; and a meta information reanalysis unit that deletes meta information corresponding to the mirror image when the inference model recording unit detects the mirror image from the video information. A video recording device characterized by:

Citation Information

Patent Citations

  • JP1973066754A