Video surveillance using neural networks
A neural network-based system addresses scalability and accuracy issues in video surveillance by training an inference engine with operator-labeled data, enabling efficient event detection and reducing reliance on human operators.
Patent Information
- Application Number
- JP2024159872
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2018-07-05
- Filing Date
- 2024-09-17
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2039-07-05
AI Technical Summary
Traditional video surveillance systems relying on human operators for monitoring video feeds face scalability and accuracy issues due to the overwhelming number of feeds, leading to potential overburdening and reduced detection efficiency.
Implementing a neural network with an inference engine trained using operator-labeled video segments to detect events, allowing for deployment of trained inference engines at monitoring locations to augment human operators and improve scalability and accuracy.
The neural network enhances the scalability and accuracy of video surveillance systems by autonomously detecting events, reducing the burden on human operators and improving event detection capabilities.
Smart Images

Figure 0007718023000001 
Figure 0007718023000002 
Figure 0007718023000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to video surveillance, and more particularly to video surveillance using neural networks. [Background technology]
[0002] Traditionally, video surveillance systems have been used for security monitoring by large organizations, such as large corporations, government agencies, and educational institutions. Such video surveillance systems typically employ video cameras to cover the area being monitored and return video feeds to a central monitoring facility, such as a security office. The central monitoring facility typically includes one or more monitoring stations staffed by one or more human operators who view the monitored video feeds and flag events of interest. In some instances, the monitoring stations enable the human operators to log events of interest and take appropriate corrective action, such as raising an alarm, contacting first responders, etc.
[0003] In more recent years, as the cost of video surveillance cameras has decreased, the use of video surveillance systems in other settings has increased. For example, it has become common to see homes, small businesses, parks, and common areas equipped for monitoring with video surveillance systems. For example, such video surveillance systems may utilize low-cost cameras and / or any other imaging sensors to monitor areas of interest. These cameras often include a network interface that allows them to be connected to a network, enabling the cameras to transmit their respective video feeds to one or more remote monitoring facilities. These remote monitoring facilities also utilize monitoring stations staffed with human operators who view the monitored video feeds, flag events of interest, and take appropriate action in response to the events of interest. [Brief explanation of the drawings]
[0004] [Figure 1] FIG. 1 is a block diagram of an example video surveillance system.
[0005] [Figure 2] FIG. 1 is a block diagram of a first example video surveillance system including a trainable neural network constructed to support video surveillance according to the teachings of the present disclosure.
[0006] [Figure 3] FIG. 1 is a block diagram of a second example video surveillance system including a trainable neural network and a trained inference engine constructed to support video surveillance according to the teachings of the present disclosure.
[0007] [Figure 4] FIG. 10 is a block diagram of a third example video surveillance system including a trainable neural network and a trained inference engine constructed to support video surveillance according to the teachings of the present disclosure.
[0008] [Figure 5] 5 shows a flowchart representing example machine-readable instructions that may be executed to implement the example video surveillance system of FIGS. [Figure 6] 5 shows a flowchart representing example machine-readable instructions that may be executed to implement the example video surveillance system of FIGS. [Figure 7] 5 shows a flowchart representing example machine-readable instructions that may be executed to implement the example video surveillance system of FIGS. [Figure 8] 5 shows a flowchart representing example machine-readable instructions that may be executed to implement the example video surveillance system of FIGS.
[0009] [Figure 9]FIG. 10 is a block diagram of an example processor platform configured to execute the example machine-readable instructions of one or more of FIGS. 5-8 to implement an example monitoring station included in the example video surveillance system of FIGS. 2-4.
[0010] [Figure 10] FIG. 9 is a block diagram of an example processor platform configured to execute the example machine-readable instructions of one or more of FIGS. 5-8 to implement the example database included in the example video surveillance system of FIGS. 2-4.
[0011] [Figure 11] FIG. 10 is a block diagram of an example processor platform configured to execute the example machine-readable instructions of one or more of FIGS. 5-8 to implement the example neural network included in the example video surveillance system of FIGS. 2-4.
[0012] The drawings are not to scale. Wherever possible, the same reference numbers are used throughout the drawings and the accompanying written description to refer to the same or like parts, elements, etc. DETAILED DESCRIPTION OF THE INVENTION
[0013]
[0003] Example methods, apparatus, systems, and articles of manufacture (e.g., physical storage media) for implementing video surveillance using neural networks are disclosed herein. The example video surveillance system disclosed herein includes a database for storing operator-labeled video segments (e.g., as operator-labeled video segment records). The operator-labeled video segments include reference video segments and corresponding reference event labels that describe the reference video segments. The disclosed example video surveillance system also includes a neural network including a first instance of an inference engine and a training engine for training the first instance of the inference engine based on a training set of operator-labeled video segments retrieved from the database. In the disclosed example, the first instance of the inference engine infers events from the operator-labeled video segments included in the training set. The disclosed example video surveillance system further includes a second instance of the inference engine for inferring events from monitored video feeds, the second instance of the inference engine being based on (e.g., initially a duplicate of) the first instance of the inference engine.
[0014] In some disclosed examples, the reference event label indicates whether the corresponding reference video segment depicts the defined event.
[0015] In some disclosed examples, a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents a type of event.
[0016] In some disclosed examples, a video surveillance system includes a monitoring station, in some disclosed examples, the monitoring station including a display for presenting a first one of the monitored video feeds and a monitoring interface for generating an operator event label based on an operator determination corresponding to a monitored video segment of the first one of the monitored video feeds.
[0017] In some such disclosed examples, the database communicates with the monitoring station to receive a first monitored video segment of the monitored video segments and a first operator event label of the operator event labels corresponding to the first monitored video segment of the monitored video segments. In some such examples, the database stores the first monitored video segment of the monitored video segments and the corresponding first operator event label of the operator event labels as a first reference video segment of the reference video segments and a corresponding first reference event label of the reference event labels included in the first operator-labeled video segment of the operator-labeled video segments.
[0018] Additionally or alternatively, in some such disclosed examples, the monitoring station further implements a second instance of the inference engine. For example, the second instance of the inference engine can output an inferred event for a second one of the monitored video segments of the first one of the monitored video feeds presented by the display of the monitoring station. In some such examples, the monitoring interface of the monitoring station generates a second one of the operator event labels from the operator's judgment detected for the second one of the monitored video segments. In some such examples, the monitoring station further includes a comparator for comparing the inferred event with the second one of the operator event labels to obtain updated training data. In some disclosed examples, the neural network is in communication with the monitoring station to receive the updated training data, and the neural network training engine is for retraining the first instance of the inference engine based on the updated training data.
[0019] These and other example methods, apparatus, systems and articles of manufacture (eg, physical storage media) for implementing video surveillance using neural networks are disclosed in more detail below.
[0020] Video surveillance systems are commonly used by large organizations, such as large corporations, government agencies, and educational institutions, for security monitoring. In more recent years, as the cost of video surveillance cameras has decreased, the use of video surveillance systems in other settings has increased. For example, homes, small businesses, parks, and common areas equipped for video surveillance monitoring have become commonplace. Such new video surveillance systems may utilize low-cost cameras and / or any other imaging sensors that can connect to a network, such as the Internet, one or more cloud services accessible via the Internet and / or other networks, and the like. Such network access enables the cameras and / or other imaging sensors to transmit their respective video feeds to one or more remotely located monitoring facilities. These remote monitoring facilities typically utilize monitoring stations staffed by human operators who view the monitored video feeds, flag events of interest, and take appropriate action in response to the events of interest. As the cost of video surveillance cameras / sensors and associated network technology continues to decrease, it is expected that the use of video surveillance will continue to increase, possibly exponentially. However, utilizing monitoring stations staffed by human operators to monitor the video feeds generated by a video surveillance system can limit the scalability of video surveillance and affect the accuracy with which events are detected. This is especially true when human operators are overburdened by the number of video feeds that must be monitored.
[0021] A video surveillance system implemented using a neural network disclosed herein provides a technical solution to the scalability and accuracy problems associated with traditional video surveillance systems that utilize human operators for video feed monitoring. The example video surveillance system disclosed herein includes a neural network having an inference engine that is trained to detect or infer events in monitored video feeds. As disclosed in further detail below, the neural network is trained using a training set of reference video segments having reference event labels determined based on judgments by the human operator. For example, the reference event labels may describe whether the corresponding reference video segment represents a defined event (e.g., security breach, human presence, package arrival, any specified / predetermined event, etc.). In some examples, the reference event labels also include a description of the type of event of interest (e.g., security breach, human presence, package arrival, etc.) and an indication of whether the corresponding reference video segment represents the described event.
[0022] As described in further detail below, some disclosed examples of video surveillance systems deploy instances of a trained inference engine at one or more monitoring locations to infer events from each monitored video feed. For example, the trained inference engines may operate (e.g., in parallel, asynchronously, collaboratively, etc.) to infer whether one or more trained events (e.g., security breaches, human presence, package arrival, etc.) are displayed in the corresponding video feeds they monitor. In some such examples, the trained inference engines may be deployed for execution by or in conjunction with monitoring stations staffed by human operators, thereby augmenting the monitoring performed by the human operators. In some such examples, inference events output by a trained inference engine operating in conjunction with a manned monitoring station can be compared with corresponding decisions made by the human operators to determine updated training data that can be used to improve the operation of the inference engine. For example, a neural network in a video surveillance system can receive updated training data, retrain its inference engine, and then redeploy instances of the retrained inference engine to one or more monitoring locations. In some examples, trained / retrained inference engine instances can be deployed to unmanned monitoring locations or can be deployed to manned locations to provide additional capacity, thereby allowing the capacity of the video surveillance system to easily scale as demand increases. These and other aspects of video surveillance using neural networks disclosed herein are described in further detail below.
[0023] Referring now to the figures, a block diagram of an example video surveillance system 100 is shown in FIG. 1. The example video surveillance system 100 of FIG. 1 includes example imaging sensors 105A-D, an example network 110, and example monitoring stations 115A-B. In the illustrated example of FIG. 1, the imaging sensors 105A-D are configured to monitor areas of interest, such as, but not limited to, areas in one or more commercial businesses, small businesses, government offices, educational institutions, homes, parks, common areas, etc., and / or any combination thereof. The example imaging sensors 105A-D can include any number, type, and / or combination of imaging sensors. For example, the imaging sensors 105A-D can be implemented with one or more video cameras, smartphones, photodiodes, photodetectors, etc.
[0024] 1 , imaging sensors 105A-D include network interfaces capable of communicating with example network 110. Example network 110 can be implemented by any number, type, and / or combination of networks. For example, network 110 can be implemented by the Internet and / or one or more cloud services accessible via the Internet. In some examples, imaging sensors 105A-D include network interfaces capable of accessing network 110 via one or more wireless access points (e.g., cellular access points / base stations, wireless local area network access points, Bluetooth access points, etc.), wired access points (e.g., Ethernet access points, cable communication links, etc.), or any combination thereof.
[0025] 1, imaging sensors 105A-D transmit their respective video feeds to monitoring stations 115A-B, each of which includes a network interface capable of communicating with network 110. In this manner, monitoring stations 115A-B communicate with imaging sensors 105A-D via network 110. The term "communication," including variations thereof, as used herein encompasses direct and / or indirect communication through one or more intermediary components and does not require direct physical (e.g., wired) and / or constant communication, but additionally includes periodic or aperiodic intervals and selectable communication of one-time events.
[0026] The illustrated example monitoring stations 115A-B may be implemented by any system / device capable of presenting monitored video feeds and accepting user input regarding the monitored video feeds. For example, the monitoring stations 115A-B may be implemented by any number, type, and / or combination of computing systems / devices, such as one or more computers, workstations, smartphones, tablet computers, personal digital assistants (PDAs), etc. In some examples, the monitoring stations 115A-B are implemented by a processor platform, such as the example processor platform 900 shown in FIG. 9, which is described in more detail below.
[0027] 1 , monitoring stations 115A-B include example respective displays 120A-B and example respective monitoring interfaces 125A-B to enable video feeds reported by imaging sensors 105A-D to be monitored by human operators 130A-B. In the illustrated example, each display 120A-B of monitoring stations 115A-B is configured to present one or more video feeds reported by imaging sensors 105A-D. For example, displays 120A-B can present one video feed, cycle between presenting multiple video feeds one at a time, present multiple video feeds simultaneously in a tiled manner, etc.
[0028] The monitoring interfaces 125A-B of each of the monitoring stations 115A-B are configured to accept input from a human operator 130A-B that reflects a determination made by the human operator 130A-B regarding whether an event is depicted in the video feed presented by the display 120A-B. For example, the monitoring interfaces 125A-B may include input buttons or keys (labeled "Y" and "N" in FIG. 1 ) that allow the human operator 130A-B to indicate whether a video segment of the monitored video feed depicts or otherwise displays an event of interest (e.g., "Y" indicates that the event of interest is depicted / displayed, and "N" indicates that the event of interest is not depicted / displayed). For example, the human operators 130A-B may be trained to detect events of interest, such as, but not limited to, security breaches, human presence, and package arrivals. In some such examples, the "Y" and "N" inputs of the monitoring interfaces 125A-B can be used by the human operators 130A-B to indicate whether an event of interest is depicted or otherwise displayed in a video segment of the monitored video feed. In some examples, the monitoring interfaces 125A-B can further include input / output functionality (labeled "I / O" in FIG. 1 ) to provide greater flexibility by allowing the human operators 130A-B to input descriptions of types of events. In such examples, the human operators 130A-B are not limited to monitoring only one or more events of interest, but can input descriptions of any type of event of interest. In this manner, this additional functionality allows events of interest to change over time.In some examples, a human operator 130A-B can use an "I / O" interface to input a description of the type of event being monitored in a particular video segment of a monitored video feed, and can use "Y" and "N" inputs to indicate whether the event is actually depicted or otherwise displayed in the monitored video segment.
[0029] 1 is shown as including four imaging sensors 105A-D, one network 110, and two monitoring stations 115A-B, the video surveillance system 100 is not limited to this. Rather, the example video surveillance system 100 may include any number of imaging sensors 105A-D, any number of networks 110, and any number of monitoring stations 115A-B.
[0030] A block diagram of a first example video surveillance system 200 that implements video surveillance using a neural network in accordance with the teachings of the present disclosure is shown in Figure 2. The example video surveillance system 200 of Figure 2 includes the example imaging sensors 105A-D, example network 110, and example monitoring stations 115A-B of the video surveillance system 100 of Figure 1. As such, aspects of these elements of video surveillance system 200 have been described above in conjunction with the description of Figure 1.
[0031] The example video surveillance system 200 of FIG. 2 also includes an example neural network 205 implemented in accordance with the teachings of the present disclosure and an example database 210. The example database 210 includes one or more network interfaces capable of communicating with the example network 110. In the illustrated example of FIG. 2, the database 210 communicates with monitoring stations 115A-B over the network 110 and receives video segments from the monitoring stations 115A-B with corresponding event labels that describe the video segments. The video segments received from the monitoring stations 115A-B are obtained from monitored video feeds. The event labels indicate whether the corresponding video segments depict or otherwise display an event of interest based on judgment information entered by human operators 130A-B. In some examples, the event labels of the corresponding video segments indicate whether the video segments depict or otherwise display a defined event. In some examples, the event label for the corresponding video segment includes a description of or otherwise indicates the type of event and whether the video segment depicts or otherwise displays the type of event.
[0032] In the illustrated example, the database 210 treats video segments with corresponding event labels received from the monitoring stations 115A-B as operator-labeled video segments used to train the neural network 205. For example, the database 210 treats the video segments received from the monitoring stations 115A-B as reference video segments 215 of the example operator-labeled video segments, and treats the corresponding event labels received from the monitoring stations 115A-B as reference event labels 218 corresponding to the reference video segments 215 of the operator-labeled video segments. Further, the database 210 generates records 220 of the example operator-labeled video segments, each record 220 including the reference video segment 215 and the corresponding reference event label 218 of the operator-labeled video segment represented by the record 220. In this manner, each record 220 of the operator-labeled video segments includes the reference video segment 215 and the corresponding reference event label 218 for the operator-labeled video segment represented by that record 220. Thus, in some examples, the database 210 implements a means for obtaining operator-labeled video segments, where the operator-labeled video segments include reference video segments and corresponding reference event labels that describe the reference video segments. Other means for obtaining operator-labeled video segments may include, but are not limited to, a computing device, server, cloud-based service, website, etc., configured to collect and combine video segments with corresponding event labels that describe the video segments to form operator-labeled video segments.
[0033] The illustrated example database 210 also includes an example record storage 225 for storing records 220 of operator-labeled video segments created by the database 210 from the video segments and corresponding event labels received from the monitoring stations 115A-B. The record storage 225 of the illustrated example database 210 may be implemented by any number and / or types of storage technologies, memory technologies, etc. For example, the database 210 may be implemented by any computing system / device, such as the example processor platform 1000 shown in FIG. 10 , described in further detail below. In such an example, the record storage 225 may be implemented by one or more of the example mass storage device 1028 and / or volatile memory 1014 of the example processor platform 1000.
[0034] The illustrated example neural network 205 includes an example inference engine 230 for inferring events from video segments, such as video segments from monitored video feeds acquired from imaging sensors 105A-D. In the illustrated example, the inference engine 230 is implemented by a convolutional neural network (CNN) inference engine including one or more weight layers (also referred to as neurons) trained to infer events from video segments. For example, the inference engine 230 may be constructed to include an input layer for accepting one or more input video segments 235 as input data, one or more hidden layers for processing the input data, and an output layer for providing one or more outputs indicating whether a given input video segment 235 represents or otherwise displays one or more events that the inference engine 230 is trained to detect. In some examples, the output of the inference engine 230 may additionally or alternatively provide a likelihood (e.g., probability) that a given input video segment 235 represents or otherwise displays one or more events that the inference engine 230 is trained to detect. Although the inference engine 230 of the neural network 205 in the illustrated example is implemented by a CNN, other neural network solutions can be used to implement the inference engine 230.
[0035] The illustrated example neural network 205 also includes an example training engine 240, an example comparator 245, and an example training data retriever 250 to train the inference engine 230 to infer events from input video segments 235. In the illustrated example, the training data retriever 250 retrieves from the database 210 a set of operator-labeled video segment records 220 to be used as training data for training the inference engine 230. For example, the training data retriever 250 may send a request to the database 210 for the set of operator-labeled video segment records 220. In some examples, the request includes the number of records 220 requested. In some examples, the request additionally or alternatively includes the type of event represented or otherwise indicated by the operator-labeled video segment records 220 included in the training set. In such an example, database 210 may respond to the request by retrieving the requested set of operator-labeled video segment records 220 from record storage 225 and transmitting the retrieved set of records 220 to training data retrieval unit 250.
[0036] In the illustrated example, after the requested set of training records 220 is retrieved from database 210, training data retriever 250 applies the training records 220 to inference engine 230 to train inference engine 230 to infer events from operator-labeled video segments included in the training set of records 220. For example, for a given one of the training records 220, training data retriever 250 applies the reference video segment of that training record 220 to inference engine 230 as an input video segment 235. Training data retriever 250 also applies the corresponding reference event label of that training record to the inferred event decision output from inference engine 230 as an example training event label 255 to be compared by comparator 245. Ideally, when trained, the inferred event decision output from inference engine 230 will match (e.g., produce zero error) the training event label 255 corresponding to the input video segment 235. However, while inference engine 230 is being trained, comparator 245 may detect errors between training event labels 255 and the inferred event decisions output from inference engine 230. In the illustrated example, the output of comparator 245 is provided to training engine 240, which feeds back the errors in any suitable manner to update the weight layers of inference engine 230 to the accuracy with which inference engine 230 infers events from input segments 235. For example, training engine 240 is shown as a backpropagation unit that performs backpropagation to train inference engine 230. However, any other suitable training mechanism may be implemented by training engine 240. Thus, in some examples, training engine 240 implements means for training a first instance of an inference engine based on a training set of operator-labeled video segments, where the first instance of the inference engine is for inferring events from operator-labeled video segments included in the training set.Other means for training the first instance of the inference engine based on the training set of operator-labeled video segments may include, but are not limited to, a computing device, server, cloud-based service, website, etc. configured to obtain the training set of operator-labeled video segments and apply the training set to any type of machine learning inference engine to train the inference engine.
[0037] In some examples, the training data retriever 250 continues to apply different ones of the training records 220 to the inference engine 230 until the comparator 245 indicates that a desired inference accuracy has been achieved. For example, the inference accuracy can be specified as a percentage threshold for correct event detection (e.g., corresponding to the percentage of the number of input reference video segments for which the inference engine 230 correctly infers whether a corresponding event is present as represented by a reference event label), a percentage threshold for false event detection (e.g., corresponding to the percentage of the number of input reference video segments for which the inference engine 230 erroneously infers that a corresponding event is present when an event is actually not present as indicated by the reference event label corresponding to the reference video segment), a threshold for missed event detection (e.g., corresponding to the percentage of the number of input reference video segments for which the inference engine 230 erroneously infers that a corresponding event is not present when an event is actually present as indicated by the reference event label corresponding to the reference video segment), etc.
[0038] 2 is shown as including four imaging sensors 105A-D, one network 110, two monitoring stations 115A-B, one neural network 205, and one database 210, video surveillance system 200 is not limited to this. Rather, example video surveillance system 200 may include any number of imaging sensors 105A-D, any number of networks 110, any number of monitoring stations 115A-B, any number of neural networks 205, and any number of databases 210.
[0039] A block diagram of a second example video surveillance system 300 that implements video surveillance using a neural network in accordance with the teachings of the present disclosure is shown in FIG. 3. The example video surveillance system 300 of FIG. 3 includes the example imaging sensors 105A-D, example network 110, and example monitoring stations 115A-B of the video surveillance systems 100 and 200 of FIGS. 1 and 2. Accordingly, aspects of these elements of the video surveillance system 300 are described above in conjunction with the description of FIGS. 1 and 2. The example video surveillance system 300 of FIG. 3 also includes the example neural network 205 and example database 210 of the video surveillance system 200 of FIG. 2. Accordingly, aspects of these elements of the video surveillance system 300 are described above in conjunction with the description of FIG. 2.
[0040] 3 , the neural network 205 also includes an example deployer 305 that deploys instances of the trained inference engine 230 for inferring events from monitored video feeds. Thus, in the illustrated example, the inference engine 230 corresponds to a first (or reference) instance of the inference engine 230, and the deployer deploys other instances of the inference engine 230 that are based on (e.g., initially clones of) the first (or reference) instance of the inference engine 230. For example, the deployer 305 deploys a second example instance 310A of the inference engine 230 for execution by or in conjunction with the example monitoring station 115A, and a third example instance 310B of the inference engine 230 for execution by or in conjunction with the example monitoring station 115B.
[0041] In some examples, the deployer 305 deploys instances of the inference engine 230, such as the second instance 310A and the third instance 310B, as data representing trained weight layers obtained by training a first instance of the inference engine 230 included in the neural network 205. In such examples, the deployer 305 downloads (e.g., via the network 110) the data representing the trained weight layers to instances of the inference engine 230, such as the second instance 310A and the third instance 310B, that are already present at the target monitoring locations. In some examples, the deployer 305 deploys instances of the inference engine 230, such as the second instance 310A and the third instance 310B, as downloadable executable files (e.g., downloaded via the network 110) to be executed by computing devices, such as the monitoring stations 115A-B. Thus, in some examples, the deployer 305 implements a means for deploying an instance of an inference engine for inferring events from monitored video feeds, where the deployed instance of the inference engine is based on (e.g., initially a clone of) the trained instance of the inference engine. Other means for deploying an instance of the inference engine may include, but are not limited to, a computing device, server, cloud-based service, website, etc. configured to obtain and deploy copies of instances of the trained inference engine.
[0042] 3, the second instance 310A of the inference engine 230 and the third instance 310B of the inference engine 230 are part of respective reinforced inference engines 315A and 315B executed by or otherwise implemented to operate in conjunction with the respective monitoring stations 115A-B. In the illustrated example, the reinforced inference engine 315A is executed by or otherwise implemented by the monitoring station 115A to infer events from the monitored video feeds processed by the monitoring station 115A. For example, the second instance 310A of the inference engine 230 included in the reinforced inference engine 315A accepts video segments of the video feeds monitored by the monitoring station 115A and outputs inferred events for the monitored video segments (e.g., an indication of whether a particular event is depicted in each of the monitored video segments). The reinforced inference engine 315A also includes an example comparator 320A for determining updated training data by comparing the inference events output by the second instance 310A of the inference engine 230 for the corresponding monitored video segments with respective operator event labels generated by the monitoring station 115A from operator decisions (e.g., input by the human operator 130A via the monitoring interface 125A, as described above) detected by the monitoring station 115A for the corresponding monitored video segments. The comparator 320A reports this updated training data to the neural network 205 (e.g., via the network 110).
[0043] 3, a reinforced inference engine 315B is executed or otherwise implemented by monitoring station 115B to infer events from monitored video feeds processed by monitoring station 115B. For example, a third instance 310B of inference engine 230 included in reinforced inference engine 315B accepts video segments of the video feed monitored by monitoring station 115B and outputs inferred events for the monitored video segments (e.g., indications of whether a particular event is depicted in each of the monitored video segments). Reinforced inference engine 315B also includes an example comparator 320B for determining updated training data by comparing the inferred events output by the third instance 310B of inference engine 230 for the corresponding monitored video segments with respective operator event labels generated by monitoring station 115B from operator judgments detected by monitoring station 115B for the corresponding monitored video segments (e.g., input by human operator 130B via monitoring interface 125B, as described above). Comparator 320B reports this updated training data to neural network 205 (eg, via network 110).
[0044] 3, the neural network 205 uses updated training data received from the reinforced inference engines 315A-B implemented by each monitoring station 115A-B to retrain the first instance of the inference engine 230 included in the neural network 205 to improve event inference accuracy. For example, the neural network 205 may retrain the first instance of the inference engine 230 at periodic intervals based on one or more events (e.g., when a threshold amount of updated training data is received from the reinforced inference engines 315A-B implemented by each monitoring station 115A-B), operator input, or the like, or any combination thereof. In some examples, the deployment unit 305 of the illustrated example neural network 205 then redeploys the retrained instance of the inference engine 230 to one or more target monitoring locations. For example, the deployer 305 may redeploy a retrained instance of the inference engine 230 to update / replace a second instance 310A implemented by the inference engine 230 of the monitoring station 115A and / or a third instance 310B of the inference engine 230 implemented by the monitoring station 115B.
[0045] 3 is shown as including four imaging sensors 105A-D, one network 110, two monitoring stations 115A-B implementing two reinforced inference engines 315A-B, one neural network 205, and one database 210, the video surveillance system 200 is not limited in this respect. Rather, the example video surveillance system 200 may include any number of imaging sensors 105A-D, any number of networks 110, any number of monitoring stations 115A-B implementing any number of reinforced inference engines 315A-B, any number of neural networks 205, and any number of databases 210.
[0046] A block diagram of a third example video surveillance system 400 implementing video surveillance using a neural network in accordance with the teachings of the present disclosure is shown in FIG. 4. The example video surveillance system 400 of FIG. 4 includes the example imaging sensors 105A-D, example network 110, and example monitoring stations 115A-B of the video surveillance systems 100, 200, and 300 of FIGS. 1-3. Accordingly, aspects of these elements of the video surveillance system 400 are described above in conjunction with the descriptions of FIGS. 1-3. The example video surveillance system 400 of FIG. 3 also includes the example neural network 205 and example database 210 of the video surveillance systems 200 and 300 of FIGS. 2-3. Accordingly, aspects of these elements of the video surveillance system 400 are described above in conjunction with the descriptions of FIGS. 2-3.
[0047] In the example video surveillance system 400 shown in FIG. 4 , the neural network 205 deployment unit 305 also deploys instances of the trained inference engine 230 for monitoring video feeds without coordination with the monitoring stations 115A-B. For example, in the video surveillance system 400, the neural network 205 deployment unit 305 can deploy instances of the trained inference engine 230 to unmanned monitoring locations. Additionally or alternatively, the neural network 205 deployment unit 305 can deploy instances of the trained inference engine 230 to monitoring locations that have monitoring stations, such as the monitoring stations 115A-B, but for operation independent of the monitoring stations. For example, in the video surveillance system 400 of FIG. 4 , the neural network 205 deployment unit 305 deploys a fourth instance 410A of the inference engine 230 and a fifth instance 410B of the inference engine 230 to monitor video feeds independent of the monitoring stations 115A-B. Thus, additional instances of the inference engine 230 can be deployed to increase the monitoring capacity in the video surveillance system 400 in a cost-effective manner.
[0048] 4 is shown as including four imaging sensors 105A-D, one network 110, two monitoring stations 115A-B implementing two reinforced inference engines 315A-B, one neural network 205, one database 210, and two separate instances 410A-B of inference engine 230, but video surveillance system 200 is not limited in this regard. Rather, example video surveillance system 200 may include any number of imaging sensors 105A-D, any number of networks 110, any number of monitoring stations 115A-B implementing any number of reinforced inference engines 315A-B, any number of neural networks 205, any number of databases 210, and any number of instances 410A-B of inference engine 230.
[0049] Additionally, while the illustrated example video surveillance systems 200, 300, and 400 include imaging sensors 105A-D, surveillance monitoring using neural networks as disclosed herein is not limited to video surveillance. For example, the neural network techniques disclosed herein can be adapted for use with other monitoring sensors. For example, video surveillance systems 100, 200, and / or 300 may include other sensors in addition to or in place of imaging sensors 105A-D. Such other sensors may include, but are not limited to, motion sensors, heat / temperature sensors, sound sensors (e.g., microphones), electromagnetic sensors, etc. In such examples, these sensors, possibly in conjunction with one or more of monitoring stations 115A-B, transmit respective data feeds over network 110 for monitoring by one or more of inference engines 230, 310A, 310B, 410A, and / or 410B.
[0050] Although example ways of implementing video surveillance systems 100, 200, 300, and 400 are illustrated in Figures 1-4, one or more of the elements, processes, and / or devices illustrated in Figures 1-4 may be combined, divided, rearranged, omitted, removed, and / or implemented in any other way. Furthermore, example video surveillance systems 100, 200, 300, and / or 400 of Figures 1-4, example imaging sensors 105A-D, example network 110, example monitoring stations 115A-B, example neural network 205, example database 210, example reinforced inference engines 315A-B, and / or example instances 310A-B and / or 410A-B of inference engine 230 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, example video surveillance systems 100, 200, 300 and / or 400, example imaging sensors 105A-D, example network 110, example monitoring stations 115A-B, example neural network 205, example database 210, example reinforced inference engines 315A-B, and / or any of example instances 310A-B and / or 410A-B of inference engine 230 may be implemented by one or more analog or digital circuits, logic circuits, programmable processors, application specific integrated circuits (ASICs), programmable logic devices (PLDs), and / or field programmable logic devices (FPLDs).When any of the apparatus or system claims of this patent are read to cover pure software and / or firmware implementations, at least one of the example video surveillance systems 100, 200, 300 and / or 400, example imaging sensors 105A-D, example network 110, example monitoring stations 115A-B, example neural network 205, example database 210, example reinforced inference engines 315A-B, and / or example instances 310A-B and / or 410A-B of inference engine 230 is expressly defined herein as including a non-transitory computer-readable storage device or storage disk, such as a memory, digital versatile disk (DVD), compact disk (CD), Blu-ray® disk, etc., that includes software and / or firmware. Still further, example video surveillance systems 100, 200, 300 and / or 400 may include one or more elements, processes and / or devices in addition to or instead of those shown in FIGS. 1 through 4, and / or may include more than one of any or all of the elements, processes and devices shown.
[0051] Flowcharts illustrating example machine-readable instructions for implementing example video surveillance systems 100, 200, 300, and / or 400, example imaging sensors 105A-D, example network 110, example monitoring stations 115A-B, example neural network 205, example database 210, example reinforced inference engines 315A-B, and / or example instances 310A-B and / or 410A-B of inference engine 230 are shown in Figures 5 through 8. In these examples, the machine-readable instructions include one or more programs for execution by a processor, such as processors 912, 1012, and / or 1112 shown in example processor platforms 900, 1000, and 1100 described below in connection with Figures 9 through 11. One or more programs, or portions thereof, may be embodied in software stored on a non-transitory computer-readable storage medium, such as a CD-ROM, floppy disk, hard drive, digital versatile disk (DVD), Blu-ray® disk, or memory associated with processor 912, 1012 and / or 1112, although the entire program or programs and / or portions thereof may alternatively be executed by devices other than processor 912, 1012 and / or 1112 and / or may be embodied in firmware or dedicated hardware (e.g., implemented in an ASIC, PLD, FPLD, discrete logic, etc.). Furthermore, although the example programs are described with reference to the flowcharts shown in FIGS. 5 through 8, many other ways of implementing the example video surveillance systems 100, 200, 300 and / or 400, the example imaging sensors 105A-D, the example network 110, the example monitoring stations 115A-B, the example neural network 205, the example database 210, the example reinforced inference engines 315A-B, and / or the example instances 310A-B and / or 410A-B of the inference engine 230 may alternatively be used.5 through 8, the order of execution of the blocks may be changed, and / or some of the described blocks may be changed, eliminated, combined, and / or subdivided into multiple blocks. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., discrete and / or integrated analog and / or digital circuits, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), comparators, operational amplifiers (op amps), logic circuits, etc.) configured to perform corresponding operations without executing software or firmware.
[0052] 5-8 may be implemented using coded instructions (e.g., computer- and / or machine-readable instructions) stored on a non-transitory computer- and / or machine-readable medium, such as a hard disk drive, flash memory, read-only memory, compact disk, digital versatile disk, cache, random access memory, and / or any other storage device or storage disk on which information is stored for any length of time (e.g., long-term, permanent, short-term, temporary buffer, and / or caching of information). As used herein, the term non-transitory computer-readable medium is expressly defined to include any type of computer-readable storage device and / or storage disk, to exclude propagating signals, and to exclude transmission media.
[0053] The terms "including" and "comprising" (and all forms and tenses thereof) are used herein as open-ended terms. Thus, when something is recited in a claim after any form of "including" or "comprising" (e.g., comprises, includes, comprising, including, etc.), it should be understood that the additional elements, terms, etc. may be present without departing from the scope of the corresponding claim. As used herein, the phrase "at least" is open-ended in the same way that the terms "including" and "comprising" are open-ended when used as a transition in a claim preamble. Also, as used herein, the terms "computer-readable" and "machine-readable" are considered equivalent unless otherwise noted.
[0054] An example program 500 that may be executed to implement one or more example monitoring stations 115A-B included in the example video surveillance systems 100, 200, 300, and 400 of FIGS. 1-4 is shown in FIG. 5. For convenience and without loss of generality, execution of the example program 500 is described from the perspective of the example monitoring station 115A operating in the example video surveillance system 400 of FIG. 4. With reference to the above figures and associated written description, the example program 500 of FIG. 5 begins execution at block 505, in which the monitoring station 115A accesses a video feed received from one of the imaging sensors 105A-D via the network 110 and presents the accessed video feed via the display 120A of the monitoring station 115A. At block 510, the monitoring station 115A detects an operator decision entered by a human operator 130A via the monitoring interface 125A of the monitoring station 105A. As described above, the operator's decision detected in block 510 indicates whether an event of interest is depicted or otherwise displayed in the monitored video feed. For example, the monitoring station 115A may segment the accessed video feed into monitored video segments based on any segmentation criteria. For example, the monitoring station 115A may segment the monitored video feed into consecutive video segments having a given time length (e.g., 15 seconds, 30 seconds, 60 seconds, etc., or any time length) based on transitions in the detected video feed and characteristics of the imaging sensors 105A-D providing the feed (e.g., sensor sweep rate, sensor capture rate, etc.). In such an example, the monitoring interface 125A associates the entered operator's decision with the particular monitored video segment being presented by the display 120A at that time.In some examples, if no operator judgment input is detected while a given video segment is being presented, monitoring interface 125A determines that the operator judgment is that no event of interest is depicted during the given video segment. As described above, in some examples, the input operator judgment is a Yes or No indication as to whether a predefined event of interest is depicted in the monitored video segment. However, in some examples, the input operator judgment may also include a description of the type of event of interest and a Yes or No indication as to whether the described event of interest is depicted in the monitored video segment.
[0055] In block 515, the monitoring interface 125A of the monitoring station 115A determines whether an input operator decision is detected. If an operator decision is detected (block 515), in block 520, the monitoring interface 125A accesses an event label associated with the current monitored video segment and updates the event label to reflect the input operator decision. For example, the event label may indicate whether a predefined event of interest is depicted in the corresponding monitored video segment. In some examples, the event label includes a description of the type of event of interest and an indication of whether the described event of interest is depicted in the corresponding monitored video segment. In some examples, if no operator decision input is detected while a given video segment is being presented, in block 520, the monitoring interface 125A updates the event label for the video segment to indicate that no event of interest is depicted in the corresponding video segment. As shown in the example of FIG. 5, monitoring station 105A also provides database 210 with the video segments of the video accessed in block 505 and the corresponding event labels accessed in block 520, enabling database 210 to generate records 220 of the operator-labeled video segments described above.
[0056] In the illustrated example of FIG. 5 , in block 525, the monitoring interface 125A of the monitoring station 115A determines whether the input operator judgment detected for the current monitored video segment corresponds to the detection of an event for which an alarm should be triggered. If an alarm should be triggered (block 525), in block 530, the monitoring interface 125A causes the alarm to be triggered. For example, in block 530, the monitoring interface 125A may automatically trigger an audio and / or visual alarm and contact an emergency response system to summon first responders, etc. In block 535, the monitoring station 115A determines whether video surveillance monitoring should continue. If video surveillance monitoring should continue (block 535), processing returns to block 505 and subsequent blocks, allowing the monitoring station 115A to continue monitoring the current video feed and / or other video feeds received by the monitoring station 115A from the imaging sensors 105A-D via the network 110.
[0057] An example program 600 that may be executed to implement the example neural network 205 and example database 210 included in the example video surveillance systems 200, 300, and 400 of Figures 2-4 is shown in Figure 6. For convenience and without loss of generality, execution of the example program 600 will be described in terms of the example neural network 205 and example database 210 operating in the example video surveillance system 400 of Figure 4. With reference to the above figures and related written description, the example program 600 of Figure 6 begins execution at block 605, where the database 210 stores as records 220 monitored video segments, corresponding event labels, and operator-labeled video segments received from monitoring stations 115A-B over the network 110, as described above. At block 610, the neural network 205 trains a first (or reference) instance of the example inference engine 230 to infer events from the video segments, as described above, using the records 220 of operator-labeled video segments generated and stored by the database 210. At block 615, the neural network 205 deploys the trained instance of the inference engine 230 to the monitoring site to infer events from the monitored video feeds, as described above.
[0058] An example program 700 that may be executed to implement the example database 210 included in the example video surveillance systems 200, 300, and 400 of FIGS. 2-4 and / or to perform the processing of block 605 of FIG. 6 is shown in FIG. 7. For convenience and without loss of generality, execution of the example program 700 will be described in terms of the example database 210 operating in the example video surveillance system 400 of FIG. 4. With reference to the above-mentioned figures and related written description, the example program 700 of FIG. 7 begins execution at block 705, where the database 210 receives monitored video segments and corresponding event labels from monitoring stations 115A-B over network 110, as described above. As discussed above, the event labels reflect judgments entered by a human operator regarding whether an event of interest is depicted or otherwise indicated in the corresponding monitored video segments. In block 710, as described above, the database 210 generates a record 220 of the operator-labeled video segment received from the video segment and the corresponding event label received in block 705, and stores the record 220 in the example record storage 225 of the database 210.
[0059] At block 715, database 210 determines whether a request for a set of training data has been received from neural network 205. As described above, in some examples, the request includes the number of records 220 requested to be included in the set of training data. In some examples, the request additionally or alternatively includes the type of event represented or otherwise indicated by records 220 included in the set of training data. If a request for training data is received (block 715), database 210 retrieves a training set of records 220 from record storage 225 that satisfies the request and outputs the training set of records 220 to neural network 205 (e.g., via network 110) to facilitate training of the neural network.
[0060] FIG. 8 illustrates an example program 800 that may be executed to implement the example neural network 205 included in the example video surveillance systems 200, 300, and 400 of FIGS. 2-4 and / or to perform the processing in blocks 610 and 615 of FIG. 6 . For convenience and without loss of generality, execution of the example program 800 will be described in terms of the example neural network 205 operating in the example video surveillance system 400 of FIG. 4 . With reference to the above-mentioned figures and related written description, the example program 800 of FIG. 8 begins execution at block 805, in which the training data retriever 250 of the example neural network 205 requests and obtains a set of training records 220 of operator-labeled video segments from the database 210. As discussed above, in some examples, the request includes the requested number of records 220 to include in the set of training data. In some examples, the request additionally or alternatively includes the type of event represented or otherwise indicated by the records 220 included in the set of training data.
[0061] At block 810, the example training engine 240 and the example comparator 245 train the example inference engine 230 of the neural network 205 by using the acquired training set of records 220 to infer events from reference video segments included in the training set of records 220, as described above. At block 815, the deployment unit 305 of the example neural network 205 deploys an instance of the trained inference engine 230 to one or more target monitoring locations to infer events from monitored video feeds, as described above. For example, at block 815, the deployment unit 305 may deploy example instances 310A-B of the trained inference engine 230 for execution by or in conjunction with example monitoring stations 105A-B. Additionally or alternatively, in some examples, the deployment unit 305 may deploy example instances 410A-B of the trained inference engine 230 to monitoring locations to perform video surveillance monitoring independent of the monitoring stations 105A-B.
[0062] At block 820, the neural network 205 obtains (e.g., via the network 110) updated training data determined by one or more of the monitoring stations 105A-B executing or operating in conjunction with the trained inference engine 230 instances 310A-B. For example, as described above, the monitoring stations 105A-B may determine the updated training data by comparing the inferred events output by the trained inference engine 230 instances 310A-B for the corresponding monitored video segments with the respective operator event labels generated by the monitoring stations 115A-B from operator judgments input by the human operators 130A-B for the corresponding monitored video segments. At block 825, the neural network retrains the first (or reference) instance of the inference engine 230 using the updated training data, as described above. At block 830, the deployer 305 redeploys the retrained inference engine 230 instances to one or more target monitoring locations, as described above. In block 835, the neural network 205 determines whether to continue retraining the inference engine 230. If so, processing returns to block 820 and subsequent blocks to allow the neural network 205 to continue retraining the inference engine 230 based on updated training data received from the monitoring stations 105A-B.
[0063] Figure 9 is a block diagram of an example processor platform 900 configured to execute the instructions of Figures 5, 6, 7, and / or 8 to implement the example monitoring stations 115A-B of Figures 1-4. For convenience and without loss of generality, the example processor platform 900 will be described in terms of implementing the example monitoring station 115A. The processor platform 900 may be, for example, a server, a personal computer, a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad), a personal digital assistant (PDA), an Internet appliance, or any other type of computing device.
[0064] The processor platform 900 of the illustrated example includes a processor 912. The processor 912 of the illustrated example is hardware. For example, the processor 912 can be implemented by one or more integrated circuits, logic circuits, microprocessors, or controllers of any desired family or manufacturer. The hardware processor 912 can be a semiconductor-based (e.g., silicon-based) device.
[0065] The processor 912 of the illustrated example includes a local memory 913 (e.g., a cache). The processor 912 of the illustrated example communicates with a main memory, including a volatile memory 914 and a nonvolatile memory 916, via a link 918. The link 918 may be implemented by a bus, one or more point-to-point connections, etc., or a combination thereof. The volatile memory 914 may be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), and / or any other type of random access memory device. The nonvolatile memory 916 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 914, 916 is controlled by a memory controller.
[0066] The processor platform 900 of the illustrated example also includes an interface circuit 920. The interface circuit 920 may be implemented by any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB), and / or a PCI Express interface.
[0067] In the illustrated example, one or more input devices 922 are connected to the interface circuitry 920. The input devices 922 allow a user to input data and commands into the processor 912. The input devices may be implemented, for example, by audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, trackpads, trackballs, trackbars (such as Isopoint), voice recognition systems, and / or any other human-machine interface. Many systems, such as the processor platform 900, may also allow a user to control a computer system and provide data to the computer using physical actions, including, but not limited to, hand and body movements, facial expressions, and face recognition. In some examples, the input data devices 922 implement the example monitoring interface 125A.
[0068] One or more output devices 924 are also connected to the interface circuitry 920 of the illustrated example. The output device(s) 924 can be implemented, for example, by a display device (e.g., a light emitting diode (LED)), an organic light emitting diode (OLED), a liquid crystal display, a cathode ray tube display (CRT), a touch screen, a tactile output device, a printer, and / or a speaker). Thus, the interface circuitry 920 of the illustrated example typically includes a graphics driver card, a graphics driver chip, or a graphics driver processor. In some examples, the output device(s) 924 implements the example display 120A.
[0069] The interface circuitry 920 of the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, and / or network interface cards to facilitate data exchange with external machines (e.g., any type of computing device) over a network 926 such as the example network 110 (e.g., an Ethernet connection, a digital subscriber line (DSL), a telephone line, a coaxial cable, a cellular telephone system, etc.).
[0070] The processor platform 900 of the illustrated example also includes one or more mass storage devices 928 for storing software and / or data. Examples of such mass storage devices 928 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray® disk drives, redundant array of independent disks (RAID) systems, and digital versatile disk (DVD) drives.
[0071] Encoded instructions 932 corresponding to the instructions of Figures 5, 6, 7 and / or 8 may be stored on mass storage device 928, volatile memory 914, non-volatile memory 916, local memory 913, and / or a removable tangible computer-readable storage medium such as a CD or DVD 936.
[0072] Figure 10 is a block diagram of an example processor platform 1000 configured to execute the instructions of Figures 5, 6, 7, and / or 8 to implement the example database 210 of Figures 2-4. Processor platform 1000 may be, for example, a server, a personal computer, a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad), a PDA, an Internet appliance, or any other type of computing device.
[0073] The processor platform 1000 of the illustrated example includes a processor 1012. The processor 1012 of the illustrated example is hardware. For example, the processor 1012 can be implemented by one or more integrated circuits, logic circuits, microprocessors, or controllers of any desired family or manufacturer. The hardware processor 1012 can be a semiconductor-based (e.g., silicon-based) device.
[0074] The processor 1012 of the illustrated example includes a local memory 1013 (e.g., a cache). The processor 1012 of the illustrated example communicates with a main memory, including a volatile memory 1014 and a non-volatile memory 1016, via a link 1018. The link 1018 may be implemented by a bus, one or more point-to-point connections, etc., or a combination thereof. The volatile memory 1014 may be implemented by SDRAM, DRAM, RDRAM, and / or any other type of random access memory device. The non-volatile memory 1016 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1014, 1016 is controlled by a memory controller.
[0075] The processor platform 1000 of the illustrated example also includes an interface circuit 1020. The interface circuit 1020 may be implemented by any type of interface standard, such as an Ethernet interface, a USB and / or a PCI Express interface.
[0076] In the illustrated example, one or more input devices 1022 are connected to the interface circuitry 1020. The input devices 1022 allow a user to input data and commands into the processor 1012. The input devices may be implemented, for example, by audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, trackpads, trackballs, trackbars (such as Isopoint), voice recognition systems, and / or any other human-machine interface. Many systems, such as the processor platform 1000, may also allow a user to control and provide data to a computer system using physical actions, including, but not limited to, hand and body movements, facial expressions, and face recognition.
[0077] One or more output devices 1024 are also connected to the interface circuitry 1020 of the illustrated example. The output device(s) 1024 may be implemented, for example, by a display device (e.g., an LED, OLED, LCD display, CRT display, touch screen, tactile output device, printer, and / or speaker). Accordingly, the interface circuitry 1020 of the illustrated example typically includes a graphics driver card, a graphics driver chip, or a graphics driver processor.
[0078] The interface circuitry 1020 of the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, and / or network interface cards to facilitate data exchange with external machines (e.g., any type of computing device) over a network 1026 such as the example network 110 (e.g., an Ethernet connection, DSL), a telephone line, a coaxial cable, a cellular telephone system, etc.).
[0079] The processor platform 1000 of the illustrated example also includes one or more mass storage devices 1028 for storing software and / or data. Examples of such mass storage devices 1028 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disc drives, RAID systems, and DVD drives. In some examples, the mass storage device 1028 may implement the example record storage 225. Additionally or alternatively, in some examples, the volatile memory 1014 may implement the example record storage 225.
[0080] Encoded instructions 1032 corresponding to the instructions of Figures 5, 6, 7 and / or 8 may be stored on mass storage device 1028, volatile memory 1014, non-volatile memory 1016, local memory 1013, and / or a removable tangible computer-readable storage medium such as a CD or DVD 1036.
[0081] Figure 11 is a block diagram of an example processor platform 1100 configured to execute the instructions of Figures 5, 6, 7, and / or 8 to implement the example neural network 205 of Figures 2-4. Processor platform 1100 may be, for example, a server, a personal computer, a mobile device (e.g., a mobile phone, a smartphone, a tablet such as an iPad), a PDA, an Internet appliance, or any other type of computing device.
[0082] The illustrated example processor platform 1100 includes a processor 1112. The illustrated example processor 1112 is hardware. For example, the processor 1112 can be implemented by one or more integrated circuits, logic circuits, microprocessors, or controllers of any desired family or manufacturer. The hardware processor 1112 can be a semiconductor-based (e.g., silicon-based) device. In this example, the processor 1112 implements the example inference engine 230, the example training engine 240, the example comparator 245, the example training data searcher 250, and / or the example deployer 305.
[0083] The processor 1112 of the illustrated example includes a local memory 1113 (e.g., a cache). The processor 1112 of the illustrated example communicates with a main memory, including a volatile memory 1114 and a non-volatile memory 1116, via a link 1118. The link 1118 may be implemented by a bus, one or more point-to-point connections, etc., or a combination thereof. The volatile memory 1114 may be implemented by SDRAM, DRAM, RDRAM, and / or any other type of random access memory device. The non-volatile memory 1116 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1114, 1116 is controlled by a memory controller.
[0084] The processor platform 1100 of the illustrated example also includes an interface circuit 1120. The interface circuit 1120 may be implemented by any type of interface standard, such as an Ethernet interface, a USB and / or a PCI Express interface.
[0085] In the depicted example, one or more input devices 1122 are connected to the interface circuit 1120. The input devices 1122 allow a user to input data and commands into the processor 1112. The input devices may be implemented, for example, by audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, trackpads, trackballs, trackbars (such as Isopoint), voice recognition systems, and / or any other human-machine interface. Many systems, such as the processor platform 1100, may also allow a user to control and provide data to a computer system using physical actions, including, but not limited to, hand and body movements, facial expressions, and face recognition.
[0086] One or more output devices 1124 are also connected to the interface circuitry 1120 of the illustrated example. The output device(s) 1124 may be implemented, for example, by a display device (e.g., an LED, OLED, LCD display, CRT display, touch screen, tactile output device, printer, and / or speaker). Accordingly, the interface circuitry 1120 of the illustrated example typically includes a graphics driver card, a graphics driver chip, or a graphics driver processor.
[0087] The interface circuitry 1120 of the illustrated example also includes communication devices such as transmitters, receivers, transceivers, modems, and / or network interface cards to facilitate data exchange with external machines (e.g., any type of computing device) over a network 1126 such as the example network 110 (e.g., an Ethernet connection, DSL), a telephone line, a coaxial cable, a cellular telephone system, etc.).
[0088] The processor platform 1100 of the illustrated example also includes one or more mass storage devices 1128 for storing software and / or data. Examples of such mass storage devices 1128 include floppy disk drives, hard drive disks, compact disk drives, Blu-ray disc drives, RAID systems, and DVD drives.
[0089] Encoded instructions 1132 corresponding to the instructions of Figures 5, 6, 7 and / or 8 may be stored on mass storage device 1128, volatile memory 1114, non-volatile memory 1116, local memory 1113, and / or a removable tangible computer-readable storage medium such as a CD or DVD 1136.
[0090] The above disclosure provides examples of neural network-based video surveillance. The following additional examples are disclosed herein, including subject matter such as a video surveillance system for implementing neural network-based video surveillance, at least one computer-readable storage medium comprising instructions that, when executed by at least one processor, cause the at least one processor to implement neural network-based video surveillance, a means for implementing neural network-based video surveillance, and a video surveillance method for performing neural network-based video surveillance. The disclosed examples can be implemented individually and / or in one or more combinations.
[0091] From the foregoing, it is understood that example methods, apparatus, systems, and articles of manufacture (e.g., physical storage media) for implementing video surveillance using neural networks are disclosed herein. The disclosed examples include a neural network having an inference engine trained to detect or infer events in monitored video feeds. The neural network inference engine is trained using a training set of reference video segments having reference event labels indicating whether the reference video segments represent defined events. The trained inference engine is then deployed to one or more monitoring locations and operates (e.g., in parallel, asynchronously, cooperatively, etc.) to infer whether one or more trained events (e.g., security breach, human presence, package arrival, etc.) are displayed in the corresponding video feeds it monitors. In some examples, trained / retrained inference engine instances can be deployed to unmanned monitoring locations or can be deployed to manned locations to provide additional capacity, thereby allowing the video surveillance system's capacity to easily scale as demand increases.
[0092] The above disclosure provides examples of implementing video surveillance using neural networks. Further examples of implementing video surveillance using neural networks are disclosed below. The disclosed examples can be implemented individually and / or in one or more combinations.
[0093] Example 1 is a video surveillance system including a database for storing operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments. The system of Example 1 also includes a neural network including a first instance of an inference engine and a training engine for training the first instance of the inference engine based on a training set of operator-labeled video segments retrieved from the database, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set. The system of Example 1 further includes a second instance of the inference engine for inferring events from the monitored video feed, the second instance of the inference engine being based on the first instance of the inference engine.
[0094] Example 2 includes the subject matter of Example 1, where the reference event label indicates whether the corresponding reference video segment depicts the defined event.
[0095] Example 3 includes the subject matter of Examples 1 and / or 2, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event.
[0096] Example 4 includes the subject matter of one or more of Examples 1-3, and further includes a monitoring station, the monitoring station including a display for presenting a first monitored video feed of the monitored video feeds, and a monitoring interface for generating an operator event label based on an operator determination corresponding to a monitored video segment of the first monitored video feed of the monitored video feeds.
[0097] Example 5 includes the subject matter of Example 4, wherein the database is in communication with the monitoring station to receive a first monitored video segment of the monitored video segments and a first operator event label of the operator event labels corresponding to the first monitored video segment of the monitored video segments, and the database stores the first monitored video segment of the monitored video segments and the corresponding first operator event label of the operator event labels as a first reference video segment of the reference video segments and a corresponding first reference event label of the reference event labels included in the first operator-labeled video segment of the operator-labeled video segments.
[0098] Example 6 includes the subject matter of example 5, wherein the monitoring station further implements a second instance of the inference engine.
[0099] Example 7 includes the subject matter of Example 6, wherein the second instance of the inference engine outputs an inferred event for a second one of the monitored video segments of the first one of the monitored video feeds, the monitoring interface generates a second one of the operator event labels from a detected operator judgment for the second one of the monitored video segments, and the monitoring station further includes a comparator that compares the inferred event with the second one of the operator event labels to obtain updated training data.
[0100] Example 8 includes the subject matter of Example 7, wherein the neural network is in communication with the monitoring station to receive the updated training data, and the training engine retrains the first instance of the inference engine based on the updated training data.
[0101] Example 9 includes at least one non-transitory computer-readable storage medium including computer-readable instructions that cause one or more processors to at least: train a first instance of an inference engine based on a training set of operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; and deploy a second instance of the inference engine for inferring events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine.
[0102] Example 9 includes the subject matter of Example 10, where the reference event label indicates whether the corresponding reference video segment depicts the defined event.
[0103] Example 11 includes the subject matter of Examples 9 and / or 10, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event.
[0104] Example 12 includes the subject matter of one or more of Examples 9 to 11, wherein the computer-readable instructions, when executed, cause the one or more processors to obtain a first reference video segment among the reference video segments and a corresponding first reference event label among the reference event labels from the monitoring station.
[0105] Example 13 includes the subject matter of Example 12, wherein the computer-readable instructions, when executed, cause the one or more processors to deploy the second instance of the inference engine to the monitoring station.
[0106] Example 14 includes the subject matter of Example 13, wherein the second instance of the inference engine is a duplicate of the first instance of the inference engine when the second instance of the inference engine is initially deployed to a monitoring station.
[0107] Example 15 includes the subject matter of Example 13, wherein the monitoring station obtains updated training data by comparing (i) inferred events output by the second instance of the inference engine for segments of the monitored video feed and (ii) operator event labels generated by the monitoring station for the segments of the monitored video feed, and the computer-readable instructions, when executed, further cause the one or more processors to retrain the first instance of the inference engine based on the updated training data.
[0108] Example 16 is an apparatus including means for obtaining operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments. The apparatus of Example 16 also includes means for training a first instance of an inference engine based on a training set of the operator-labeled video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set. The apparatus of Example 16 further includes means for deploying a second instance of the inference engine that infers events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine.
[0109] Example 17 includes the subject matter of Example 16, where the reference event label indicates whether the corresponding reference video segment depicts the defined event.
[0110] Example 18 includes the subject matter of Examples 16 and / or 17, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event.
[0111] Example 19 includes the subject matter of one or more of Examples 16-18, wherein the means for obtaining the record of the operator-labeled video segment is for obtaining a first reference video segment from the reference video segments and a corresponding first reference event label from the reference event labels from the monitoring station.
[0112] Example 20 includes the subject matter of Example 19, wherein the means for deploying the second instance of the inference engine is for deploying the second instance of the inference engine to the monitoring station.
[0113] Example 21 includes the subject matter of Example 20, wherein the second instance of the inference engine is a duplicate of the first instance of the inference engine when the second instance of the inference engine is initially deployed to a monitoring station.
[0114] Example 22 includes the subject matter of Example 20, wherein the monitoring station is for obtaining updated training data by comparing (i) inferred events output by the second instance of the inference engine for segments of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segments of the monitored video feed, and the means for training the first instance of the inference engine further is for retraining the first instance of the inference engine based on the updated training data.
[0115] Example 23 is a method of video surveillance comprising: executing instructions with at least one processor to train a first instance of an inference engine based on a training set of operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set. The video surveillance method of Example 23 also includes executing instructions with at least one processor to deploy a second instance of the inference engine to infer events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine.
[0116] Example 24 includes the subject matter of Example 23, where the reference event label indicates whether the corresponding reference video segment depicts the defined event.
[0117] Example 25 includes the subject matter of Examples 23 and / or 24, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event.
[0118] Example 26 includes the subject matter of one or more of Examples 23 to 25, and wherein accessing the record of the operator-labeled video segment includes obtaining a first reference video segment from the reference video segments and a corresponding first reference event label from the reference event labels from a monitoring station.
[0119] Example 27 includes the subject matter of Example 26, wherein deploying the second instance of the inference engine includes deploying the second instance of the inference engine to the monitoring station.
[0120] Example 28 includes the subject matter of Example 27, wherein the second instance of the inference engine is a duplicate of the first instance of the inference engine when the second instance of the inference engine is initially deployed to a monitoring station.
[0121] Example 29 includes the subject matter of Example 27, wherein the monitoring station obtains updated training data by comparing (i) inferred events output by the second instance of the inference engine for a segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and wherein training the first instance of the inference engine further includes retraining the first instance of the inference engine based on the updated training data.
[0122] Although certain example methods, apparatus, and articles of manufacture are disclosed herein, the scope of coverage of this patent is not limited thereto. Rather, this patent covers all methods, apparatus, and articles of manufacture fairly falling within the scope of the claims of this patent. Other possible claims (Item 1) 1. A video surveillance system comprising: a database for storing operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments; A neural network, a first instance of an inference engine; a training engine for training the first instance of the inference engine based on a training set of the operator-labeled video segments retrieved from the database, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; and a neural network including a second instance of the inference engine that infers events from the monitored video feed, the second instance of the inference engine being based on the first instance of the inference engine; 1. A video surveillance system comprising: (Item 2) 2. The video surveillance system of item 1, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 3) A video surveillance system as described in item 1, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event. (Item 4) The method further comprises a monitoring station, the monitoring station comprising: a display for presenting a first one of the monitored video feeds; a monitoring interface that generates an operator event label based on an operator determination corresponding to a monitored video segment of the first monitored video feed of the monitored video feeds; Item 1. The video surveillance system of item 1, comprising: (Item 5) 5. The video surveillance system of claim 4, wherein the database communicates with the monitoring station to receive a first monitored video segment from among the monitored video segments and a first operator event label from among the operator event labels corresponding to the first monitored video segment from among the monitored video segments, and the database stores the first monitored video segment from among the monitored video segments and the corresponding first operator event label from among the operator event labels as a first reference video segment from among the reference video segments included in the first operator-labeled video segment from among the operator-labeled video segments and a corresponding first reference event label from among the reference event labels. (Item 6) Item 6. The video surveillance system of item 5, wherein the monitoring station further implements the second instance of the inference engine. (Item 7) 7. The video surveillance system of claim 6, wherein the second instance of the inference engine outputs an inferred event for a second one of the monitored video segments of the first one of the monitored video feeds, the monitoring interface generates a second one of the operator event labels from an operator judgment detected for the second one of the monitored video segments, and the monitoring station further includes a comparator that compares the inferred event with the second one of the operator event labels to obtain updated training data. (Item 8) 8. The video surveillance system of claim 7, wherein the neural network communicates with the monitoring station to receive the updated training data, and the training engine retrains the first instance of the inference engine based on the updated training data. (Item 9) At least one non-transitory computer-readable storage medium containing computer-readable instructions, the computer-readable instructions causing one or more processors to at least: training a first instance of an inference engine based on a training set of operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; deploying a second instance of the inference engine for inferring events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine; At least one non-transitory computer-readable storage medium that executes the method. (Item 10) Item 10. At least one storage medium according to item 9, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 11) At least one storage medium as described in Item 9, wherein a first reference event label among reference event labels corresponding to a first reference video segment among reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among reference video segments represents the above type of event. (Item 12) At least one storage medium as described in item 9, wherein the computer-readable instructions, when executed, cause the one or more processors to obtain a first reference video segment among the reference video segments and a corresponding first reference event label among the reference event labels from a monitoring station. (Item 13) Item 13. At least one storage medium according to item 12, wherein the computer-readable instructions, when executed, cause the one or more processors to deploy the second instance of the inference engine to the monitoring station. (Item 14) Item 14. At least one storage medium according to item 13, wherein when the second instance of the inference engine is initially deployed in a monitoring station, the second instance of the inference engine is a duplicate of the first instance of the inference engine. (Item 15) Item 14. At least one storage medium described in Item 13, wherein the monitoring station obtains updated training data by comparing (i) inferred events output by the second instance of the inference engine for the segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and the computer-readable instructions, when executed, further cause the one or more processors to retrain the first instance of the inference engine based on the updated training data. (Item 16) 1. An apparatus comprising: means for obtaining an operator-labeled video segment, the operator-labeled video segment including a reference video segment and a corresponding reference event label that describes the reference video segment; means for training a first instance of an inference engine based on a training set of the operator-labeled video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; means for deploying a second instance of the inference engine that infers events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine; An apparatus comprising: (Item 17) Item 17. The apparatus of item 16, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 18) The apparatus of item 16, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event. (Item 19) Item 17. The apparatus of item 16, wherein the means for obtaining the operator-labeled video segment is for obtaining a record of a first reference video segment from the reference video segments and a corresponding first reference event label from the reference event labels from a monitoring station. (Item 20) 20. The apparatus of claim 19, wherein the means for deploying the second instance of the inference engine is for deploying the second instance of the inference engine to the monitoring station. (Item 21) Item 21. The apparatus of item 20, wherein the second instance of the inference engine is a duplicate of the first instance of the inference engine when the second instance of the inference engine is initially deployed in the monitoring station. (Item 22) 21. The apparatus of claim 20, wherein the monitoring station is for obtaining updated training data by comparing (i) inferred events output by the second instance of the inference engine for a segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and wherein the means for training the first instance of the inference engine further is for retraining the first instance of the inference engine based on the updated training data. (Item 23) 1. A video surveillance method comprising: training a first instance of an inference engine based on a training set of operator-labeled video segments by executing instructions with at least one processor, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; deploying a second instance of the inference engine for inferring events from monitored video feeds by executing instructions with the at least one processor, the second instance of the inference engine being based on the first instance of the inference engine; A method for providing (Item 24) Item 24. The method of item 23, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 25) The method described in item 23, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments expresses the type of event. (Item 26) Item 24. The method of item 23, wherein accessing the record of the operator-labeled video segment includes obtaining a first reference video segment from the reference video segments and a corresponding first reference event label from the reference event labels from a monitoring station. (Item 27) 27. The method of claim 26, wherein deploying the second instance of the inference engine includes deploying the second instance of the inference engine to the monitoring station. (Item 28) Item 28. The method of item 27, wherein the second instance of the inference engine is a duplicate of the first instance of the inference engine when the second instance of the inference engine is initially deployed to the monitoring station. (Item 29) 28. The method of claim 27, wherein the monitoring station obtains updated training data by comparing (i) inferred events output by the second instance of the inference engine for the segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and wherein training the first instance of the inference engine further includes retraining the first instance of the inference engine based on the updated training data. (Item 2-1) 1. A video surveillance system comprising: a database for storing operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments; A neural network, a first instance of an inference engine; a training engine for training the first instance of the inference engine based on a training set of the operator-labeled video segments retrieved from the database, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; and a neural network including a second instance of the inference engine that infers events from the monitored video feed, the second instance of the inference engine being based on the first instance of the inference engine; 1. A video surveillance system comprising: (Item 2-2) The video surveillance system described in item 2-1, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 2-3) A video surveillance system as described in item 2-1, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event. (Item 2-4) The method further comprises a monitoring station, the monitoring station comprising: a display for presenting a first one of the monitored video feeds; a monitoring interface that generates an operator event label based on an operator determination corresponding to a monitored video segment of the first monitored video feed of the monitored video feeds; A video surveillance system according to any one of items 2-1 to 2-3, comprising: (Item 2-5) The video surveillance system described in items 2-4, wherein the database communicates with the monitoring station to receive a first monitored video segment among the monitored video segments and a first operator event label among the operator event labels corresponding to the first monitored video segment among the monitored video segments, and the database stores the first monitored video segment among the monitored video segments and the corresponding first operator event label among the operator event labels as a first reference video segment among the reference video segments included in the first operator-labeled video segment among the operator-labeled video segments and a corresponding first reference event label among the reference event labels. (Item 2-6) The video surveillance system of items 2-5, wherein the monitoring station further implements the second instance of the inference engine. (Item 2-7) The video surveillance system described in items 2-6, wherein the second instance of the inference engine outputs an inferred event for a second monitored video segment of the monitored video segments of the first monitored video feed of the monitored video feed, the monitoring interface generates a second operator event label from the operator event labels from a detected operator judgment for the second monitored video segment of the monitored video segments, and the monitoring station further includes a comparator that compares the inferred event with the second operator event label from the operator event labels to obtain updated training data. (Item 2-8) A video surveillance system as described in items 2-7, wherein the neural network communicates with the monitoring station to receive the updated training data, and the training engine retrains the first instance of the inference engine based on the updated training data. (Item 2-9) 1. An apparatus for performing video surveillance, said apparatus comprising: means for obtaining operator-labeled video segments, the operator-labeled video segments including a reference video segment and a corresponding reference event label that describes the reference video segment; means for training a first instance of an inference engine based on a training set of the operator-labeled video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; means for deploying a second instance of the inference engine that infers events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine; An apparatus comprising: (Item 2-10) The apparatus described in items 2-9, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 2-11) The device described in Items 2-9, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event. (Item 2-12) An apparatus described in any one of items 2-9 to 2-11, wherein the means for obtaining the operator-labeled video segment is for obtaining a first reference video segment among the reference video segments and a corresponding first reference event label among the reference event labels from a monitoring station. (Item 2-13) The apparatus of items 2-12, wherein the means for deploying the second instance of the inference engine is for deploying the second instance of the inference engine to the monitoring station. (Item 2-14) The apparatus of items 2-13, wherein when the second instance of the inference engine is initially deployed to the monitoring station, the second instance of the inference engine is a duplicate of the first instance of the inference engine. (Item 2-15) The apparatus described in items 2-13, wherein the monitoring station is for obtaining updated training data by comparing (i) inferred events output by the second instance of the inference engine for a segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and the means for training the first instance of the inference engine further includes retraining the first instance of the inference engine based on the updated training data. (Item 2-16) 1. A video surveillance method comprising: training, with at least one processor, a first instance of an inference engine based on a training set of operator-labeled video segments, the operator-labeled video segments including reference video segments and corresponding reference event labels that describe the reference video segments, the first instance of the inference engine inferring events from the operator-labeled video segments included in the training set; deploying, with the at least one processor, a second instance of the inference engine that infers events from monitored video feeds, the second instance of the inference engine being based on the first instance of the inference engine; A method for providing (Item 2-17) The method of claim 2-16, wherein the reference event label indicates whether the corresponding reference video segment represents a defined event. (Item 2-18) The method described in item 2-16, wherein a first reference event label among the reference event labels corresponding to a first reference video segment among the reference video segments indicates (i) a type of event and (ii) whether the first reference video segment among the reference video segments represents the type of event. (Item 2-19) A method described in any one of items 2-16 to 2-18, further comprising a step of obtaining a first reference video segment among the reference video segments and a corresponding first reference event label among the reference event labels from a monitoring station. (Item 2-20) The method of claim 2-19, wherein the step of deploying the second instance of the inference engine includes the step of deploying the second instance of the inference engine to the monitoring station. (Item 2-21) The method of item 2-20, wherein when the second instance of the inference engine is initially deployed to the monitoring station, the second instance of the inference engine is a replica of the first instance of the inference engine. (Item 2-22) The method described in items 2-20, wherein the monitoring station obtains updated training data by comparing (i) inferred events output by the second instance of the inference engine for a segment of the monitored video feed with (ii) operator event labels generated by the monitoring station for the segment of the monitored video feed, and the step of training the first instance of the inference engine further includes a step of retraining the first instance of the inference engine based on the updated training data. (Item 2-23) A program that causes a computer to execute the method according to any one of items 2-16 to 2-22.
Claims
1. an interface circuit; computer-readable instructions; at least one processor circuit; the at least one processor circuit is programmed with the computer-readable instructions; deploying a first machine learning model on a first device, the first machine learning model being trained using first training data to detect at least one event depicted in a first video segment; after at least a threshold amount of second training data becomes available, retraining the first machine learning model based on the second training data to obtain a second machine learning model, the second training data being based on output data from the first device, the output data being associated with execution of the first machine learning model by the first device; The apparatus deploys the second machine learning model to a second device different from the first device.
2. The apparatus of claim 1 , wherein the output data is from a plurality of devices that each executed the first machine learning model, the plurality of devices including the first device.
3. 10. The apparatus of claim 1, wherein one or more of the at least one processor circuit retrains the first machine learning model (i) after the at least threshold amount of second training data becomes available and (ii) after periodic intervals have elapsed.
4. 10. The apparatus of claim 1, wherein one or more of the at least one processor circuit retrains the first machine learning model (i) after the at least threshold amount of second training data becomes available, and (ii) after periodic intervals, and (iii) after receiving user input.
5. 5. The apparatus of claim 1, wherein one or more of the at least one processor circuit trains the first machine learning model to output values representing respective likelihoods that corresponding events are represented in the input video segment.
6. The apparatus of claim 1 , wherein the at least one event corresponds to the arrival of a package.
7. The device of claim 1 , wherein the at least one event corresponds to the presence of a human being.
8. an interface circuit for downloading the machine learning model; computer-readable instructions; at least one processor circuit; the at least one processor circuit is programmed with the computer-readable instructions; executing the machine learning model to generate first output data corresponding to a first input video segment, the machine learning model being trained using first training data to detect at least one event depicted in the training video segment; causing the interface circuit to report second training data, the second training data based on the first input video segment, the first output data, and a label specifying whether the first input video segment represents the at least one event; and executing a retrained instance of the machine learning model to generate second output data corresponding to a second input video segment, the retrained instance of the machine learning model being based on the second training data.
9. 9. The device of claim 8, wherein one or more of the at least one processor circuit segments a video feed based on characteristics of an image sensor to obtain the first input video segment and the second input video segment, and the image sensor captures the video feed.
10. The apparatus of claim 9 , wherein the characteristics include a sweep rate of the image sensor.
11. The apparatus of claim 9 , wherein the characteristics include an imaging speed of the image sensor.
12. 12. The apparatus of claim 8, wherein the label specifying whether the first input video segment depicts the at least one event is based on user input.
13. One or more of the at least one processor circuit: causing the first input video segment to be displayed on a display; and The apparatus of claim 12 , further comprising: generating the label based on the user input.
14. 14. The apparatus of claim 8, wherein the first output data includes an indication of whether the machine learning model has inferred that the first input video segment represents the at least one event, and wherein one or more of the at least one processor circuit determines the second training data based on a comparison of the indication (i) of whether the machine learning model has inferred that the first input video segment represents the at least one event with the label (ii) indicating whether the first input video segment represents the at least one event.
15. At least one processor circuit includes at least executing the downloaded machine learning model to generate first output data corresponding to a first input video segment, the machine learning model having been trained using first training data to detect at least one event depicted in the training video segment; reporting second training data, the second training data being based on the first input video segment, the first output data, and a label specifying whether the first input video segment represents the at least one event; and and executing a retrained instance of the machine learning model to generate second output data corresponding to a second input video segment, wherein the retrained instance of the machine learning model is based on the second training data.
16. 16. The computer program product of claim 15, further comprising causing one or more of the at least one processor circuit to perform a procedure for segmenting a video feed based on characteristics of an image sensor to obtain the first input video segment and the second input video segment, the image sensor capturing the video feed.
17. The computer program product of claim 16 , wherein the characteristics include a sweep rate of the image sensor.
18. The computer program product of claim 16 , wherein the characteristics include an imaging speed of the image sensor.
19. one or more of the at least one processor circuit; displaying the first input video segment on a display; 19. A computer program product according to claim 15, further comprising: a step of generating the label based on a user input.
20. 20. The computer program product of claim 15, wherein the first output data includes an indication of whether the machine learning model has inferred that the first input video segment represents the at least one event, and wherein the computer program product causes one or more of the at least one processor circuit to perform a procedure of determining the second training data based on a comparison of the indication (i) of whether the machine learning model has inferred that the first input video segment represents the at least one event with the label (ii) indicating whether the first input video segment represents the at least one event.
21. At least one non-transitory computer readable medium having stored thereon a computer program according to any one of claims 15 to 20.
Citation Information
Patent Citations
Program for correcting morpheme analysis result
JP2006252091A
Information transfer device, leaning system, information transfer method, and program
JP2016191973A
Image managing device, image managing method and program
JP2017028688A
Video surveillance system, video processing apparatus, video processing method, and video processing program
JP2017225122A
Recording / analyzing system for accidental event
WO2005101346A1