System and method for dynamic sampling of video data
The system dynamically adjusts video sampling rates based on event types using AI algorithms to enhance the efficiency and clarity of video playback, addressing the inefficiencies of constant frame rate capture in surveillance systems.
Patent Information
- Application Number
- PCT/US2025/011525
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2025-01-14
- Publication Date
- 2025-07-31
AI Technical Summary
Existing video surveillance systems capture video at a constant frame rate, leading to inefficiencies when interesting events occur sporadically over long periods, as viewers may miss important events due to high-speed playback or fatigue, and current solutions like increasing playback speed can result in missed events.
A system and method for dynamically adjusting video sampling rates based on event types, using AI algorithms to detect and classify events, and varying sampling rates accordingly to focus on interesting portions, allowing for efficient review of relevant content.
This approach enables targeted playback of important events by dynamically adjusting sampling rates, reducing the time spent reviewing uninteresting footage and enhancing the clarity of interesting segments, thus improving the efficiency and effectiveness of video analysis.
Smart Images

Figure US2025011525_31072025_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR DYNAMIC SAMPLING OF VIDEO DATACROSS-REFERENCE TO RELATED APPLICATION[0001| This application claims the benefit of IIS Provisional Application No. 63 / 623,371 filed January 22, 2024, which is hereby incorporated by reference to the extent not inconsistent.BACKGROUND
[0002] Video flies generally include multiple individual still images (‘‘frames’’) that when displayed in a stream create the appearance of a moving image. In common implementations, the time interval between consecutive frames is constant so that displayed images will appear to move smoothly when the video is played. Camera devices are commonly placed in a fixed location to catch different types of events, such as in the ease of a surveillance camera. For example, a ‘'movement event” may be captured that involves any motion detected in the camera’s field of view. This may occur because of an object moving through the field of view, or possibly when the camera itself is moved. In other examples, the camera may capture a person moving through its field of view (a “person event”), a vehicle passing through the camera's field of view (a “vehicle event”), and / or a pet passing into the field of view (a “pet event”). These kinds of events, as well as others, may happen at any time during a long period of recorded video. The events may also happen in a short time frame (in seconds) within a long period (of hours or days).
[0003] Video is usually captured at a constant frame sampling rate. For example, at 20 frames-per-secotid (FPS), the time between each two consecutive frame is 1 / 20 of a second. If events only occur sporadically during the video, very few “interesting” events (lasting maybe only a few seconds or minutes) may happen over a period of hours or days. This makes it inefficient to view or search the video since much of what the camera has captured is not interesting (e.g. does not include any noteworthy events). One common approach addressing this issue is to increase the playback speed of the video by a factor of 2, 15, or 50, or more. This can reduce the total time needed io replay the video, but it may mean missing an important event because the event may occur so quickly that a person viewing the video at high speed may miss it, particularly in instances where the viewer has been watching many hours of such video and f atigue becomes a factor.SUMMARY
[0004] Disclosed is a system and method for generating and / or displaying a video using dynamic video sampling rates based on the content of the video.[0005| In one aspect, the disclosed method includes automatically adjusting the sampling rate according to specific criteria. The criteria may include, for example, different types of events such as movement events, a person event, a vehicle event, and the like. In one aspect, the method optionally includes obtaining video data comprising multiple individual image frames and detecting events captured in the video data. For example, video captured by a camera may be recorded videos in advance in volatile or nonvolatile memory. Examples include an SD card, a solid-state hard drive, a mechanical hard drive, a Network Attached Storage (NAS) device, or a remote cloud storage service.
[0006] In another aspect, the disclosed method includes locating and or extracting or exporting portions of video data associated with the different events. For example, a movement event may be detected when an object passes through the field of view of the camera. In another aspect, a movement event may occur when the camera moves, hi another aspect, the disclosed method may include classifying the event according to predetermined criteria. For example, a person event may have happened at 3:30:35pm and ended at 3:31 : 10pm when a person walked through the field of view of the recording camera. A vehicle event may have happened between 4:30:05pm and 4:30:28pm when a vehicle passed through the viewing area of the camera,
[0007] In another aspect, Artificial Intelligence (Al) algorithms may he included, in the system of the present disclosure and may be configured to determine when a person, vehicle, or other type of object has passed, or is passing, through the field of view defined by a camera. These algorithms may, for example, include a neural network, such as a Convolutional Neural Network (CNN), a transformer model, or other similar decisionmaking algorithm.
[0008] In another aspect, the disclosed method may include accepting input specifying examples of existing types of objects. In another aspect the disclosed system may be configured to accept input defining new types of objects. This input may, for example, include images of the different types of objects. Training sets including images and / or videos for different classes or types of objects may be accepted as input to the control logic as part of a process of training the control logic to detect similar or related objects with related characteristics that are optionally classified by type (e.g. intrusive deer as opposed to animalsgenerally; cars rather than vehicles generally which could also inchide vans, motorcycles, or delivery trucks; or particular “unfamiliar” people as opposed to “familiar" people, or just people in general).
[0009] In another aspect, the method optionally includes sampling frames from the different portions of video data associated with different events at differing sampling rates that may vary according to the type of event detected in each portion. The sampling rates are optionally predetermined according to configuration data in the control logic which may be modified according to input accepted from a user, or optionally according to training input and: or updates to the A.l model.
[0010] In another aspect, the method optionally includes displaying the sampled portions of the video data at a predetermined frame rate. For example, where a default sampling rate may be only one frame of every 100 frames, or one frame of every 50 frames. In another example, for a “person” event, the system may be configured to sample every other frame, or to include all frames to thus increase the clarity of the image and provide smoother playback later. For “vehicle”, “car”, or other events where action happens quickly, the system may sample every frame to maintain clarity.
[0011] In another aspect, the system of the present disclosure may provide for viewing the sampled (and / or the unsampled “raw”) video. In one aspect, the method of the present disclosure includes creating a playback video file, or video stream, that optionally includes the sampled video. The sampled video may be replayed at a predetermined framerate of 20, 30, or 60 frames per second (or more). The result is a video playback experience that reduces time spent watching video that is of no particular interest. For example, surveillance footage of unauthorized persons entering a restricted area for 60 seconds of a 24 hour period can be quickly located and reviewed without a user being required to preview all 24 hours of video. Thus the system of the present disclosure effectively focuses the resulting video content based on predetermined event criteria that is optionally tailored to a viewer’s interest. When a portion of the recorded raw video includes interesting e vents, more of those frames are included in the sampled video output thus capturing more important details and ignoring uninteresting periods of video that is possibly useless to the viewer.
[0012] The disclosed method may be implemented in control logic in software that is executed using one or more processors of one or more compu ters. In another aspect, the disclosed method may be implemented using control logic in hardware such as programs, Al models, instructions, or logic gates embedded in Application Specific Integrated Circuits(ASICs), in Field Programmable Gate .Arrays (FPGAs), or microcontrollers, Al chips, or other custom built hardware using combinations of logic gates or other electronic components. The disclosed method may be implemented according to control logic in any suitable configuration of hardware and / or software. The control logic may be included in a camera, in a local computing device such as a local server, or in a remote computing device such as in the case of a cloud computing platform.[0013| Further forms, objects, features. aspects, benefits, advantages, and embodiments of the present invention will become apparent from a detailed description and drawings provided herewith.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] FIG. 1 is a component diagram illustrating components that may be included in a system of the present disclosure.
[0015] FIG. 2 is a flowchart illustrating aspects of the disclosed method for dynamic sampling of video data.
[0016] FIG. 3 is a flowchart illustrating one example of actions the disclosed system may take in generating dynamically sampled video data,
[0017] FIG, 4 is a diagram illustrating multiple examples of events found in video data according to the present disclosure.
[0018] FIG. 5 is a flowchart illustrating another example of actions the disclosed system may take in generating dynamically sampled video data.DET AIL ED DESCRI PTION
[0019] Illustrated in Fig. I at 100 is one example of components that may be included in a system of the present disclosure. The system may include a camera 107 (or multiple cameras) operable to capture images within a field-of-view or viewing area 106 defined by the camera. In another aspect, the field of view may be thought of as the detection zone, or detection region for the camera 107.
[0020] The camera 107 may include control logic 105 which may implement some or all of the disclosed control logic. The control logic may be useful for operation of the camera to capture and sample video imagery according to the present disclosure. Control logic 105 may include, or be electrically connected to, a processor, memory, or other logic circuitry, and / or sensors such as Charged Couple Devices (CCDs) or Complimentary Metal Oxide Semiconductor (CMOS) image sensors, which may be used to capture light visible or invisible to the human eye.
[0021] Control logic 105 may also include configuration settings and logical operations which may be used by the control logic to determine what type of object is within the field- of-view 106. Some exemplary types of objects that may be detected and categorized by the control logic of the present disclosure are illustrated in Fig. 1. These include, but are not limited to, vehicles such as a car 1 10, or a truck 114, people 11 I , animals 112, or other objects 1 13.
[0022] The control logic may include one or more rules with criteria or thresholds for determining that objects are present, or that activities are occurring, within the field of view of the camera (and by extension, in the video data the camera captures) that may be of particular interest. For example, when the control logic is applied to image data collected by camera 107, the algorithms of the present disclosure may indicate movement has occurred within the cameras field-of-view 106. In another aspect, the control logic may be operable to determ ine a type of object that is present in the field of view 106. This may be useful for determining that a delivery truck arrived, that a person is at the door, and whether the person or truck is familiar or unfamiliar,
[0023] In another aspect, Artificial Intelligence (Al) algorithms or decision-making models may be included in the control logic 105 of the present disclosure and may be configured to determine that an object passed through the field-of-view 106 of the camera, and possibly the type of object that it is. These Al models or algorithms may be operable to classify the type of object automatically preferably without human intervention. The Alalgorithms may, for example, include one or more neural networks, such as a Convolutional Neural Network (CNN), or other similar decision making algorithm. Transformers optimized for decision making using captured video feeds as input may also be implemented in control logic 105.
[0024] In another aspect the control logic may be configured to maintain a collection of familiar objects, unfamiliar objects, or any combination thereof. This data may be updated automatically over time as the decision-making paradigm of the control logic 105, and any optional Al algorithms that may be included therein, is improved with additional input and training.
[0025] In one example, an algorithm configured to make comparisons using rules, thresholds, or other criteria maintained by the control logic 105, may be used to determine when image data retrieved by the camera changes over time. These changes may be analyzed by the algorithm in real time, or later after the video has been saved and retrieved for analysis, to determine when changes occurring within the field-of-view 106 match the criteria in the algorithm which are configured to indicate that movement has occurred. Predetermined values may be assigned on a gradient with predefined maximum and minimum values, In one example, no change may be defined as a value of 0.0, while a high degree of change may be assigned a maximum value of 1 .0. Threshold values may then be assigned on a gradient between these extremes to specify when enough change has occurred to indicate movement.
[0026] In another aspect, the camera 107 may capture multiple individual frames over time such as at the rate of 30 frames per second, or 60 frames per second, or any other suitable rate. The control logic 105 of the camera may be configured to present the individual frames captured by (he camera as input to the algorithm to determine if movement has occurred, and optionally to determine a category for the events taking place, or for the objects involved in those events. The algorithm may, for example, compare the pixels at or near corresponding positions of the individual frames for one or more successive frames to determine if the pixel data has changed according to the criteria in the rules. These individual pixel “deltas” may be grouped together, filtered, and / or compared over time to determine whether the changes in the pixel data are sufficient to trigger an alert that movement has been detected, and they may also be used to determine a category for the activity that has occurred.In another aspect, the changes in pixel data may be compared and may be the basis for determining changes in lighting, shadows, positions of objects, types of objects, familiarity, and the like. The results of these comparisons performed according to the disclosed methodmay be further input for the control logic as it determines whether shapes and configurations of objects in relation to one another indicate an event has taken place, and / or whether this event includes familiar faces, objects, vehicles, and the like. The overall result of these comparisons may be assigned values according to the predetermined gradients between a maximum and minimum value. The control logic may be configured to trigger a change to a different sampling rate if one or more of the compari son outcomes are outside of predefined triggering ranges or thresholds specified in the control logic.
[0027] For example, a camera positioned at a garage door may be mounted at a high angle and positioned with a field-of-view of the parking area adjacent the garage door. The parking area tha t is within the field-of-view of the camera may change very little over time except when a vehicle passes, or stops within the cameras field-of-view, As the vehicle enters the field-of-view, the pixel data within the frames captured by the camera shifts from the nominal black or gray colors of the parking lot to accommodate the color and shape of the vehicle. A small change of a few pixels may be insufficient to trigger a sampling rate adjustment, but a growing number of pixels changing rapidly within a few frames, or an overall increase in the number of pixels that have changed may be sufficient to trigger a movement event, or to trigger an event indicating that an object is now in viewy or other event. The changing pixel data may also be used to determ ine what type of object is moving into the field of view according to the shape, color, markings, lights, type of activity', or other characteristics perceived by the control logic.
[0028] The control logic may include criteria for determining not only that motion is occurring in the video data, but that a particular type of activity is taking place, or that a particular type of object is within the field-of-view of the camera. Control logic 105 may be configured with pattern recognition algorithms operable to detect packages, faces, animals, vehicles, and the like. These pattern recognition algorithms may include multiple predetermined matching criteria specific to individual portions of the image, or to the image overall, or to specific configurations of shapes or arrangements of shadows, or other image features. Any suitable method for determining the type of objects w ithin the field-of-view of the camera may be useful by the control logic criteria for categorizing objects appearing, disappearing, or otherwise moving within the field-of-view 106.
[0029] In another aspect, pattern recognition algorithms may include criteria specific to determine whether an object appearing in the image is familiar or unfamiliar. The criteria may be trained to automatically detect familiar persons, vehicles, animals, and the like basedon past experience and / or input from a user. For example, the system may determine that a vehicle is a delivery truck from a well-known delivery service. In another example, the system may determine at first that a delivery vehicle is present, but that it is unfamiliar to the system. This scenario might occur when a new delivery service begins operations in the area. Over time, with repeated visits, and repeated movements by the driver carrying a package to the door, or other common movement, the system may automatically adjust to update this delivery service vehicle as “familiar” where before it was classified as unfamiliar.
[0030] The process of automatically adjusting an object or activity from unfamiliar to familiar may be implemented by the system of the present disclosure for objects of any type including people, animals, and the like. For example, a new friend may be classified by the control logic 105 as an unfamiliar person, and later classified as a. familiar person upon repeated visits. In another aspect, the control logic 105 may be configured to selectively only update a person, vehicle, animal, or other object, as familiar, based on input from a user. In some instances, it may be desirable to require user input classifying an object, as familiar to avoid misclassification.
[0031] In another aspect, user input may be obtained via an input device of computing device 101 . The user may be presented with visual and / or auditory output from camera 107 during playback, and may be given the option to change the classification for a detected object. The user may be given the option to specify whether the category that was determined by the rules in control logic 105 matches the image / sound output of the camera. For example, a rule in control logic 105 may be triggered based on image data indicating that a package has been delivered. Upon inspecting the image, the user 104 may determine that an object that is presently within the field-of-view of camera 106 is not a package but is instead some other type of object, and / or that the object is either familiar or unfamiliar to the user.
[0032] In another aspect, a remote service platform 102 may include one or more computers 103 with one or more processors configured to execute or implement the control logic 105. In this example, control logic 105 may be wholly, or partially, stored in, and maintained in, the service platform 102. The control logic as disclosed herein throughout, may be partially in the camera 107, partially in the remotes service platform 102, or any combination thereof. In one example, the disclosed Al algorithms may be maintained in the service platform 102 and may be configured to train Al models to implement the decisionmaking and classification aspects disclosed herein. These trained models or other aspects of the Al algorithms may be delivered to the camera 107 via a communication link 120.
[0033] In another aspect, the video data including multiple individual image frames may be sent from the camera 107 to the remote service platform 102 for processing according to the present disclosure. The processed video including the disclosed dynamically sampled frame rates is optionally delivered from the service platform 102 to the computing device 101 via a communication link 121.
[0034] Ln another aspect, the camera 107 may include all control logic 105, and thus the dynamic sampling of the present disclosure may be executed by the camera. The dynamically sampled output may be provided to the service platform 102 for storage. This sampled output received from the camera may optionally be delivered to the computing device 101 via the communication link 121 ,
[0035] In another aspect, the camera 107 may include some or all of the control logic 105, and thus the dynamic sampling of the present disclosure may be executed by the camera 107, service platform 102, or any combination thereof. The dynamically sampled output may be provided to the computing device 101 via an optional communication link 122 between the computing device 101 and the camera 107. In this example, the camera 107 may communicate directly with the computing device 101. Some of the input received from (he camera 107 may optionally be passed to the service platform 102 by the computing device 101 if needed.
[0036] The user interface presented by computing device 101 may offer an option to indicate that an event type chosen by the control logic 105 is correct or is incorrect, and / or to adjust oilier configuration data or settings of the control logic. This input accepted by the system of the present disclosure may be obtained by the service platform 102 and passed along to the camera 107, or delivered from (he computing device 101 directly to the camera 107, The service platform 102 may be configured to adjust one or more of the rule criteria threshold values accordingly to better match the video data to the desired event related determinations. These newly calculated values may be sent back to camera 107 and automatically installed in control logic 105 so that future determinations of event type, familiarity vs unfamiliarity, or other automatic determinations the system may make can be more accurate.
[0037] FIG. 2 is a flowchart illustrating aspects of (he disclosed method for dynamic sampling of video data. The system of the present disclosure may be arranged and configured to obtain video data at 202, and to detect events in the video data at 203, At 204, video datamay be sampled according to the detected events at 203, and the resulting video may be displayed at 205.
[0038] Obtaining the video at 202 optionally includes receiving, accepting, or capturing video data comprising multiple individual image frames. The video may be obtained from a fil e of a predetermined fixed length with a fixed number of individual frames saved in the file . In another example, the video may be obtained from a continuous stream of frames passing over a communication link such as tn a streaming video feed obtained via a telecast, and the like. In one aspect, this may include capturing video data using a camera, either before or during the dynamic sampling of the present disclosure. The multiple individual images may have been generated or captured and stored as video data that was captured at a predetermined and / or fixed number of frames captured per unit of time, usually expressed as a number of “frames per second”.
[0039] In another aspect, video may be accepted from another computer system tor processing by the system of the present disclosure. A remote server may send the captured video to a server or other computer of the present system for processing, and the resulting dynamically sampled video may be sent back after processing is complete. In another aspect, the video data may be generated rather than captured by a camera, such as in the case of computer-generated video graphics or other Computer-Generated Imagery (CGI).
[0040] Fig. 3 illustrates at 300 one example of actions the system may take in determining events are present in the video data, sampling the video data accordingly, and displaying the resulting video data (corresponding with 203, 204, and 205, respectively). The disclosed process optionally includes automatically adjusting the sampling rate according to specific criteria such as the event type. The sampling rate tor frames from different portions of the video data may vary according to the type of event detected in each portion. Portions of the video may thus be prepared for display according to the events found in the video data, and this dynamic sampling may occur in real time as the video data is streamed, or later by accessing a recording after the events have occurred.
[0041] At 302, the system of the present disclosure is arranged and configured to read video data, such as might be obtained at 202. If an event is found at 303, the type of event is determined according to the present disclosure at 304. In another aspect, the control logic may include object detection capabilities that may be implemented, for example, in an object detection module implemented in hardware and / or software. Object detection of the present disclosure optionally includes determining if a detected object is a person, pet, vehicle, orother category of object. This optionally includes accessing a database of familiar shapes and / or activities to determine if an event includes a familiar object. The database of information about familiar shapes may also include defining aspects of portions of shapes which may be used by the control logic to develop an overall assessment of what type of object is present. In another aspect, the control logic of the present disclosure may be arranged and configured to classify the type of even t based on cues obtained from the video data. For example, an event classification module implemented in hardware and / or software may be included and may be used to classify each event with a specific event type.
[0042] In another aspect, the object detection module may be configured to determine differences between objects that are important and should be tracked versus objects that are unimportant. Unimportant objects may be in the background of the image data while objects of greater importance may be more toward the foreground of an image. For example, the object detection module may be configured to differentiate between a light pole that is always present in the image and of little importance, and a temporary and recently appearing broadcast tower set up by a news crew which may have a similar shape, color, and orientation as the light pole, but be much more interesting than the light pole for a variety of reasons. In another aspect, the object detection module may be configured to trigger an event when cars move around in the field-of-view, but not when trash cans, tables, chairs and the like are moved. These are a few nonexclusive examples of how the system of the present disclosure may automatically differentiate between foreground or background objects or events, or between objects or events that matter and those that do not matter.
[0043] In another aspect, the event detection according to the present disclosure may include accepting input specifying objects as familiar, unfamiliar, interesting, or uninteresting, and the like. In another aspect, event detection may include accepting input defining a new category for an object that the control logic of the present disclosure is unable to classify. In another aspect, the control logic of the present disclosure may be configured to automatically determine that: an object is a familiar or unfamiliar object based on frequency of appearance of that object in the video data. This may Include automatically modifying data about an object, or a type of object, to define that object as familiar after a predetermined number of appearances in the video feed. For example, the system may accept input defining a threshold value indicating that if the same object appears more than a threshold value number of times, the object should now be classified as familiar rather than unfamiliar. In another aspect, the system may generate a notification that a particular object, or type ofobject, has appeared numerous times and may need to be classified or marked as familiar. These aspects may be useful in training the control logic to better determine the types of objects and events that are of interest and differentiate them from those that are not of interest.
[0044] A sampling rate may be determined at 305. When an event is found in the video data at 303 according to the present disclosure, the system optionally determines a new sampling rate at 305 based on criteria such as the type of event found (if any is found) ai 304, The new sampling rate is optionally different for different types of events, and may be changed to a rate that is higher or lower than the default rate when an event is found.
[0045] The sampling rate is optionally associated with a general or specific type of event found to be occurring at 304. in the present disclosure, the sampling rate optionally defines the number of individual frames of video to retrieve from the video data per unit of recorded time. For example, a default sampling rate for portions of the video data that do not include events of interest may be up to one frame from every tenth of a second of raw video, up to one frame per second of raw video data, up to one frame from every five seconds, up to one frame from every 20 seconds, or more. A lower sampling rate results in reduced clarity and further reductions in playback time for the resulting sampled video. In one instance, when the raw video data includes 30 frames per second of recorded video, then these example sampling rates result in a reduction in the overall video data by a factor of 3, 30, 150, and 600 respectively.[00461 The system optionally samples frames from the video data ai 306 at the new sampling rate, and adds those sampled frames to output video data at 307. As used herein, the sampling process may involve choosing a frame by any suitable method. For example, if the sampling rate is one frame per second, the system may default to capturing the first frame of every second from the original raw video data, or the la st frame of every second of raw data. or the median (middle) frame, or any other suitable frame.
[0047] Ln another aspect, adding the sampled frames to output video data may include saving them to a memory, such as a buffer or output stream operable to store sampled video frames for later processing, hi another aspect, separate buffers or output streams may be created and managed in a memory by the control logic of the present disclosure to maintain separate storage locations and / or data structures for video data associated with each type of event.
[0048] T he sampled frames may optionally be displayed at 308 on a display device, and optionally saved at 309. The sampled frames may be saved in volatile or nonvolatile memory storage such as on a local memory chip, in a hard drive, or on a remote data store via a computer network, or any combination thereof If more video data is available to process at 311 , the actions repeat starting at 302, If no farther video data is available, the video may optionally be displayed at ,312 (optionally instead of or in spite of displaying it at ,308),
[0049] As noted above, if an event is found at 303, the type of event is determined and the data sampled accordingly. If an event is not, found at 303, the video data is sampled according to the default sampling rate, and this sampled video output is added to the sampled frames at 307, Processing thus moves on as discussed above.
[0050] The actions disclosed in Fig. 3 may be performed sequentially on multiple streams or files of video data, or they may be performed in parallel on many streams or files at the same time by multiple processors, execution threads, or other control logic. Some or all of the actions may thus occur simultaneously as multiple combinations of the disclosed action sequences are executed at the same time. Thus, multiple different video feeds or files may be processed sim ultaneously .
[0051] In another aspect, the system of the present disclosure, (as shown in Fig. 3 ) optionally may not be required to process all of the video data to determine event metadata before the sampling can occur. Metadata defining the precise location, length, or type of a given event may thus be either unknown, or unavailable, in advance of beginning the sampling aspects of the disclosed method. The reading, sampling, and output actions may occur sequentially for a given body of video data as the start and end of events may be determined as the data is obtained. This may result in a time lag between obtaining the video and producing the sampled output as the system may buffer the video data for some period of time in order to look ahead to determine the start, end, and / or type of events in the stream. In another aspect, generating and / or accessing event metadata may not be necessary as the actions taken by the system may be taken on individual files extracted from video data where each separate video file only incl udes raw unsampled video with types of events that were previously determined .
[0052] However, it can be advantageous io determine metadata about events in the video data. The system of the present disclosure is operable to detect events captured in the video data. Fig. 4 illustrates at 400 some examples of detected events located by the system of the present disclosure in video data. Video data 402 comprising one or more individual frames isillustrated as having a start time of 403 and an end time of 404. In one instance, the start time and end time are dictated by the start and end of the video file that the video data 402 is stored in. In another instance, the start time and end times may be arbitrarily applied to a continuous stream of video.
[0053] Detecting events (tor example, at 203) optionally includes preparing metadata information indicating the types of events found, and the start and end frames for each event. This metadata information may be used in preparing one or more output streams or files that include portions of the video data specific to each event. In another aspect, the system may be configured to extract one or more video frames between the start and end of a given event for each event found in the video data. In this example, the video for each specific event may be extracted from the video data and saved to a separate file or broadcast via a different video stream. In another aspect, other computers may be configured to only receive video associated with specific event types and to optionally ignore other video of other event types.
[0054] Looking at Fig. 4, the metadata may include information about the types of events found, and wha t frames of video from the raw video data 402 are associated with or define each event. For example, the metadata may indicate the location in the raw data where a given event may be found, the number of individual frames of video to retrieve from the video data for a given event, and / or the sampling rate to apply.
[0055] At 403, sampling begins at a default sampling rate and optionally includes accessing one or more individual frames of video from the video data. The number of individual frames of video data to retrieve optionally depends on the sampling rate. The default sampling rate may continue to be applied at 408 as long as no event is found at 404. Frames sampled at the default rate may continue to be applied to (he output video data (e.g. output video stream, video file, and the like) until an event is found. While no event is ongoing, the control logic optionally executes actions 403, 404, 408, and 409 repeatedly to generate the video output at the default sampling rate.
[0056] The dynamic sampling process of the present disclosure may include using software configuring the one or more processors to analyze the video data according to the present disclosure to determine when an event has occurred. This optionally includes determining the start of an event defined by one or more frames of video such as a starting frame with a first timestamp associated with the video, and an ending frame with a second timestamp associated with the video that is after the first timestamp in the video data. Theactions of determinin g the start and end for an event may be performed multiple times for a given set of video data to detect multiple events for the recorded video data.
[0057] Examples of detected events are illustrated at 400. Multiple events 405, 408, 41 1, 414, 417, 420, and 423 are shown. Event 405 has a start time at 406, and an end time at 407, event 408 has a start time at 409 and an end time at 410, and event 41 1 starts at 412 and ends at 413. Similarly, event 414 defines a start time 415 and an end time at 416, while event 417 defines start and end times 418 and 419. Event 420 is defined by a start time 421 and an end time 422, and event 423 is defined by a start time 424 and an end time 425. Thus it may be said that an event is defined by the start and end times, or that the event defines the start and end times, or optionally that the events have start and end times-all of which are synonymous.
[0058] In another aspect, portions of the video may be categorized based on whether or not each portion of video includes an event at all as defined by the disclosed control logic. In this example, multiple uncategorized periods of time are shown. For example the time period at 430 between the beginning of the video data at 403, and the beginning of event 405, the time period at 431 between the end of event 405 and the beginning of event 408, and the time period at 432 between the end of event 408 in the beginning of event 414, are all examples of uncategorized video data. Other examples are shown at 433, 434, and 435. These uncategorized portions may optionally be sampled at the default sampling rate.
[0059] In ano ther aspect, multiple e vents may be detec ted for a given period of time. For example, event 414 starts at 415 and ends at 416, however, event 411 may also be detected while event 414 is going on. Event 411 may include a start time at 412 and another at 413, both of which are after the start time of 415 for event 414, and before the end time of 416.
[0060] For example, event 414 may be classified as a movement, vehicle, person, or other event when the system of foe present disclosure determines that a delivery truck is in the viewing area of the camera capturing the video data 402. While the delivery truck is in view, and the packages to be deli vered are unloaded, a car, person, or other object may pass through the viewing area thus triggering a possibly different type of event (411) that has occurred (or is occurring) while another event (414) is ongoing.
[0061] In another example, event 417 begins at 418 while event 414 is still in progress, and continues on to 419 after event 414 is finished at 416. A wide variety of different scenarios may .result in this outcome. For example, a first individual may walk into view of the camera at 415 and remain in view until 416. A second individual may then walk into viewof the camera, perhaps conversing for a time with the first person (or not), and may stay within the field-of-view of the camera after the first person has left the viewing area.
[0062] In another example, the system of the present disclosure may be configured to detect when the camera is changing positions, or is moved, such as in the case of a surveillance camera that continuously and / or periodically rotates, zooms, or pans to adjust the viewing area of the camera over time. This process may be formed automatically by the camera, or may be actuated by input from a user using an input device such as a joystick, or other selector for repositioning the camera. The system of the present disclosure may detect at 415 that the camera has begun to move, and that the camera does not stop moving until 416, This camera motion may trigger logic in the disclosed control logic to change the sampling rate, such as to increase it. so that images captured during the camera movement may be less blurry than what may be otherwise provided when sampling according to the default sampling rate.
[0063] While that motion is in progress, the system may also detect, the presence of, for example, a vehicle passing through at 411 , and a delivery truck unloading packages at 417. The delivery truck, in this example, may have already been in position unloading packages, while the camera was in motion and may have been undetectable by the camera until the camera was repositioned to capture the track within the now updated field-of-view.
[0064] One example of the disclosed control logic that uses the event metadata of Fig. 4 is illustrated in Fig. 5 at 500, Video is obtained at 202 according to the present disclosure, and events may be determined in the video data according to what is shown at 400 in Fig. 4. One or more portions of the metadata generated a t 400 may be accessed or obtained at 502, The metadata may be accessed aspects of the control logic, such as, for example, one or more sampling modules 503 that optionally includes at least one sampling module 510. The sampling modules may be implemented in hardware, in software, or any combination thereof.
[0065] Actions that a sampling module 510 may take include determining the sampling rate based on the metadata al 51 1. The metadata optionally defines values indicating the start and end of a portion of the video associated with an event, the type of event it is associated with, and / or the sampling rate. In another aspect, the event type from the metadata 502 may be used as input to a process by which the sampling rate may be determined. For example, sampling rates may be a configurable array, map, or other data store that associates an event type with a sampling rate. The system of the present disclosure may be configured to accept input changing these sampling rates. In one example, when the metadata indicates a default(no event) type, the system may obtain the default sampling rate from the data store associating event types with sampling rates. Other event types are optionally accessed in a similar manner. The video is optionally sampled at 512 according to the sampling rate determined.
[0066] As illustrated, multiple sampling modules may be used by the control logic to handle video for mul tiple events simultaneously, or at about the same time. All of the sampling modules may sample data which is optionally ordered in time at 504 to provide for seamless viewing of the sampled data at 205. The sampled video may also be saved at 505, optionally before it is viewed, although it may be saved afterward as well.
[0067] In another aspect, the disclosed system and method optionally includes the capability to display the sampled portions of the video data (such as at 205). In one aspect, the sampled data may be viewed at a predetermined fixed or variable frame rate, such as up to 15 fps, up to 30 fps, up to 60 fps, or more. Displaying the sampled video optionally includes displaying the frames obtained from the video data for each individual sampled portion of the video. In one example, the sampled portions of the video may be displayed in temporal order, which is to say, in order with respect to time, starting with the earliest video data and displaying the sampled frames in order until the most recent frames are displayed, Ln another aspect, the system of the present disclosure may provide or include a user interface configured to accept input selecting a sampled portion of the video data to display. In another aspect, displaying the video at 205 optionally includes displaying the original unsampled video for the corresponding sampled portion of the video. The system of the present disclosure may provide or include a user interface configured to accept input selecting an original unsampled video to display instead of, or along with, the sampled portions of the video data, CLAUSES
[0068] The following numbered clauses set out examples of the disclosed concepts that may be useful in understanding the present disclosure:
[0069] Example 1 : A method for dynamic sampl ing of video data using one or more processors,
[0070] Example 2: The method of any other example including automatically adjusting the sampling rate according to specific criteria.
[0071] Example 3: The method of any other example including obtaining video data comprising multiple individual image frames.
[0072] Example 4: The method of any other example including detecting events captured in the video data.
[0073] Example 5: The method of any other example including obtaining portions of video data associated with events detected in the video data.
[0074] Example 6: The method of any other example including sampling frames of the portions of video data at varying sampling rates according to the type of event detected in each portion,
[0075] Example 7: The method of any other example including displaying the sampled portions of the video data at a predetermined frame rate.
[0076] Example 8: The method of any other example including capturing video data using a camera.
[0077] Example 9; The method of any other example wherein capturing video data includes capturing multiple individual images.
[0078] Example 10: The method of any other example wherein multiple individual images captured in the video data is captured a t a predetermined and / or fixed number of images (“frames”) captured per unit of time.
[0079] Example 1 1 : The method of any other example including categorizing portions of the video base on whether or not each portion of the video includes and event.
[0080] Example 12: The method of any other example including determining the start of an event defined by one or more frames of a video.
[0081] Example 13: The method of any other example including determining the end of an event defined by one or more frames of a video,
[0082] Example 14: The method of any other example including extracting one or more video frames between the start and end of a given event.
[0083] Example 15: The method of any other example including using the one or more processors to analyze the video data to determine when an event has occurred.
[0084] Example 16; The method of any other example including using an event classification module implemented in hardware and / or software to classify an event type,
[0085] Example 17: The method of any other example including repeating the actions of determining the start and end for an event multiple times to detect multiple events for the recorded video data.
[0086] Example 18: The method of any other example including accessing a database of familiar shapes to determine if an invent includes a familiar object.
[0087] Example 19: The method of any other example including determining if a detected object is a person, pet, vehicle, or other category of object.
[0088] Example 20: The method of any other example including accepting input defining a new category of object.
[0089] Example 21 : The method of any other example wherein data about familiar objects includes one or more images defining aspects of familiar shapes.
[0090] Example 22: The method of any other example including accepting input specifying familiar objects.
[0091] Example 23: The method of any other example including accepting input verifying that a portion of the video data includes familiar objects.
[0092] Example 24: The method of any other example including automatically determ ining that an object is a familiar or unfamiliar object based on frequency of appearance of that object in the video data.[0093 j Example 25: The method of any other example automatically modifying data about an object to define that object as familiar after a predetermined number of appearances in the video feed.
[0094] Example 26: The method of any other example wherein the system automatically modifies data about an object to define that object as familiar based on input received from a user interface.
[0095] Example 27: The method of any other example including accepting input defining an object as familiar.
[0096] Example 28: The method of any other example including determining that a movement event is the result of the camera being moved.
[0097] Example 29: The method of any other example including determining that a movement event is the result of an object passing through the field-of-view of a stationery camera.
[0098] Example 30: The method of any other example including differentiating movement of the camera from the movement of an objec t passing through the field-of-view of the camera.
[0099] Example 31 : The method of any other example wherein the control logic configured to determine the event type is in the camera.
[0100] Example 32: The method of any other example including sending the video data to a remote server.
[0101] Example 33: The method of any other example wherein the control logic configured to determine the event type is on a remote server.
[0102] Example 34: The method of any other example wherein the control logic for determining an event has occurred includes artificial intelligence.
[0103] Example 35: The method of any other example wherein the control logic for determining the type of that event has occurred includes artificial intelligence.
[0104] Example 36: The method of any other example including preparing portions of the video for display according to the events found in the video data.
[0105] Example 37: The method of any other example including accessing one or more individual frames of video from the video data.
[0106] Example 38: The method of any other example including determining the number of individual frames of video to retrieve from the video data.
[0107] Example 39: The method of any other example including determining a sampling rate defining the number of individual frames of video to retrieve from the video data per unit of recorded time.
[0108] Example 40: The method of any other example including ob taining frames from the recorded video data at a default sampling rate.
[0109] Example 41 : The method of any other example including changing the sampling rate of the video obtained from the recorded video data.
[0110] Example 42: The method of any other example wherein the sampling rate is increased when tin event is found,
[0111] Example 43: The method of any other example including determining a new sampling rate based on the category of event found.
[0112] Example 44: The method of any other example wherein the new sampling rate is different for different types of events.
[0113] Example 45: The method of any other example including changing the sampling rate to a person event rate when a person is detected in the video data.
[0114] Example 46: The method of any other example including changing the sampl ing rate to a unfamiliar person sampling rate when an unfamiliar person or object is detected in the video data.
[0115] Example 47: The method of any other example including changing the samplingrate to a pet sampling rate when a pet is detected in the video data.
[0116] Example 48: The method of any other example including changing the sampling rate to a movement sampling rate when movement is detected in the video data.
[0117] Example 49: The method of any other example including displaying the frames obtained from the video data for each individual sampled portion of the video.
[0118] Example 50: The method of any other example including displaying the sampled portions of the video in succession.
[0119] Example 51 : The method of any other example including displaying the sampled portions of the video at a predeterm ined fixed frame rate.
[0120] Example 52: The method of any other example including accepting input selecting a sampled portion of the video to display.[01211 Example 53: The method of any other example including displaying the original unsampled video for an individual sampled portion of the video.GLOSSARY OF DEFINITIONS AND ALTERNATIVES
[0122] While the invention is illustrated in the drawings and described herein, this disclosure is to be considered as illustrative and not restrictive in character. The present disclosure is exemplary in nature and all changes, equivalents, and modifications that come within the spirit of the inven tion are included. The detailed description is included herein to discuss aspects of the examples illustrated in the drawings for the purpose of promoting an understanding of the principles of the invention. No limitation of the scope of the invention is thereby intended. Any alterations and further modifications in the described examples, and any further applications of the principles described herein are contemplated as would normally occur to one skilled in the art to which the invention relates. Some examples are disclosed in detail, however some features (hat may not be relevant may have been left out for the sake of clarity.
[0123] Where there are references to publications, patents, and patent applications cited herein, they are understood to be incorporated by reference as if each individual publication, patent, or patent application were specifically and individually indicated io be incorporated by reference and set forth in its entirely herein.
[0124] Singular forms “a”, “an”, ‘The”, and the like include plural referents unless expressly discussed otherwise. As an illustration, references to “a device” or “the device” include one or more of such devices and equivalents thereof.
[0125] Directional terms, such as "up", "down", "top" "bottom'', "fore", "aft", "lateral'', "longitudinal”, "radial", "circumferential", etc., are used herein solely for the convenience ofthe reader in order to aid in the reader's understanding of the illustrated examples. The use of these directional terms does not in any manner limit the described, illustrated, and / or claimed features to a specific direction and / or orientation.
[0126] Multiple related items illustrated in the drawings with the same part number which are differentiated by a letter for separate individual instances, may be referred to generally by a distinguishable portion of the full name, and / or by the number alone. For example, if multiple “laterally extending elements’' 90A, 90B. 90C, and 90D tire illustrated in the drawings, the disclosure may refer to these as “laterally extending elements 90A-90D,” or as '‘laterally extending elements 90,” or by a distinguishable portion of the full name such as “elements 90”.
[0127] The language used in the disclosure are presumed to have only their plain and ordinary' meaning, except as explicitly defined below. The words used in the definitions included herein are to only have their plain and ordinary meaning. Such plain and ordinary meaning is inclusive of all consistent dictionary definitions from the most, recently published Webster’s and Random House dictionaries. As used herein, the following definitions apply to the following terms or to common variations thereof (e.g., singular / plural forms, past / present tenses, etc.):
[0128] “About” with reference to numerical values generally refers to plus or minus 10% of the stated value. For example, if the stated value is 4.375, then use of the term “about 4.375” generally means a range between 3.9375 and 4.8125.
[0129] “Activate” generally is synonymous with “providing power to”, or refers to “enabling a specific function” of a circuit or electronic device tha t already has power.
[0130] “Alert” generally refers to an audible and or visual message intended to inform a system's users or administra tors about a change in the operating conditions of the system or about an error condition of the system. In a graphical user interface, the alert may be displayed as a small window containing a message and / or photo detailing the alert information and parameters. In some examples, the alert may include a button (virtual or physical) to click in order to dismiss the alert. In other examples, the alert may be strictly audible and based on preset parameters. In a further example, the alert may be transmitted to a remote device for analysis. Other synonymous terms for alert include alarm and / or notification.
[0131] “And / or” is inclusive here, meaning “and” as well as “or”. For example. “P and or Q” encompasses, P, Q, and P with Q; and, such “P and / or Q” may include other elements as well.
[0132] “Artificial Intelligence’'’ generally refers to using a computer algorithm, or set of instructions, to simulate human intelligence processes by computer systems. Specific applications of .Al include expert systems, natural language processing, speech recognition and machine vision.
[0133] “Camera” generally refers to an apparatus or assembly that records i mages of a viewing area or field-of~view on a medium or in a memory. The images may be still images comprising a single frame or snapshot of the viewing area, or a series of frames recorded over a period of time that may be displayed in sequence to create the appearance of a moving image. Any suitable media may be used to store, reproduce, record, or otherwise maintain the images.
[0134] “Communication Link” generally refers to a connection between two or more communicating entities and may or may not include a communications channel between the communicating entities. The communication between the communicating entities may occur by any suitable means. For example the connection may be implemented as an actual physical link, an electrical link, an electromagnetic link, a logical link, or any other suitable linkage facilitating communication.
[0135] In the case of an actual physical link, communication may occur by multiple components in the communication link configured to respond to one another by physical movement of one element in relation to another. In the case of an electrical link, the communication link may be composed of multiple electrical conductors electrically connected to form the communication link.
[0136] In the case of an electromagnetic link, the connection may be implemented by sending or receiving electromagnetic energy at any suitable frequency, thus allowing communications to pass as electromagnetic waves. These electromagnetic waves may or may not pass through a physical medium such as an optical fiber, or through free space, or any combination thereof. Electromagnetic waves may be passed at any suitable frequency including any frequency in the electromagnetic spectrum.
[0137] A communication link may include any suitable combination of hardware which may include software components as well. Such hardware may include routers, switches, networking endpoints, repeaters, signal strength enters, hubs, and the like.
[0138] In the case of a logical link, the communication link may be a conceptual linkage between the sender and recipient such as a transmission station in the receiving station.Logical link may include any combination of physical, electrical, electromagnetic, or other types of communication links.
[0139] “Computer” generally refers to any computing device configured to compute a result from any number of input values or variables. A computer may include a processor for performing calculations to process input or output, A computer may include a memory for storing values to be processed by the processor, or for storing the results of previous processing.
[0140] A computer may also be configured to accept input and output from a wide array of input and output devices for receiving or sending values. Such devices include other computers, keyboards, mice, visual displays, printers, industrial equipment, and systems or machinery of all types and sizes. For example, a computer can control a network or network interface to perform various network communications upon request. The network interface may be part of the computer, or characterized as separate and remote from the computer.
[0141] A computer may be a single, physical, computing device such as a desktop computer, a laptop computer, or may be composed of multiple devices of the same type such as a group of servers operating as one device in a networked cluster, or a heterogeneous combination of different computing devices operating as one computer and linked together by a communication network. The communication network connected to the computer may also be connected to a wider network such as the internet. Thus a computer may include one or more physical processors or other computing devices or circuitry, and may also include any suitable type of memory.
[0142] A computer may also be a virtual computing platform having an unknown or fluctuating number of physical processors and memories or memory devices, A computer may thus be physically located in one geographical location or physically spread across several widely scatered locations with multiple processors linked together by a communication network to operate as a single computer,
[0143] T he concept of “computer” and “processor” within a computer or computing device also encompasses any such processor or computing device serving to make calculations or comparisons as part of the disclosed system. Processing operations related to threshold comparisons, rules comparisons, calculations, and the like occurring in a computer may occur, for example, on separate servers, the same server with separate processors, or ona virtual computing environment having an unknown number of physical processors as described above.
[0144] A computer may be optionally coupled to one or more visual displays and / or may include an integrated visual display. Likewise, displays may be of the same type, or a heterogeneous combination of different visual devices. A computer may also include one or more operator input: devices such as a keyboard, mouse, touch screen, laser or infrared pointing device, or gyroscopic pointing device to name just a few representative examples. Also, besides a display, one or more other output devices may be included such as a printer, plotter, industrial manufacturing machine, 3D printer, and the like. As such, various display, input and output device arrangements are possible.
[0145] Multiple computers or computing devices may be configured to communicate with one another or with other devices over wired or wireless communication links to form a network. Network communications may pass through various computers operating as network appliances such as switches, routers, firewalls or other network devices or interfaces before passing over other larger computer networks such as the internet. Communications can also be passed over (he network as wireless data transmissions carried over electromagnetic waves through transmission lines or free space. Such communications Include using WiFi or other Wireless Local Area Network (WLAN) or a cellular transmitter / receiver to transfer data.
[0146] “Control Logic” generally refers to hardware or software configured to implement an automatic decision making process by which inputs are considered, and corresponding outputs are generated. The output may be used for any suitable purpose such as to provide specific commands to machines or processes specifying specific actions to take. Examples of control logic incl ude computer programs executed by a processor io accept commands from a user and generate output according to the logic implemented in the program as executed by the processor. In another example, control logic may be implemented as a series of logic gates, rnicroconlrollers, and the like, electrically connected together in a predetermined arrangement so as to accept input from other circuits or computers and produce an output according to the rules implemented in the logic circuits.
[0147] “Controller” or “control circuit” generally refers (o a mechanical or electronic device configured, to control the behavior of another mechanical or electronic device. A controller or "‘control circuit” is optionally configured to provide signals or other electricalimpulses that may be received and interpreted by the controlled device to indicate how it should behave.
[0148] “Data” generally refers to one or more values of qualitative or quantitative variables that are usually the result of measurements. Data may be considered “atomic” as being finite individual units of specific information. Data can also be thought of as a value or set of values that includes a frame of reference indicating some meaning associated with the values. For example, the number “2” alone is a symbol that absent some context is meaningless. The number “2” may be considered “data” when it is understood to indicate, for example, (he number of items produced in an hour.
[0149] Data may be organized and represented in a structured format. Examples include a tabular representation using rows and columns, a tree representation with a set of nodes considered to have a parent-children relationship, or a graph representation as a set of connected nodes to name a few.
[0150] The term “data” can refer to unprocessed data or “raw data” such as a collection of numbers, characters, or other symbols representing individual facts or opinions. Data may be collected by sensors in controlled or uncontrolled environments, or generated by observation, recording, or by processing of other data. The word “data” may be used in a plural or singular form. The older plural form “datum” may be used as well.
[0151] “Database” also referred to as a “data store”, “data repository”, or “knowledge base” generally refers to an organized collection of data. The data is typically organized to model aspects of the real world, in a way that supports processes obtaining information about the world from the data. Access to the data is generally provided by a “Database Management System” (DBMS) consisting of an individual computer software program or organized set of software programs that allow user to interact with one or more databases providing access to data stored in the database (although user access restrictions may be put in place to limit access to some portion of the data). The DBMS provides various functions that allow entry, storage and retrieval of large quantities of information as well as ways to manage how that information is organized. A database is not generally portable across different DBMSs, but different DBMSs can interoperate by using standardized protocols and languages such as Structured Query Language (SQL), Open Database Connectivity (ODBC), Java Database Connectivity (JDBC), or Extensible Markup Language (XML) to allow a single application to work with more than one DBMS.
[0152] Databases and their corresponding database management systems are often classified according to a particular database model they support. Examples include a DBMS that relies on the “relational model” for storing data, usually referred to as Relational Database Management Systems (RDBMS). Such systems commonly use some variation of SQL to perform functions which include querying, formatting, administering, and updating an RDBMS. Other examples of database models include the “object’' model, chained model (such as in the case of a “blockchain” database), the “object-relational” model, the “tile”, “indexed file” or “flat-file” models, the “hierarchical” model, the “network” model, the “document” model, (he “XML” model using some variation of XML, the “entity-attribute- value” model, and others.
[0153] Examples of commercially available database management systems include PostgreSQL provided by the PostgreSQL Global Development Group; Microsoft SQL Server provided by the Microsoft Corporation of Redmond, Washington, USA: MySQL and various versions of the Oracle DBMS, often referred to as simply “Oracle” both separately offered by the Oracle Corporation of Redwood City, California, USA; the DBMS generally referred to as “SAP” provided by SAP SE of Walldorf, Germany; and the D22 DBMS provided by the International Business Machines Corporation (IBM ) of Armonk, New York, USA.
[0154] The database and the DBMS software may also be referred to collectively as a “database”. Similarly, the term “database” may also collectively refer to the database, the corresponding DBMS software, and a physical computer or collection of computers. Thus the term “database” may refer to the data, software for managing the data, and / or a physical computer that includes some or all of the data and / or the software for managing the data.
[0155] “Detection Zone” generally refers to an area within which an object may be detected. The detection zone may be either two dimensional and or three dimensional and may be defined by one or more sensors operable to detect objects within the detection zone, or by a control circuit that is responsive to the sensors.
[0156] “Display device” generally refers to any device capable of being controlled by an electronic circuit or processor to display information in a visual or tactile. A display device may be configured as an input device taking input from a user or other system (e.g. a touch sensitive computer screen), or as an output device generating visual or tactile information, or the display device may configured to operate as both an input or output device al the same time, or at different times.
[0157] T he output may be two-dimensional, three-dimensional, and / or mechanical displays and includes, but is not limited to, the following display technologies: Cathode ray tube display (CRT), Light-emitting diode display (LED), Electroluminescent display (ELD), Electronic paper. Electrophoretic Ink (E-fok), Plasma display panel (PDP), Liquid crystal display (LCD), High-Performance Addressing display (HPA), Thin-film transistor display (TFT), Organic light-emitting diode display (OLED), Surface-conduction electron-emitter display (SED), Laser TV, Carbon nanotubes, Quantum dot display . Interferometric modulator display (1M0D), Swept- volume display, Varifocal mirror display, Emissive volume display. Laser display. Holographic display. Light field displays, Volumetric display. Ticker tape, Split-flap display, Flip-disc display (or flip-dot display), Rollsign, mechanical gauges with moving needles and accompanying indicia. Tactile electronic displays (aka refreshable Braille display), Optacon displays, or any devices that either alone or in combination are configured to provide visual feedback on the status of a system, such as the “check engine” l ight, a “low altitude” warning light, an array of red, yellow, and green indicators configured to indicate a temperature range.
[0158] “Input Device” generally refers to any device coupled to a computer that is configured to receive input and deliver the input to a processor, memory, or other part of the computer. Such input devices can include keyboards, mice, trackballs, touch sensitive pointing devices such as touchpads, or touchscreens. Input devices also include any sensor or sensor array for detecting environmental conditions such as temperature, light, noise, vibration, humidity, and the like.
[0159] “Means For*’ in a claim invokes 35 U.S.C. 112(f), literally encompassing the recited function and corresponding structure and equivalents thereto. Its absence does not, unless there otherwise is insufficient structure recited for that claim element. Nothing herein or elsewhere restricts the doctrine of equivalents available to the patentee.
[0160] “Memory” generally refers to any storage system or device configured to retain data or information. Each memory may include one or more types of solid-state electronic memory, magnetic memory , or optical memory , just to name a few. Memory may use any suitable storage technology, or combination of storage technologies, and may be volatile, nonvolatile, or a hybrid combination of volatile and nonvolatile varieties. By way of nonlimiting example, each memory may incl ude solid-state electronic Random Access Memory (.RAM), Sequentially Accessible Memory (SAM) (such as the First-In, First-Out (FIFO) variety or the Last-ln-First-Out (LIFO) variety ). Programmable Read Only Memory(PROM), Electronically Programmable Read Only Memory (EPROM), or Electrically Erasable Programmable Read Only Memory (EEPROM).
[0161] Memory can refer to Dynamic Random Access Memory (DRAM) or any variants, including static random access memory (SRAM), Burst SRAM or Synch Burst SRAM (BSRA.M), Fast Page Mode DRAM (FPM DRAM), Enhanced DRAM (EDRAM), Extended Data Output RAM (EDO RAM), Extended Data Output DRAM (EDO DRAM), Burst Extended Data Output DRAM (REDO DRAM), Single Data Rate Synchronous DRAM (SDR SDRAM), Double Data Rate SDRAM (DDR SDRAM), Direct Rambus DRAM (DRDRAM), or Extreme Data Rate DRAM (XDR DRAM).
[0162] Memory can also refer to non-volatile storage technologies such as non-volatile read access memory (NVRAM), flash memory, non-volatile static RAM (nvSRAM), Ferroelectric RAM (FeRAM), Magnetoresistive RAM (MRAM), Phase-change memory (PRAM), conductive-bridging RAM (CBRAM), Silicon-Oxide-Nitride-Oxide-Silicon (SONOS), Resistive RAM (RRAM), Domain Wall Memory (DWM) or “Racetrack” memory, Nano-RAM (NRAM), or Millipede memory. Other non-volatile types of memory include optical disc memory (such as a DVD or CD ROM), a magnetically encoded hard disc or hard disc platter, floppy disc, tape, or cartridge media. The concept of a '‘memory” includes the use of any suitable storage technology or any combination of storage technologies.
[0163] “Module” or “Engine” generally refers to a collection of computational or logic circuits implemented in hardware, or lo a series of logic or computational instructions expressed in executable, object, or source code, or any combination thereof, configured to perform tasks or implement processes. A module may be implemented in software maintained in volatile memory in a computer and executed by a processor or other circuit. A module may be implemented as software stored in an erasable-programmable nonvolatile memory and executed by a processor or processors. A module may be implanted as software coded into an .Application Specific Information Integrated Circuit (ASIC). A module may be a collection of digital or analog circuits configured to control a machine lo generate a desired outcome.
[0164] Modules may be executed on a single computer with one or more processors, or by multiple computers with multiple processors coupled together by a network. Separate aspects, computations, or functionality performed by a module may be executed by separateprocessors on separate computers, by the same processor on the same computer, or by different computers at different times.
[0165] “Multiple” as used herein is synonymous with the term “plurality” and refers to more than one, or by extension, two or more.
[0166] “Network” or “Computer Network” generally refers to a telecommunications network that allows computers to exchange data. Computers can pass data to each other along data connections by transforming data into a collection of datagrams or packets. The connections between computers and the network may be established using either cables, optical fibers, or via electromagnetic transmissions such as for wireless network devices.
[0167] Computers coupled to a network may be referred to as “nodes” or as “hosts” and may originate, broadcast, route, or accept data from fee network. Nodes can include any computing device such as personal computers, phones, servers as well as specialized computers that operate to maintain the flow of data across the network, referred to as “network devices”. Two nodes can be considered “networked together” when one device is able to exchange information wife another device, whether or not they have a direct connection to each other.
[0168] Examples of wired network connections may include Digital Subscriber Lines (DSL), coaxial cable lines, or optical fiber lines. The wireless connections may include BLUETOOTH, Worldwide Interoperability for Microwave Access (WiMAX), infrared channel or satellite band, or any wireless local area network (Wi-Fi) such as those implemented using the Institute of Electrical and Electronics Engineers’ (IEEE) 802.11 standards (e.g. 802. 11(a), 802.11(b), 802, 1 1(g), or 802.1 l(n) to name a tew). Wireless links may also include or use any cellular network standards used to communicate among mobile devices including 1G, 2G, 3G, or 4G. The network standards may qualify as 1G, 2G, etc. by fulfilling a specification or standards such as the specifications maintained by international Telecommunication Union (ITU). For example, a network may be referred to as a “3G network” if it meets the criteria in the International Mobile Telecommunicalions-2000 (IMT- 2000) specification regardless of what it may otherwise be referred to. A network may be referred to as a “4G network” if it meets the requirements of the International Mobile Telecommunications Advanced (IMT Advanced) specification. Examples of cellular network or other wireless standards include AMPS, GSM, GPRS, UMTS, LTE, LTE Advanced, Mobile WiMAX, and WiMAX-Advanced.
[0169] Cellular network standards may use various channel access methods such as FDMA, TDMA, CDM A, or SDMA. Different types of data may be transmitted via different links and standards, or the same types of data may be transmitted via different links and standards,
[0170] The geographical scope of the network may vary widely. Examples include a body area network (BAN), a personal area network (PAN), a low power wireless Personal Area Network using IPv6 (6L0WPAN), a local-area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), or the Internet.
[0171] A network may have any suitable network topology defining the number and use of the network connections. The network topology may be of any suitable form and may include point-to-point, bus, star, ring, mesh, or tree. A network may be an overlay network which is virtual and is configured as one or more layers that use or “lay on top of’ other networks.
[0172] A network may utilize different communication protocols or messaging techniques including layers or stacks of protocols. Examples include the Ethernet protocol, the internet protocol suite (TCP / IP), the ATM (Asynchronous Transfer Mode) technique, the SONET (Synchronous Optical Networking) protocol, or the SDE1 (Synchronous Digital Elierarchy) protocol. The TCP / IP internet protocol suite may include application layer, transport layer, internet layer (including, e.g., IPv6), or the link layer.
[0173] “Neural Network” generally refers to a collection of cooperating computational nodes implemented in hardware and / or software that use a mathematical or computational model for information processing based on a connectionistic approach to computation. A neural network may be an adaptive system that changes its structure based on external or internal information that flows through the network. The connections between nodes may be “weighted” to achieve specific outcomes given a wide range of inputs. A more positive weight reflects a more relevant or more “excitatory” connection, while a more negative weight reflects a more uninteresting or more “inhibitory” connections. All inputs to each node are modified according to the weights and summed. This activity is referred to as a linear combination. Finally, an activation function is generally used by each node to control the amplitude of the output. For example, an acceptable range of outpu t is usually between 0 and 1 , or it could be " I and 1 . The output of each node may then be fed as input to othernodes, and thus the overall network of nodes may be able to solve complex problems and / or to adapt to changes in the input over time.
[0174] These artificial networks may be used for predictive modeling, adaptive control and applications where they can be trained via a dataset. Self-learning resulting from experience can occur within networks, which can derive conclusions from a complex and seemingly unrelated set of information.
[0175] “Optionally’* as used herein means discretionary; not required; possible, but not compulsory'; left to personal choice.
[0176] “Output Device” generally refers to any device or collection of devices that is controlled by computer to produce an output. This includes any system, apparatus, or equipment receiving signals from a computer to control the device to generate or create some type of output. Examples of output devices include, but are not limited to, screens or monitors displaying graphical output, any projector a projecting device projecting a two- dimensional or three-dimensional image, any kind of printer, plotter, or similar device producing either two-dimensional or three-dimensional representations of the output fixed in any tangible medium (e.g. a laser printer printing on paper, a lathe controlled to machine a piece of metal, or a three-dimensional printer producing an object). An output device may also produce intangible output such as, for example, data stored in a database, or electromagnetic energy transmitted through a medium or through free space such as audio produced by a speaker controlled by the computer, radio signals transmitted through free space, or pulses of light passing through a fiber-optic cable.[017'7] “Personal computing device” generally refers to a computing device configured lor use by individual people. Examples include mobile devices such as Personal Digital Assistants (PDAs), tablet computers, wearable computers installed in items worn on the human body such as in eye glasses, watches, laptop computers, portable music / video players, computers in automobiles, or cellular telephones such as smart phones. Personal computing devices can be devices that are typically not mobile such as desk top computers, game consoles, or server computers. Personal computing devices may include any suitable input-output devices and may be configured to access a network such as through a wireless or wired connection, and / or via other network hardware.
[0178] “Portion” means a part of a whole, either separated from or Integrated with It.
[0179] “Predominately” as used herein is synonymous with greater than 50%.
[0180] “Rule” generally refers to a conditional statement with at least two outcomes. A rule may be compared to available data which can yield a positive result Call aspects of the conditional statement of the rule are satisfied by the data), or a negative result (at least one aspect of the conditional statement of the rule is not satisfied by the data). One example of a rule is shown below as pseudo code of an “if then / else” statement that may be coded in a programming language and executed by a processor in a computer: if (clouds . are Grey ( ) and( clouds . numberOf Clouds > I C C } ) then prepare for rain;} else {Prepare for sunshine;
[0181] “Sampling Rate” generally refers to a number of samples taken per second. This is common in multiple areas of interest such as in converting analog wave (a continuously changing value or collection of values) to digital input (a discrete wave ). Two equivalent units for sampling rate are samples per second (sps) or Hertz (Hz).
[0182] “Substantially” generally refers to the degree by which a quantitative representation may vary from a stated reference without resulting in an essential change of the basic, function of the subject matter at issue. The term “substantially” is util ized herein to represent the inherent degree of uncertainty that may be attributed to any quantitative comparison, value, measurement, and / or other representation.
[0183] “Transformer” generally refers to a deep learning architecture that implements a parallel multi-head attention mechanism. Transformers may be applied to text or classify image input.
[0184] Transformers generally include an initial step by which the input is apportioned or broken up into manageable pieces. For text, tokenizers may be applied to convert text into tokens. In the ease of image or video input, image processing may be applied to convert an image to a collection or sequence of flattened image patches.
[0185] The transformer architecture optionally also includes a single embedding layer, which converts the portions and positions of the portions into vector representations, one ormore transformer layers, which carry out repeated transformations on the vector representations, extracting more and more image context or linguistic information (and these generally consist of alternating attention and feedforward layers), and optionally, an unembedding layer, which converts the final vector representations back to a probability distribution over the different portions,
[0186] I 'ext may be split into n-grams encoded as tokens and each token converted into a vector via a table lookup. At each layer, each token is then contextualized within foe scope of the context window with other (unmasked) tokens via a parallel multi-head attention mechanism allowing the signal for key tokens to be amplified and less important tokens to be diminished.
[0187] the final vector representations back to a probability distribution over the tokens.
[0188] “Triggering a Rule” generally refers to an outcome that follows when all elements of a conditional statement expressed in a rule are satisfied. In this context, a conditional statement may result in either a positive result (all conditions of the rule are satisfied by the data), or a negative result (at least one of the conditions of the rule is not satisfied by the data) when compared to available data. The conditions expressed in the rule are triggered if all conditions are met causing program execution to proceed along a different path than if the rule is not triggered.
[0189] “User Interface” generally refers an aspect of a device or computer progrant that provides a means by which the user and a device or computer program interact, in particular by coordinating the use of input devices and software, A user interface may be said to be “graphical” in nature in that the device or software executing on the computer may present images, text, graphics, and the like using a display device to present output meaningful to the user, and accept input from the user in conjunction with the graphical display of the output.
[0190] “Viewing Area”, “Field of View”, or “Field of Vision” is the extent of the observable world that is seen at any given moment. In ease of optical instruments, cameras, or sensors, it is a solid angle through which a detector is sensitive to electromagnetic radiation that include light visible io the human eye, and any other form of electromagnetic radiation that may be invisible to humans.
[0191] “Wi-Fi” generally refers to a family of wireless network protocols that are based on the IEEE 802.11 family of standards, Wi-Fi networks are commonly used for local area networking of devices so that these devices may communicate with each other and with a broader computer network such as the internet, Wi-Fi protocols define how enabled devicesmay exchange data wirelessly via radio waves. Wi-Fi wireless connections may be useful for providing wireless communications links between desktop and laptop computers, cameras, tablet computers, smartphones, smart TVs, printers, smart speakers, and the like with wireless network access devices to connect them to the Internet.
[0192] Wi-Fi uses multiple parts of the IEEE 802 protocol family and is designed to be operable seamlessly with wired communication protocols, such as Ethernet. Compatible devices can network through wireless access points to each other as well as to wired devices and the Internet. The different versions of Wi-Fi are specified by various IEEE 802.11 protocol standards, with different radio technologies determining radio bands, and (he maximum ranges, and data rates that may be achieved. For example, Wi-Fi uses the 2.4 gigahertz (120 mm wavelength) UHF and 5 gigahertz (60 mm wavelength) SHF radio bands, which may be subdivided into multiple channels.
[0193] The radio frequencies typically used by Wi-Fi transmitters and receivers have relatively high absorption rates and work best for line-ol-sight communication links. Many common obstructions such as walls, pillars, home appliances, etc. may greatly reduce range, but interference between different networks in crowded environments is usually minimal. In one example, a Wi-Fi network access point may have a range of about 65 feet indoors, or as much as 500 feet outdoors. Wireless network access points may include a single transmitter / receiver to cover a single room to a multiple transmittere-'receivers spread over square miles of area to provide overlapping access to client devices.
Claims
CLAIMSWhat is claimed is:I , A method, comprising: obtaining video data using one or processors of one or more computers; detecting events captured in the video data using the one or more processors; determining a type of event defined by one or more portions of the video; sampling the video data according to the type of even t detected: separating the video data into separate portions wherein each portion includes an event of a predetermined type using the one or more processors; sampling the video data at varying sampling rates to create one or more sampled portions of the video data; and displaying sampled portions of the video data, wherein the video is displayed at a predetermined frame rate, and wherein the sampling rate is increased when an event is found, 2, The method of claim 1, comprising: assigning a default event type for portions of the video that do not include events of interest,3. The method of claim 1 , comprising: automatically adjusting the sampling rate according to specific criteria. 4, The method of claim 1, comprising: sampling frames of one or more portions of video data at varying sampling rates according to the type of event detected in each portion,5. The method of claim i, comprising: capturing the video data using a camera and obtaining the captured video data from the camera,6. The method of claim 5, control logic configured to determine the event type is implemented in hardware and / or software that is in the camera,7. The method of claim 1 , wherein the video data is captured at a predetermined and / or fixed number of images (“frames”) captured per unit of time.
8. The method of claim i, comprising: determining a start of an event defined by one or more frames of a video; and determining an end of an event defined by one or more frames of a video.
9. The method of claim 8, comprising:extracting one or more video frames between the start and end of an event.
10. The method of claim 8, comprising: repeating determining the start and determining the end of the event multiple times io detect multiple events for the video data.
11. The method of claim 1, comprising: using the one or more processors to analyze the video data to determine when an event has occurred.
12. The method of claim I, comprising: using an event classification module implemented in hardware and / or software to classify an event type for a portion of the video data.The method of claim 1 , comprising: accessing a database of familiar shapes, faces, vehicles, and / or other objects to determine if an invent includes a familiar object.
14. The method of claim 1 , comprising: differentiating movement of a camera capturing the video data from movement of an object passing through a field-of-view defined by the camera.
15. The method of claim 1, comprising: sending the video data to a remote server.
16. The method of claim 1, wherein control logic configured to determine the event type is executed in hardware and / or software that is on a remote server.
17. The method of claim 1 , comprising: accessing one or more individual frames of video from the video data.
18. The method of claim i, comprising: determining a number of individual frames of video to retrieve from the video data.
19. The method of claim 18, comprising: determining a sampling rate defining the number of individual frames of video to retrieve from the video data per unit of recorded time.
20. The method of claim 1 , comprising: obtaining frames from the video data at a default sampling rate. 21 . The method of claim 1, comprising: changing foe sampling rate of foe video obtained from the video data.
22. The method of claim 1. comprising: determining a new sampling rate based on a category or type of event found.
23. The method of claim I, comprising: changing the sampling rate to a person event rate when a person is detected in the video data.
24. The method of ciaim 1. comprising: changing the sampling rate to a unfamiliar person sampling rate when an unfamiliar person or object is detected in the video data.
25. The method of claim 1 , comprising: displaying the sampled portions of the video at a predetermined fixed frame rate.
Citation Information
Patent Citations
Providing Method For Video Contents and Electronic device supporting the same
KR1020180013325A
Method and monitoring camera for detecting intrusion in real time based image using artificial intelligence
KR102021441B1
Logistic movement optimization system of chain conveyor
KR1020250067558A
Apparatus for Electrolytic Reduction and Method for Electrolytic Reduction
KR102182475B1
Method and system for tracking an object in a defined area
US20180107880A1