Systems and methods for ai-based control of a security system at a location
The AI-based framework dynamically adjusts camera views and sequences video clips to address the limitations of traditional security systems, providing comprehensive event capture and tracking in security systems.
Patent Information
- Application Number
- PCT/US2025/011903
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-22
- Filing Date
- 2025-01-16
- Publication Date
- 2025-07-31
AI Technical Summary
Traditional security systems face challenges in optimally capturing and sequencing events across multiple camera views, leading to poor detail and tracking of events due to fixed fields of view, which can compromise safety.
An AI-based framework dynamically determines and adjusts camera fields of view to ensure comprehensive event capture, combines video clips from multiple cameras into a sequential timeline, and generates interactive displays for event tracking.
Enhances security system efficiency by ensuring thorough event capture and sequencing, allowing users to navigate and understand event patterns across multiple camera views effectively.
Smart Images

Figure US2025011903_31072025_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR AI-BASED CONTROL OF A SECURITY SYSTEM AT A LOCATIONCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of, and priority to, U.S. Provisional Patent Application No. 63 / 623,592 filed January 22, 2024, its entirety of which is incorporated herein by reference.FIELD OF THE DISCLOSURE
[0002] The present disclosure is generally related to a location monitoring and control system, and more particularly, to a decision intelligence (Dl)-based computerized framework for automatically and dynamically capturing and displaying security footage for a location.BACKGROUND
[0003] Security systems applied to a location involve many components that are enabled to detect activity and alert a particular set of users as to the occurrence of such detected activity'.SUMMARY OF THE DISCLOSURE
[0004] Traditionally, a comprehensive security system for a location (e.g., home, office, for example) includes a variety' of components designed to work together to provide effective protection. Such components can include, but are not limited to, surveillance cameras, digital video recorders, motion sensors, access control systems, intrusion detection systems, alarms, perimeter security, security lighting, remote monitoring, intercom systems, fire and smoke detection, environmental sensors, central monitoring stations, back-up power systems, and the like, or some combination thereof. By combining these components, a security system can create a layered approach to protect a location from various threats and provide a more comprehensive security solution.
[0005] Accordingly, security cameras (referred to as cameras herein) are a crucial component of security’ systems. Cameras can provide surveillance, by continuously monitoring a designated area, and capturing real-time video footage. Accordingly, cameras can effectively document events and activities, providing a visual record that, for example, can be used for investigations or legal purposes. Moreover, cameras can allow for remote monitoring, enabling users to view live or recorded footage from a distance via the internet or a dedicated network.
[0006] Cameras can be equipped with, paired with and / or associated with motion sensors that trigger recording and / or alerts when unusual movement is detected, helping to focus attentionon potential threats. As discussed herein, camera usage (e.g., the field of view and / or number of cameras used) can be scaled to accommodate the size and layout of different locations, ensuring comprehensive coverage.
[0007] Overall, security cameras play a vital role in enhancing the safety and security of various environments by providing a constant and vigilant "eye" on the surroundings.
[0008] To that end, the disclosed systems and methods provide a novel framework that provides end-to-end (E2E) control and management of security footage that can be captured at a location, while ensuring that the captured footage can be curated to track events across the location from multiple cameras, as well as provide a timeline of such events.
[0009] By way of background, cameras are typically mounted at specific positions at a location (e.g.. a doorbell camera on the door frame of a door). This often prevents optimal views from the camera, should events be occurring adjacent to the camera’s field of view, and / or in the periphery’ of such field of view. This can lead to poor detail and / or tracking of events that can impact the safety of the users at the location.
[0010] Accordingly, the disclosed systems and methods provide functionality’ for the disclosed framework by enabling automatically and / or dynamically determined fields of view to be determined and applied, such that the proper lens scope, framing, zoom and / or dimensionality7can be applied to ensure that an event, no matter where in the perceivable view of the camera, can be properly and efficiently captured. In some embodiments, such field of views can be selected and / or based on user input; and in some embodiments, the field of view can be automatically determined and applied via the frameworks computational analysis of where an event is occurring respective to the closest proximate camera(s), as discussed in more detail below.
[0011] For example, as discussed below, the disclosed framework can utilize a full frame image, where a plurality of cropped image frames (or video clips) can be captured, which can correlate to fields of view including, but not limited to, “tall,” “full,” “wide,” and the like, or any other type of known or to be know field of view applicable to a camera’s video / image capture. For example, if an event is detected to the “far right” of the field of view, the framework can select “wide” to ensure that the camera adequately captures the event.
[0012] By way of further background, users often struggle to navigate amongst a library7(or repository ) of independent clips when trying to understand the movements and patterns of a person or object. For example, a stranger walked on the user’s property, and such movementwas captured by three (3) cameras at the user's house. Currently, there is no way to computationally determine the sequence of the stranger's movements on the property.
[0013] Accordingly, the disclosed framework can computationally analyze the captured event footage (e.g., image / video files), and determine the time and sequence of such stranger’s movements (e.g., first walked up the driveway, then around the back to the shed, then to the side window, and the like). Accordingly, as discussed below in more detail, such combination of video clips can be combined into a single video file, such that a security footage file is created that tracks each detected event, in sequential order, from a plurality of source capturing cameras.
[0014] Moreover, in some embodiments, the video clips for a tracked event (e.g., the stranger moving about a location, for example, as discussed supra). can be further analyzed, for which an interactive display can be compiled, generated and caused to be displayed within a user interface (UI). For example, multiple timelines can be created, which can be for, but not limited to, tracked events, related events, events from a camera or set of cameras, a location, position at the location (e.g., front door), time period, identified user, and the like, or some combination thereof, can be generated and displayed. Such timelines can be displayed in a manner that enables the viewing user to track the sequential order of the timeline, which can involve selecting nodes on the timeline that are displayed as thumbnails, which enable the rendering of an associated clip with that time / event. In some embodiments, timelines can be displayed vertically; and in some embodiments, such timelines can be displayed horizontally, and or in any other manner in which the sequence of events of such timelines, as w ell as the sequence of such timelines, can be readily identifiable and understood to indicate which events occurred at the location, and in which order. As provided below, for example, such timelines can enable the navigation across multiple cameras, and the various clips captured from such cameras.
[0015] According to embodiments of the instant disclosure, it should be understood that the discussion herein that references a location can correspond to, but not be limited to, a home, office, building and / or any other ty pe of definable structure and / or geographic location for which a security system can be provided.
[0016] According to some embodiments, a method is disclosed for a Dl-based computerized framework for automatically and dynamically capturing and displaying security footage for a location. In accordance with some embodiments, the present disclosure provides a non- transitory computer-readable storage medium for carrying out the above-mentioned technical steps of the framework’s functionality. The non-transitory computer-readable storage mediumhas tangibly stored thereon, or tangibly encoded thereon, computer readable instructions that when executed by a device cause at least one processor to perform a method for automatically and dynamically capturing and displaying security footage for a location.
[0017] In accordance with one or more embodiments, a system is provided that includes one or more processors and / or computing devices configured to provide functionality in accordance with such embodiments. In accordance with one or more embodiments, functionality is embodied in steps of a method performed by at least one computing device. In accordance with one or more embodiments, program code (or program logic) executed by a processor(s) of a computing device to implement functionality7in accordance with one or more such embodiments is embodied in, by and / or on a non- transitory computer-readable medium.DESCRIPTIONS OF THE DRAWINGS
[0018] The features, and advantages of the disclosure will be apparent from the following description of embodiments as illustrated in the accompanying drawings, in which reference characters refer to the same parts throughout the various views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating principles of the disclosure:
[0019] FIG. 1 is a block diagram of an example configuration within which the systems and methods disclosed herein could be implemented according to some embodiments of the present disclosure;
[0020] FIG. 2 is a block diagram illustrating components of an exemplary system according to some embodiments of the present disclosure;
[0021] FIG. 3 illustrates an exemplary workflow according to some embodiments of the present disclosure;
[0022] FIG. 4 illustrates an exemplary workflow according to some embodiments of the present disclosure;
[0023] FIG. 5 illustrates an exemplary workflow according to some embodiments of the present disclosure;
[0024] FIG. 6 depicts an exemplary implementation of an architecture according to some embodiments of the present disclosure;
[0025] FIG. 7 depicts an exemplary7implementation of an architecture according to some embodiments of the present disclosure; and
[0026] FIG. 8 is a block diagram illustrating a computing device showing an example of a client or server device used in various embodiments of the present disclosure.DETAILED DESCRIPTION
[0027] The present disclosure will now be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of non-limiting illustration, certain example embodiments. Subject matter may, however, be embodied in a variety of different forms and, therefore, covered or claimed subject matter is intended to be construed as not being limited to any example embodiments set forth herein; example embodiments are provided merely to be illustrative. Likewise, a reasonably broad scope for claimed or covered subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, or systems. Accordingly, embodiments may, for example, take the form of hardware, software, firmware or any combination thereof (other than software per se). The following detailed description is, therefore, not intended to be taken in a limiting sense.
[0028] Throughout the specification and claims, terms may have nuanced meanings suggested or implied in context beyond an explicitly stated meaning. Likewise, the phrase “in one embodiment'’ as used herein does not necessarily refer to the same embodiment and the phrase “in another embodiment'’ as used herein does not necessarily refer to a different embodiment. It is intended, for example, that claimed subject matter include combinations of example embodiments in whole or in part.
[0029] In general, terminology' may be understood at least in part from usage in context. For example, terms, such as “and”, “or”, or “and / or,” as used herein may include a variety of meanings that may depend at least in part upon the context in which such terms are used. Typically, “or” if used to associate a list, such as A, B or C, is intended to mean A, B, and C, here used in the inclusive sense, as w'ell as A, B or C, here used in the exclusive sense. In addition, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a,” “an,” or “the,” again, may' be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may,instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0030] The present disclosure is described below with reference to block diagrams and operational illustrations of methods and devices. It is understood that each block of the block diagrams or operational illustrations, and combinations of blocks in the block diagrams or operational illustrations, can be implemented by means of analog or digital hardware and computer program instructions. These computer program instructions can be provided to a processor of a general purpose computer to alter its function as detailed herein, a special purpose computer, ASIC, or other programmable data processing apparatus, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, implement the functions / acts specified in the block diagrams or operational block or blocks. In some alternate implementations, the functions / acts noted in the blocks can occur out of the order noted in the operational illustrations. For example, two blocks shown in succession can in fact be executed substantially concurrently or the blocks can sometimes be executed in the reverse order, depending upon the functionality / acts involved.
[0031] For the purposes of this disclosure a non-transitory computer readable medium (or computer-readable storage medium / media) stores computer data, which data can include computer program code (or computer-executable instructions) that is executable by a computer, in machine readable form. By way of example, and not limitation, a computer readable medium may include computer readable storage media, for tangible or fixed storage of data, or communication media for transient interpretation of code-containing signals. Computer readable storage media, as used herein, refers to physical or tangible storage (as opposed to signals) and includes without limitation volatile and non-volatile, removable and nonremovable media implemented in any method or technology for the tangible storage of information such as computer-readable instructions, data structures, program modules or other data. Computer readable storage media includes, but is not limited to, RAM, ROM, EPROM, EEPROM, flash memory or other solid state memory technology, optical storage, cloud storage, magnetic storage devices, or any other physical or material medium which can be used to tangibly store the desired information or data or instructions and which can be accessed by a computer or processor.
[0032] For the purposes of this disclosure the term “server’' should be understood to refer to a service point which provides processing, database, and communication facilities. By way of example, and not limitation, the term “server” can refer to a single, physical processor withassociated communications and data storage and database facilities, or it can refer to a networked or clustered complex of processors and associated network and storage devices, as well as operating software and one or more database systems and application software that support the services provided by the server. Cloud servers are examples.
[0033] For the purposes of this disclosure a “network” should be understood to refer to a network that may couple devices so that communications may be exchanged, such as between a server and a client device or other types of devices, including between wireless devices coupled via a wireless network, for example. A network may also include mass storage, such as network attached storage (NAS), a storage area network (SAN), a content delivery network (CDN) or other forms of computer or machine-readable media, for example. A network may include the Internet, one or more local area networks (LANs), one or more wide area networks (WANs), wire-line type connections, wireless type connections, cellular or any combination thereof. Likewise, sub-networks, which may employ differing architectures or may be compliant or compatible with differing protocols, may interoperate within a larger network.
[0034] For purposes of this disclosure, a “wireless network” should be understood to couple client devices with a network. A wireless network may employ stand-alone ad-hoc networks, mesh networks, Wireless LAN (WLAN) networks, cellular networks, or the like. A wireless network may further employ a plurality of network access technologies, including Wi-Fi, Long Term Evolution (LTE), WLAN, Wireless Router mesh, or 2nd, 3rd, 4thor 5thgeneration (2G, 3G, 4G or 5G) cellular technology, mobile edge computing (MEC), Bluetooth, 802.1 Ib / g / n, or the like. Network access technologies may enable wide area coverage for devices, such as client devices with varying degrees of mobility’, for example.
[0035] In short, a wireless network may include virtually any ty pe of wireless communication mechanism by which signals may be communicated between devices, such as a client device or a computing device, between or within a network, or the like.
[0036] A computing device may be capable of sending or receiving signals, such as via a wired or wireless network, or may be capable of processing or storing signals, such as in memory as physical memory states, and may, therefore, operate as a server. Thus, devices capable of operating as a server may include, as examples, dedicated rack-mounted servers, desktop computers, laptop computers, set top boxes, integrated devices combining various features, such as two or more features of the foregoing devices, or the like.
[0037] For purposes of this disclosure, a client (or user, entity', subscriber or customer) device may include a computing device capable of sending or receiving signals, such as via a wiredor a wireless network. A client device may, for example, include a desktop computer or a portable device, such as a cellular telephone, a smart phone, a display pager, a radio frequency (RF) device, an infrared (IR) device a Near Field Communication (NFC) device, a Personal Digital Assistant (PDA), a handheld computer, a tablet computer, a phablet, a laptop computer, a set top box, a wearable computer, smart watch, an integrated or distributed device combining various features, such as features of the forgoing devices, or the like.
[0038] A client device may vary in terms of capabilities or features. Claimed subject matter is intended to cover a wide range of potential variations, such as a web-enabled client device or previously mentioned devices may include a high-resolution screen (HD or 4K for example), one or more physical or virtual keyboards, mass storage, one or more accelerometers, one or more gyroscopes, global positioning system (GPS) or other location-identifying type capability, or a display with a high degree of functionality, such as a touch-sensitive color 2D or 3D display, for example.
[0039] Certain embodiments and principles will be discussed in more detail with reference to the figures. With reference to FIG. 1. system 100 is depicted which includes user equipment (UE) 102 (e g., a client device, as mentioned above and discussed below in relation to FIG. 8), network 104, cloud system 106, database 108, sensors 110 and capture engine 200. It should be understood that while system 100 is depicted as including such components, it should not be construed as limiting, as one of ordinary skill in the art would readily understand that varying numbers of UEs, peripheral devices, sensors, cloud systems, databases and networks can be utilized; however, for purposes of explanation, system 100 is discussed in relation to the example depiction in FIG. 1.
[0040] According to some embodiments, UE 102 can be any type of device, such as, but not limited to, a mobile phone, tablet, laptop, sensor, smart television (TV) Internet of Things (loT) device, autonomous machine, wearable device, and / or any other device equipped with a cellular or wireless or wired transceiver. For example, UE 102 can be a security' camera and / or security’ control panel.
[0041] In some embodiments, a peripheral device (not shown) can be connected to UE 102, and can be any type of peripheral device, such as, but not limited to, a wearable device (e.g., smart ring or smart watch), printer, speaker, sensor, and the like. In some embodiments, a peripheral device can be any type of device that is connectable to UE 102 via any ty pe of known or to be known pairing mechanism, including, but not limited to, WiFi, Bluetooth™, Bluetooth Low Energy (BLE), NFC, and the like.
[0042] According to some embodiments, sensors 110 (or sensor devices 110) can correspond to any type of device, component and / or sensor associated with a location of system 100 (referred to, collectively, as ‘’sensors”). In some embodiments, the sensors 110 can be any type of device that is capable of sensing and capturing data / metadata related to a user and / or activity of the location. For example, the sensors 110 can include, but not be limited to, cameras, motion detectors, door and window contacts, temperature, heat and smoke detectors, passive infrared (PIR) sensors, time-of-flight (ToF) sensors, and the like. In some embodiments, the sensors 110 can be associated with devices associated with the location of system 100, such as, for example, lights, smart locks, garage doors, smart appliances (e.g., thermostat, refrigerator, television, personal assistants (e.g., Alexa®, Nest®, for example)), smart rings, smart phones, smart watches or other wearables, tablets, personal computers, and the like, and some combination thereof.
[0043] By way of a non-limiting example, as discussed herein, the sensors 110 can correspond to a doorbell camera affixed to the frame of a front door of a user’s home. In another example, sensors 110 can be an high definition (HD) outdoor camera affixed to a position on a user’s home (e.g., above the garage, for example).
[0044] In some embodiments, network 104 can be any type of network, such as, but not limited to, a wireless network, cellular network, the Internet, and the like (as discussed above). Network 104 facilitates connectivity of the components of system 100, as illustrated in FIG. 1.
[0045] According to some embodiments, cloud system 106 may be any type of cloud operating platform and / or network based system upon which applications, operations, and / or other forms of network resources may be located. For example, system 106 may be a sendee provider and / or network provider from where services and / or applications may be accessed, sourced or executed from. For example, system 106 can represent the cloud-based architecture associated with a location monitoring and control system provider (e.g., security system provided by Resideo®), which has associated network resources hosted on the internet or private network (e.g., network 104), which enables (via engine 200) the location management discussed herein.
[0046] In some embodiments, cloud system 106 may include a server(s) and / or a database of information which is accessible over network 104. In some embodiments, a database 108 of cloud system 106 may store a dataset of data and metadata associated with local and / or network information related to a user(s) of the components of system 100 and / or each of the components of system 100 (e.g., UE 102. sensors 110, and the services and applications provided by cloud system 106 and / or capture engine 200).
[0047] In some embodiments, for example, cloud system 106 can provide a private / proprietary management platform, whereby engine 200, discussed infra, corresponds to the novel functionality system 106 enables, hosts and provides to a network 104 and other devices / platforms operating thereon.
[0048] Turning to FIG. 6 and FIG. 7, in some embodiments, the exemplary computer-based systems / platforms. the exemplary computer-based devices, and / or the exemplary computer- based components of the present disclosure may be specifically configured to operate in a cloud computing / architecture 106 such as, but not limiting to: infrastructure as a sendee (laaS) 710, platform as a service (PaaS) 708, and / or software as a service (SaaS) 706 using a web browser, mobile app, thin client, terminal emulator or other endpoint 704. FIG. 6 and FIG. 7 illustrate schematics of non-limiting implementations of the cloud computing / architecture(s) in which the exemplary computer-based systems for administrative customizations and control of network-hosted application program interfaces (APIs) of the present disclosure may be specifically configured to operate.
[0049] Turning back to FIG. 1, according to some embodiments, database 108 may correspond to a data storage for a platform (e.g., a network hosted platform, such as cloud system 106, as discussed supra) or a plurality of platforms. Database 108 may receive storage instructions / requests from, for example, engine 200 (and associated microservices), which maybe in any type of known or to be known format, such as. for example, standard query language (SQL). According to some embodiments, database 108 may correspond to any type of known or to be known storage, for example, a memory or memory stack of a device, a distributed ledger of a distributed network (e.g., blockchain, for example), a look-up table (LUT), and / or any other ty pe of secure data repository.
[0050] Capture engine 200, as discussed above and further below in more detail, can include components for the disclosed functionality. According to some embodiments, capture engine 200 may be a special purpose machine or processor, and can be hosted by a device on network 104, within cloud system 106 and / or on UE 102. In some embodiments, engine 200 may be hosted by a server and / or set of servers associated with cloud system 106.
[0051] According to some embodiments, as discussed in more detail below, capture engine 200 may be configured to implement and / or control a plurality of services and / or microsen-ices, where each of the plurality of services / microservices are configured to execute a plurality of workflows associated with performing the disclosed location management. Non-limiting embodiments of such workflows are provided below in relation to at least FIG. 3. FIG. 4 and FIG. 5.
[0052] According to some embodiments, as discussed above, capture engine 200 may function as an application provided by cloud system 106. In some embodiments, engine 200 may function as an application installed on a server(s), network location and / or other type of network resource associated with system 106. In some embodiments, engine 200 may function as an application installed and / or executing on UE 102 and / or sensors 110. In some embodiments, such application may be a web-based application accessed by UE 102 and / or devices associated with sensors 110 over network 104 from cloud system 106. In some embodiments, engine 200 may be configured and / or installed as an augmenting script, program or application (e.g., a plug-in or extension) to another application or program provided by cloud system 106 and / or executing on UE 102 and / or sensors 110.
[0053] As illustrated in FIG. 2, according to some embodiments, capture engine 200 includes identification module 202, analysis module 204, determination module 206 and output module 208. It should be understood that the engine(s) and modules discussed herein are non- exhaustive, as additional or fewer engines and / or modules (or sub-modules) may be applicable to the embodiments of the systems and methods discussed. More detail of the operations, configurations and functionalities of engine 200 and each of its modules, and their role within embodiments of the present disclosure will be discussed below.
[0054] Turning to FIG. 3, Process 300 provides non-limiting example embodiments for enabling the automatic and dynamic determination and application of fields of view for captured security footage via a security camera (e.g., UE 102 and / or sensors 110, discussed supra). For example, as discussed herein, if an event is detected to the ‘‘bottom” of the field of view of a camera range of capturing content, the framework can select “tall” to ensure that the camera adequately captures the event.
[0055] According to some embodiments, Steps 302-306 of Process 300 can be performed by identification module 202 of capture engine 200; Steps 308, 310 and 316-318 can be performed by output module 208; Steps 312 can be performed by analysis module 204; and Step 314 can be determined by determination module 206.
[0056] According to some embodiments, Process 300 begins with Step 302 where engine 200, via sensors 110 of a location (as in FIG. 1, discussed supra) can monitor a location via a security system configured for the location. For example, a home security system at a house, where the house has affixed thereon a set of cameras around positions on the home (e.g., door, garage.backyard, kid’s bedroom, and the like). According to some embodiments, such monitoring can be performed periodically (e.g., according to a predetermined time period or interval (e.g.. every 30 seconds, for example) and / or continuously.
[0057] In Step 304, engine 200 can detect an event. For example, an object (e.g., a person or a car, for example), moves into the field of view of a security' camera at the location. The object’s detection can trigger the capture of footage (e.g., a video clip and / or series of video frames (or image frames)). Thus, in Step 306, engine 200 can cause the activation of a camera (or cameras) to capture the event for a time (e.g., which can be for an entirely of which movement is tracked within the field of view, for a predetermined period of time (e.g., n seconds), and / or any other criteria for which a camera is automated to capture footage upon an object’s movement being detected within its field of view. In some embodiments, such capture can occur at an initial field of view (e.g., default) field of view, which can be standard according to the camera’s settings, and / or preset (e.g., a rectangle and / or square, for example). Accordingly, in Step 308, event activity via the triggered camera(s) can be captured, and in Step 310, stored in database 108.
[0058] Upon detection of the event (as in Step 304) and / or upon commencement of the image capture (in Step 308), engine 200 can analyze the event activity, as in Step 312. According to some embodiments, engine 200 can implement any type of known or to be known computational analysis technique, algorithm, mechanism or technology' to perform such analysis.
[0059] In some embodiments, engine 200 may execute and / or include a specific trained artificial intelligence / machine learning model (AI / ML), a particular machine learning model architecture, a particular machine learning model type (e.g., convolutional neural network (CNN), recurrent neural network (RNN), autoencoder, support vector machine (SVM), and the like), or any other suitable definition of a machine learning model or any suitable combination thereof.
[0060] In some embodiments, engine 200 may be configured to utilize one or more AI / ML techniques chosen from, but not limited to, computer vision, feature vector analysis, decision trees, boosting, support-vector machines, neural networks, nearest neighbor algorithms, Naive Bayes, bagging, random forests, logistic regression, and the like. By way of a non-limiting example, engine 200 can implement an XGBoost algorithm for regression and / or classification to analyze the sensor data, as discussed herein.
[0061] In some embodiments and, optionally, in combination of any embodiment described above or below, a neural network technique may be one of. without limitation, feedforwardneural network, radial basis function network, recurrent neural network, convolutional netw ork (e.g.. U-net) or other suitable network. In some embodiments and. optionally, in combination of any embodiment described above or below; an implementation of Neural Netw ork may be executed as follows: a. define N eural N et work architecture / model , b. transfer the input data to the neural network model, c. train the model incrementally, d. determine the accuracy for a specific number of timesteps, e. apply the trained model to process the newly -received input data, f. optionally and in parallel, continue to train the trained model with a predetermined periodicity.
[0062] In some embodiments and, optionally, in combination of any embodiment described above or below; the trained neural network model may specify a neural network by at least a neural network topology, a series of activation functions, and connection weights. For example, the topology of a neural network may include a configuration of nodes of the neural network and connections between such nodes. In some embodiments and, optionally, in combination of any embodiment described above or below, the trained neural network model may also be specified to include other parameters, including but not limited to, bias values / functions and / or aggregation functions. For example, an activation function of a node may be a step function, sine function, continuous or piecewise linear function, sigmoid function, hyperbolic tangent function, or other type of mathematical function that represents a threshold at which the node is activated. In some embodiments and, optionally, in combination of any embodiment described above or below; the aggregation function may be a mathematical function that combines (e.g.. sum, product, and the like) input signals to the node. In some embodiments and. optionally, in combination of any embodiment described above or below, an output of the aggregation function may be used as input to the activation function. In some embodiments and, optionally, in combination of any embodiment described above or below; the bias may be a constant value or function that may be used by the aggregation function and / or the activation function to make the node more or less likely to be activated.
[0063] In Step 314, based on the analysis from Step 312, engine 200 can determine a proper field of view that corresponds to the context of the event activity. Such context correlates to where within the field of view the event activity is occurring. For example, as discussed above.if an event is detected to the “top’' of the field of view of a camera range of capturing content, the framework can select “tall” to ensure that the camera adequately captures the event.
[0064] Thus, in Step 316, engine 200 can cause the field of view to automatically modify based on the dynamically determined context of the event that is occurring. In some embodiments, as indicated via the line from (Step 316 to Step 302), such recursive actions by engine 200 can cause the field of view again be modified even for the same event (e.g., an object has moved from the “top” of the field of view to the left edge, therefore, via the recursive processing of Process 300, the field of view can change from “tall” to “wide” event for the same tracked event.
[0065] And, in Step 318, the capture of the event activity can continue until the object has exited the field of view, for which the capture can conclude, and the captured footage can be stored. In some embodiments, as mentioned above, capture can continue for n seconds after the object exited the field of view.
[0066] Turning to FIG. 4, Process 400 provides non-limiting example embodiments for the creation of an aggregate (or composite) event video file from a collection of video clips that include corresponding to a related event.
[0067] According to some embodiments, Steps 402-404 can be performed by identification module 202 of capture engine 200; Steps 406 and 410 can be performed by analysis module 204; Steps 408 and 412 can be performed by determination module 206: and Steps 414-418 can be performed by output module 208.
[0068] According to some embodiments, Process 400 begins with Step 402 where a request related to an object is received. For example, a search request can be received, whereby the search query identifies a particular object (e.g., a human), and / or a time, date and / or position around a location. For example: “identify activity on Jan. 1. 2024 that occurred on my driveway.”
[0069] In Step 404, a repository of stored events for a location (e.g., stored video clips) is accessed, and in Step 406, engine 200 can analyze the stored events. According to some embodiments, such analysis can be based on information related to the request, which can correspond to an object and / or criteria (e.g., time, date, location, as discussed supra in at least Step 402). According to some embodiments, such search analysis in Step 406 can involve parsing and / or mining the data / metadata for each stored video clip, and determining which clips correspond, at least to a threshold similarity value, to the request. In some embodiments, suchsearch analysis can (either alternatively and / or in combination) involve utilizing any of the AI / ML techniques discussed above.
[0070] In Step 408, based on the analysis of Step 406, engine 200 can determine a set of video clips that contain the digital representation of the object. Such video clips further comprise content and / or where captured in accordance with the criteria.
[0071] In Step 410, engine 200 can then analyze the identified video clips (from Step 408) according to a time sequence. That is, engine 200 can parse, mine and / or perform an AI / ML- based analysis of the videos (in a similar manner as discussed above) to determine a sequence among the events. For example, for a tracked event, in which order were the events captured and recorded. In some embodiments, such sequence can be determined based on timestamp data / metadata related to each captured event. In some embodiments, for example, computer vision techniques can be utilized to analyze the video clips to determine, in relation to a layout of the location and / or cameras used to capture the clips, how the object moved about / around the location.
[0072] For example, as discussed above, a stranger walked on the user's property, and such movement was captured by three (3) cameras at the user’s house. Currently, there is no way to computationally determine the sequence of the stranger’s movement (e.g., first walked up the driveway, then around the back to the shed, then to the side window, and the like).
[0073] In Step 414. engine 200 can function to “stitch” together the video clips in the determined sequence. Such stitching can involve the creation of a new video file, which is an aggregation of the activity performed by the object according to the criteria. In some embodiments, such new file creation can involve extracting the content from each of the video clips, compiling them in an organized manner according to the sequence, and then generating the file from such compilation. In Step 416, such new video file can be stored in database 108.
[0074] And, in Step 418, engine 200 can output, in response to the request (from Step 402), the new video file for rendering and / or displaying within a UI. In some embodiments, such video file can be output for sharing, downloading, uploading, and the like.
[0075] Turning to FIG. 5. Process 500 provides non-limiting example embodiments for the creation of an interactive display of tracked events within a UI, for which a timelme(s) can be curated and displayed based on tracked events, which can be related to footage captured from a camera and / or a set of cameras at a location.
[0076] According to some embodiments, Steps 502 and 516 can be performed by identification module 202 of capture engine 200; Steps 504 and 508 can be performed by analysis module204; Steps 506 and 510 can be performed by determination module 206; and Steps 512, 514 and 518 can be performed by output module 208.
[0077] According to some embodiments, Process 500 begins with Step 502 where engine 200 can access a repository of stored clips for a location. In some embodiments, such access can be based on a request, in a similar manner as discussed in relation to Steps 402-404, discussed supra. In some embodiments, such access can be based on the creation of a new set of video clips, which can be in accordance with a predetermined number, a specific location, a periodic request, and the like, or some combination thereof. In some embodiments, the processing of Process 500 can be performed for each set of related video clips in the repository', and / or for those that are requested for creation of a timeline, as discussed herein.
[0078] In Step 504, engine 200 can perform a search for video clips captured from a set of cameras. For example, in a similar manner as discussed above in at least in Step 404-406, a search query' can be defined by, but not limited to, an object, type of object, time period, location, position at the location, type of camera, position of camera, and the like, or some combination thereof.
[0079] Thus, for example, a search for camera footage from the outdoor camera in the backyard can be executed for a time period of 12 AM on Jan. 1, 2024 to 12 PM on Jan. 3, 2024.
[0080] Accordingly, such search can be performed in a similar manner as discussed above at least in relation to Step 406 (e g., engine 200 can parse, mine and / or perform an AI / ML-based analysis of the videos in the repository / database 108). As a result of the search, a set of videos can be identified.
[0081] In Step 506, engine 200 can group, based on the results of the search, group (or cluster) the video clips. For example, clips for camera 1 according to the time frame of the search query can be grouped, clips for camera 2 according to the time frame can be grouped, and the like.
[0082] In Step 508, engine 200 can analyze each grouping of video clips. According to some embodiments, such analysis can involve any of the AI / ML techniques discussed above.
[0083] In Step 510, as a result of the analysis of each of the videos, per grouping, engine 200 can determine a timeline of the video clips for each grouping. Such timeline can be determined according a time sequence, which can be determined via similar functionality discussed in relation to at least Steps 410-412, discussed supra.
[0084] In Step 512, engine 200 can compile and generate a timeline for each grouping, whereby the timeline can be an interactive set of interface objects (IOs), where each IO (or node) on the timeline corresponds to a specific video clip. Each IO can be interactive and displayed as athumbnail, such that the 10 is linked to the video clips location, such that interaction with the 10 (e.g., selecting) can cause the corresponding video clip to be displayed within a pop-up display for rendering. As discussed above, in some embodiments, a grouping’s timeline can be displayed vertically, horizontally, and / or in any manner that enables the viewer to understand the order of events captured by each video clip along the timeline. In Step 514, the information related to the timeline, inclusive of what is included in the timeline, as well as how it is configured for display, can be stored in database 108. Such storage can be based on the creation of a timeline data structure, whereby Step 512 involves the creation of the rimeline data structure (or executable file), and such data structure is then stored in Step 514.
[0085] In Step 516, engine 200 can receive a request for information related to the location. For example, a request for events occurring from Jan. 1 to Jan. 3. Accordingly, in Step 518, engine 200 can cause display of the corresponding timeline in a UI in response to the request.
[0086] In some embodiments, if a request for a timeline corresponds to an existing timeline, but such timeline is either incomplete or provides more detail than needed (e.g., for less time or more time than the timeline provides), then the timeline can be accessed, and modified. For example, a request is for events from Jan. 1 to Jan. 2. In some embodiments, engine 200 can execute the steps of Process 500 to perform such timeline creation and rendering. In some embodiments, an existing timeline (e.g., for Jan. 1 to Jan. 3) can be retrieved, and the data related to Jan. 2 to Jan. 3 can be trimmed from the existing timeline. In some embodiments, such trimming can involve creating a new timeline and / or modifying an existing timeline to remove the unrequested information. In some embodiments, data from an existing timeline can be extracted that corresponds to a request, such that rather than filtering or removing (e.g., trimming) information, information can be extracted for the creation of a new timeline data structure.
[0087] FIG. 8 is a schematic diagram illustrating a client device showing an example embodiment of a client device that may be used within the present disclosure. Client device 800 may include many more or less components than those shown in FIG. 8. However, the components shown are sufficient to disclose an illustrative embodiment for implementing the present disclosure. Client device 800 may represent, for example, UE 102 discussed above at least in relation to FIG. 1.
[0088] As shown in the figure, in some embodiments, Client device 800 includes a processing unit (CPU) 822 in communication with a mass memory’ 830 via a bus 824. Client device 800 also includes a power supply 826, one or more network interfaces 850. an audio interface 852.a display 854, a keypad 856, an illuminator 858, an input / output interface 860, ahaptic interface 862, an optional global positioning systems (GPS) receiver 864 and a camera(s) or other optical, thermal or electromagnetic sensors 866. Device 800 can include one camera / sensor 866, or a plurality of cameras / sensors 866, as understood by those of skill in the art. Power supply 826 provides power to Client device 800.
[0089] Client device 800 may optionally communicate with a base station (not shown), or directly with another computing device. In some embodiments, network interface 850 is sometimes know n as a transceiver, transceiving device, or netw ork interface card (NIC).
[0090] Audio interface 852 is arranged to produce and receive audio signals such as the sound of a human voice in some embodiments. Display 854 may be a liquid crystal display (LCD), gas plasma, light emitting diode (LED), or any other type of display used with a computing device. Display 854 may also include a touch sensitive screen arranged to receive input from an object such as a stylus or a digit from a human hand.
[0091] Keypad 856 may include any input device arranged to receive input from a user. Illuminator 858 may provide a status indication and / or provide light.
[0092] Client device 800 also includes input / output interface 860 for communicating with external. Input / output interface 860 can utilize one or more communication technologies, such as USB, infrared, Bluetooth™, or the like in some embodiments. Haptic interface 862 is arranged to provide tactile feedback to a user of the client device.
[0093] Optional GPS transceiver 864 can determine the physical coordinates of Client device 800 on the surface of the Earth, which typically outputs a location as latitude and longitude values. GPS transceiver 864 can also employ other geo-positioning mechanisms, including, but not limited to, triangulation, assisted GPS (AGPS), E-OTD. CL SAI, ETA, BSS or the like, to further determine the physical location of client device 800 on the surface of the Earth. In one embodiment, however. Client device 800 may through other components, provide other information that may be employed to determine a physical location of the device, including for example, a MAC address, Internet Protocol (IP) address, or the like.
[0094] Mass memory 830 includes a RAM 832, a ROM 834, and other storage means. Mass memory 830 illustrates another example of computer storage media for storage of information such as computer readable instructions, data structures, program modules or other data. Mass memory 830 stores a basic input / output system (“BIOS”) 840 for controlling low-level operation of Client device 800. The mass memory also stores an operating system 841 for controlling the operation of Client device 800.
[0095] Memory 830 further includes one or more data stores, which can be utilized by Client device 800 to store, among other things, applications 842 and / or other information or data. For example, data stores may be employed to store information that describes various capabilities of Client device 800. The information may then be provided to another device based on any of a variety of events, including being sent as part of a header (e.g., index file of the HLS stream) during a communication, sent upon request, or the like. At least a portion of the capability information may also be stored on a disk drive or other storage medium (not shown) within Client device 800.
[0096] Applications 842 may include computer executable instructions which, when executed by Client device 800, transmit, receive, and / or otherwise process audio, video, images, and enable telecommunication with a server and / or another user of another client device. Applications 842 may further include a client that is configured to send, to receive, and / or to otherwise process gaming, goods / services and / or other forms of data, messages and content hosted and provided by the platform associated with engine 200 and its affiliates.
[0097] As used herein, the terms “computer engine” and “engine” identify at least one software component and / or a combination of at least one software component and at least one hardware component which are designed / programmed / configured to manage / control other software and / or hardware components (such as the libraries, software development kits (SDKs), objects, and the like).
[0098] Examples of hardware elements may include processors, microprocessors, circuits, circuit elements (e g., transistors, resistors, capacitors, inductors, and so forth), integrated circuits, application specific integrated circuits (ASIC), programmable logic devices (PLD), digital signal processors (DSP), field programmable gate array (FPGA), logic gates, registers, semiconductor device, chips, microchips, chip sets, and so forth. In some embodiments, the one or more processors may be implemented as a Complex Instruction Set Computer (CISC) or Reduced Instruction Set Computer (RISC) processors; x86 instruction set compatible processors, multi-core, or any other microprocessor or central processing unit (CPU). In various implementations, the one or more processors may be dual-core processor(s), dual-core mobile processor(s), and so forth.
[0099] Computer-related systems, computer systems, and systems, as used herein, include any combination of hardware and software. Examples of software may include software components, programs, applications, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces.application program interfaces (API), instruction sets, computer code, computer code segments, words, values, symbols, or any combination thereof. Determining whether an embodiment is implemented using hardware elements and / or software elements may vary in accordance w ith any number of factors, such as desired computational rate, power levels, heat tolerances, processing cycle budget, input data rates, output data rates, memory resources, data bus speeds and other design or performance constraints.
[0100] For the purposes of this disclosure a module is a software, hardware, or firmware (or combinations thereof) system, process or functionality , or component thereof, that performs or facilitates the processes, features, and / or functions described herein (with or without human interaction or augmentation). A module can include sub-modules. Software components of a module may be stored on a computer readable medium for execution by a processor. Modules may be integral to one or more servers, or be loaded and executed by one or more servers. One or more modules may be grouped into an engine or an application.
[0101] One or more aspects of at least one embodiment may be implemented by representative instructions stored on a machine-readable medium which represents various logic within the processor, which when read by a machine causes the machine to fabricate logic to perform the techniques described herein. Such representations, known as “IP cores,” may be stored on a tangible, machine readable medium and supplied to various customers or manufacturing facilities to load into the fabrication machines that make the logic or processor. Of note, various embodiments described herein may, of course, be implemented using any appropriate hardw are and / or computing software languages (e.g., C++, Objective-C, Swift, Java, JavaScript, Python, Perl, QT, and the like).
[0102] For example, exemplary software specifically programmed in accordance with one or more principles of the present disclosure may be downloadable from a network, for example, a website, as a stand-alone product or as an add-in package for installation in an existing software application. For example, exemplary softw are specifically programmed in accordance with one or more principles of the present disclosure may also be available as a client-server software application, or as a web-enabled software application. For example, exemplary softw are specifically programmed in accordance w ith one or more principles of the present disclosure may also be embodied as a software package installed on a hardware device.
[0103] For the purposes of this disclosure the term “user”, “subscriber” “consumer” or “customer” should be understood to refer to a user of an application or applications as described herein and / or a consumer of data supplied by a data provider. By way of example, and notlimitation, the term "user" or ‘‘subscriber’' can refer to a person who receives data provided by the data or service provider over the Internet in a browser session, or can refer to an automated software application which receives the data and stores or processes the data. Those skilled in the art will recognize that the methods and systems of the present disclosure may be implemented in many manners and as such are not to be limited by the foregoing exemplary embodiments and examples. In other words, functional elements being performed by single or multiple components, in various combinations of hardware and software or firmware, and individual functions, may be distributed among software applications at either the client level or server level or both. In this regard, any number of the features of the different embodiments described herein may be combined into single or multiple embodiments, and alternate embodiments having fewer than, or more than, all of the features described herein are possible.
[0104] Functionality may also be, in whole or in part, distributed among multiple components, in manners now known or to become known. Thus, myriad software / hardware / firmware combinations are possible in achieving the functions, features, interfaces and preferences described herein. Moreover, the scope of the present disclosure covers conventionally known manners for carrying out the described features and functions and interfaces, as well as those variations and modifications that may be made to the hardware or software or firmware components described herein as would be understood by those skilled in the art now' and hereafter.
[0105] Furthermore, the embodiments of methods presented and described as flow charts in this disclosure are provided by way of example in order to provide a more complete understanding of the technology7. The disclosed methods are not limited to the operations and logical flow presented herein. Alternative embodiments are contemplated in which the order of the various operations is altered and in which sub-operations described as being part of a larger operation are performed independently.
[0106] While various embodiments have been described for purposes of this disclosure, such embodiments should not be deemed to limit the teaching of this disclosure to those embodiments. Various changes and modifications may be made to the elements and operations described above to obtain a result that remains within the scope of the systems and processes described in this disclosure.
Claims
CLAIMSWhat is claimed is:
1. A method comprising: monitoring, by a device, a location, the device being associated with a security system associated with the location; detecting, based on the monitoring, an event, the event detection causing a capture of a video clip by the device, the video clip being captured via an initial field of view format and comprising digital content corresponding to event activity related to the detected event; analyzing the video clip, and determining, based on the analysis, a context of the event activity; modifying the field of view from the initial field of view format to another type of field of view format based on the determined context; and further capturing the video clip via the device in the other type of field of view format.
2. The method of claim 1. wherein the modification of the field of view format to the other type of field of view format occurs at commencement of the capture of the video clip.
3. The method of claim 1, wherein the context of the event activity corresponds to a position within the initial field of view format for which the event was detected, the event corresponding to movement within the initial field of view format of an object.
4. The method of claim 1, further comprising: analyzing the further captured video clip; determining to further modify the other field of view format; and modifying the other field of view format, such that additional capture of the video clip is within the modified other view of view format.
5. The method of claim 1 , wherein the initial and other type of field of view format are selected from a group consisting of a tall format, full format and wide format.
6. The method of claim 1, further comprising:storing a video file based on the captured and further captured video clip in a database associated with the security system.
7. The method of claim 6, further comprising: identifying, from a database associated with the security system, a set of videos, the set of videos comprising the stored video file and other video files that are related to the stored video file.
8. The method of claim 7, further comprising: analyzing each of the set of videos; determining a time sequence among the set of videos; and generating a new video file, the new video file being a composite file composed of the set of video files ordered according to the time sequence.
9. The method of claim 7. further comprising: analyzing each of the set of videos; determining a time sequence among the set of videos; and generating a timeline for the set of videos, wherein the timeline is an interactive data structure that includes each video as a node on the timeline, such that interaction with a respective node, enables rendering of a respective video file from the set of videos.
10. The method of claim 1, wherein the device is a camera associated with the security system.
11. A method comprising: receiving, over a network, a request related to an object; searching a repository' of stored events for a location, the stored events being video clips captured by at least one camera of a security system; determining, based on the search, a set of video clips, the set of video clips comprising content corresponding to the content and satisfying a criteria; analyzing the set of video clips, and determining an order among the set of video clips that adheres to a time sequence;creating, based on the set of video clips and the determined order according to the time sequence, a new video file, the new video file being an ordered composite of the set of video clips; and outputting, in response to the request, the new video file for display and rendering within a user interface (UI).
12. The method of claim 1 1, wherein the criteria is associated with at least one of time, location, position, type of object, object identifier, type of camera and camera identifier.
13. A device comprising: a processor configured to: monitor a location, the device being associated with a security system associated with the location; detect, based on the monitoring, an event, the event detection causing a capture of a video clip by the device, the video clip being captured via an initial field of view format and comprising digital content corresponding to event activity related to the detected event; analyze the video clip, and determine, based on the analysis, a context of the event activity’, the context of the event activity corresponds to a position within the initial field of view format for which the event was detected, the event corresponding to movement within the initial field of view format of an object; modify the field of view from the initial field of view format to another type of field of view format based on the determined context; and further capture the video clip via the device in the other type of field of view format.
14. The device of claim 13, wherein the modification of the field of view format to the other type of field of view format occurs at commencement of the capture of the video clip.
15. The device of claim 13, wherein the processor is further configured to: analyze the further captured video clip; determine to further modify the other field of view format; andmodify the other field of view format, such that additional capture of the video clip is within the modified other view of view format.
16. The device of claim 13, wherein the initial and other type of field of view7format are selected from a group consisting of a tall format, full format and wide format.
17. The device of claim 13, wherein the processor is further configured to: store a video file based on the captured and further captured video clip in a database associated with the security system; and identify, from a database associated with the security system, a set of videos, the set of videos comprising the stored video file and other video files that are related to the stored video file.
18. The device of claim 17, wherein the processor is further configured to: analyze each of the set of videos; determine a time sequence among the set of videos; and generate a new7video file, the new video file being a composite file composed of the set of video files ordered according to the time sequence.
19. The device of claim 17, further comprising: analyze each of the set of videos; determine a time sequence among the set of videos; and generate a timeline for the set of videos, wherein the timeline is an interactive data structure that includes each video as a node on the timeline, such that interaction with a respective node, enables rendering of a respective video file from the set of videos.
20. The device of claim 13, wherein the device is a camera associated with the security system.
Citation Information
Patent Citations
Method, apparatus and system for adjusting field of view of observation, and storage medium and mobile apparatus
EP3817374A1
Image Processing System for Extending a Range for Image Analytics
US20230098829A1