User interface for security events
By integrating video and detection UI element display into the security system application, the problem of users having difficulty understanding the time relationship between video and detection UI elements is solved, improving user experience and navigation efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2026-03-27
AI Technical Summary
In existing security systems, the lack of integration between recorded video and UI elements representing detected features makes it difficult for users to understand the temporal relationship between the video and the detected UI elements, resulting in a poor user experience.
The application displays video and detects UI elements in different areas of the screen, and highlights the corresponding detected UI elements when the video is played back to a specific time position, thus integrating video and UI elements and allowing users to easily understand the time relationship.
It improves users' ability to understand events in recorded videos, enhances user experience and usability, and makes it easier for users to quickly navigate to relevant video sections.
Smart Images

Figure CN121753319A_ABST
Abstract
Description
Background Technology
[0001] Some security systems are able to use cameras and other devices to remotely monitor locations. Summary of the Invention
[0002] In some aspects, the technology described herein relates to a method comprising: providing a user interface via an application of a computing device to (i) play back a video in a first area of a screen and (ii) display a plurality of interactive elements corresponding to features detected in the video, the plurality of interactive elements being displayed in a second area of the screen different from the first area; determining, via the application, that the playback of the video has reached a first time position in the video corresponding to a first interactive element among the plurality of interactive elements displayed in the second area; and causing, via the application, a change in the appearance of the first interactive element to visually distinguish the first interactive element from the other interactive elements among the plurality of interactive elements, the change being temporary such that, as the playback of the video progresses beyond the first time position, the appearance of the first interactive element reverts to the appearance it was displayed before reaching the first time position in the video.
[0003] In some aspects, the technology described herein relates to a method comprising: receiving, via an application, first data representing a video of an event detected by a camera, second data representing at least a first feature detected in the video, and third data indicating a first time position in the video where the first feature was detected; via the application and using the first data, causing a device to play back at least a portion of the video in a first area of a screen; via the application and using the second data, causing the device to display a first user interface (UI) element indicating the first feature in a second area of the screen; via the application, determining that the playback of the video has reached the first time position; and via the application and at least in part based on the third data and that the playback of the video has reached the first time position, causing a change in the appearance of the first UI element to visually distinguish the first UI element from at least a second UI element displayed on the screen, the second UI element indicating the second feature detected in the video. Attached Figure Description
[0004] Additional examples, features, and advantages of this disclosure will become more apparent from the description herein taken in conjunction with the accompanying drawings, which are incorporated in and constitute a part of this disclosure. The drawings are not necessarily drawn to scale.
[0005] Figure 1A A first exemplary screen, which can be presented on a user device to display event logs generated by a security system, is shown according to some embodiments of the present disclosure.
[0006] Figure 1BA second exemplary screen, which can be presented on a user device to display event logs generated by a security system, is shown according to some embodiments of the present disclosure.
[0007] Figure 1C A third exemplary screen, which can be presented on a user device to display event logs generated by a security system, is shown according to some embodiments of the present disclosure.
[0008] Figure 2 Exemplary components of a security system configured according to some embodiments of this disclosure are shown, as well as exemplary interactions or data flows that may occur between such components.
[0009] Figure 3 Exemplary routines that can be executed by an application hosted on a user device, according to some embodiments of this disclosure, are shown.
[0010] Figure 4A and Figure 4B Exemplary tables or data structures, according to some embodiments of this disclosure, can be used to store records and information of various events detected by a security system.
[0011] Figure 5 This illustrates some embodiments of the invention that can be implemented by [the party that made the invention]. Figure 2 The flowchart illustrates an exemplary process employed by the remote image processing component.
[0012] Figure 6 This demonstrates how it can be achieved through an application. Figures 1A to 1C The screen shown contains exemplary routines for some of its functions.
[0013] Figure 7 This demonstrates how, according to some embodiments of the present disclosure, a list of features detected in a recorded video can be merged and sorted.
[0014] Figure 8 This is a schematic diagram of an exemplary security system that can be adopted by various aspects of this disclosure.
[0015] Figure 9 This is a schematic diagram of a computing device that can be used to implement one or more of the services of a client device, a monitoring device, and / or the security system disclosed herein, according to some embodiments of this disclosure. Detailed Implementation
[0016] To facilitate understanding of the principles of this disclosure, reference will now be made to the examples illustrated in the accompanying drawings, and these examples will be described using specific language. However, it should be understood that this is not intended to limit the scope of the examples described herein.
[0017] Some security systems offer applications (e.g., mobile applications) that allow their customers to review recorded videos of events captured by cameras monitoring their property. For example, a user can receive notifications via the application that a new event has been detected and can select one or more user interface (UI) elements to launch a video player to view the recorded video of the detected event. Some such systems can also perform computer vision (CV) processing on the recorded video to detect specific features within frames of the video, such as motion, people, faces, etc., and can present a timeline of such detections via the application, allowing the user to scroll through a list of UI elements representing the detected features. These UI elements are referred to herein as “detection UI elements.” While both the video player and the associated detection UI elements serve as independent tools providing useful information to the user, there is little integration between the two tools except for presenting them on the same screen when the user reviews the detected events. Therefore, using such systems, users may struggle to understand the relationship between the video being played back and the detection UI elements being displayed, potentially leading to a poor experience.
[0018] A security system is provided in which an application can integrate recorded video and detection UI elements for detected events in a manner that significantly improves the user experience and usability of the application. For example, in some implementations, the application can be configured to indicate (e.g., by highlighting, labeling, or otherwise marking) the corresponding detection UI element presented on the screen when the video player arrives at a frame in the video where features of the detection UI element (e.g., motion, people, faces, etc.) are detected, thereby allowing the user to easily associate the corresponding detection UI element with a specific portion of the video being played back. The user's ability to easily understand the temporal relationship between the video being played back and the detection UI elements displayed alongside such video can greatly enhance the user's ability to understand why the system records events and what the user should look for when reviewing events.
[0019] In some implementations, during playback of the recorded video, the application may further scroll the list of detected UI elements on the screen as needed, keeping the currently indicated detected UI element included within the currently displayed detected UI elements and not hidden off-screen to maintain the relevance between the detected UI element and the video being displayed, allowing the user to easily understand specific events at their property. In some implementations, such automatic scrolling can be disabled at least temporarily if the user interacts with the list of detected UI elements in at least some way, for example, by manually scrolling the list of detected UI elements. Furthermore, in some implementations, the user's selection of detected UI elements can cause the video player to jump to a frame of the recorded video that includes the detected features associated with the detected item, or to a position shortly before that frame, thereby allowing the user to quickly navigate to the relevant portion of the recorded video as the user scrolls or otherwise reviews the displayed list of detected UI elements.
[0020] Figure 1A A first exemplary screen 102, which can be presented by an application (e.g., a mobile application) according to some embodiments of the present disclosure, is shown. In some implementations, for example, screen 102 can be presented by an application 228 hosted on a user device 214 operated by a user 216, as described below. Figure 2 As described. Figure 1A As shown, screen 102 may include: a video playback window 104, in which a recorded video of a specific event being reviewed by user 216 may be displayed; and a progress bar 106, which indicates the relative time position of the currently displayed frame of the video within the recorded video clip. In some embodiments, user 216 can selectively pause or resume playback of the recorded video by clicking the video playback window 104 and / or navigate to a specific portion of the recorded video by clicking a corresponding position on the progress bar 106.
[0021] As shown, application 228 may additionally cause screen 102 to display multiple detection UI elements 108 organized chronologically, with earlier (in time) detection UI elements displayed higher on screen 102 than later (in time) detection UI elements 108. User 216 may, for example, cause application 228 to display detection UI elements 108 below video playback window 104 on screen 102 by selecting “Detect” UI element 110. As indicated, in some embodiments, UI element 110 may include a numeric identifier (e.g., “8”) indicating the number of detection UI elements 108 that can be reviewed by user 216. In some embodiments, user 216 may instead select UI element 112 to cause application 228 to display information below video playback window 104 about actions taken by monitoring agent 212 in response to an event, and / or select UI element 114 to cause application 228 to display information below video playback window 104 about other recent events also detected at monitoring location 204.
[0022] When UI element 110 has been selected (e.g.) Figure 1A As shown), and when the number of detection UI elements 108 available for review exceeds the number of detection UI elements that can be displayed in the area below the video playback window 104 on screen 102 (e.g., four detection UI elements), application 228 may allow user 216 to selectively scroll the list of detection UI elements 108 to view different groups of available detection UI elements 108. For example, in some embodiments, user 216 may drag a finger up or down on the portion of screen 102 displaying the detection UI elements 108 to scroll the list of detection UI elements 108 on screen 102 in the direction of finger movement. In some embodiments, application 228 may allow user 216 to manually scroll the list of detection UI elements 108 during video playback in video playback window 104 or when the video is paused (e.g., in response to user 216 tapping video playback window 104 or otherwise).
[0023] like Figure 1AAs shown, the detection UI elements 108 presented on screen 102 may include corresponding identifiers 116 indicating the type of detection they represent (e.g., motion, people, faces, identified faces, etc.) and corresponding time stamps 118 indicating the time of day during which video frames containing the detected features were recorded. In the illustrated example, detection UI element 108A corresponds to a video frame in which a feature (e.g., an identified face) was detected, detection UI element 108B corresponds to a video frame in which another feature (e.g., an unidentified face—a face detected but not identified as belonging to a specific individual) was detected, detection UI element 108C corresponds to a video frame in which yet another feature (e.g., a people) was detected, and detection UI element 108D corresponds to a video frame in which yet another feature (e.g., motion) was detected. Figure 1A As shown, the detection UI element 108 for identified faces (e.g., see detection UI element 108A) and unidentified faces (e.g., see detection UI element 108B) can include images of human faces. For example, such face images 120 can be obtained by cropping a region from a recorded video of the image frame corresponding to the detected face. Also, Figure 1A As shown, the detection UI element 108 for people (e.g., see detection UI element 108C) and motion (e.g., see detection UI element 108D) may include a thumbnail image 122 that shows frames of the recorded video that identify such features.
[0024] As shown directly below the video playback window 104, in some embodiments, screen 102 may also display additional information about the detected event, such as the location of the event (e.g., “Lake House”), the time and date the event was detected (e.g., “June 8, 2024, 7:45 a.m.”), the description of the event (e.g., “persons on property”), the handling of the event (e.g., “agency handling”), and a value 178 indicating the number of faces detected in the recorded video of the event.
[0025] Value 126 can, for example, notify user 216 of the presence of a detected face detection UI element 108 (e.g., detection UI elements 108A and 108B), and user 216 may want to take action, such as associating or unassociating a given face image 120 with a visitor profile by selecting UI elements 124A and 124B adjacent to the face images 120A and 120B. For example, if user 216 determines that the name indicated by the detection type identifier 116A is inaccurate, for example because of the remote image processing unit 222 (hereinafter referred to as...) Figure 2If the facial recognition processing performed (described in the text) misidentifies the person to whom facial image 120A belongs, then user 216 can select UI element 124A to unassociate facial image 120A with the visitor profile of the person indicated by the name. Additionally or alternatively, if detection type identifier 116B indicates that remote image processing unit 222 has detected an unrecognized face and user 216 determines that facial image 120B belongs to a specific person, then user 216 can select UI element 124B to associate facial image 120B with that person's visitor profile, including creating a new visitor profile for that person (if one does not already exist). (The following is in conjunction with...) Figure 2 More specifically, in some embodiments, the facial image 120 associated with the visitor profile may be used by the remote image processing unit 222 to perform facial recognition processing on video or other images acquired by a camera at monitoring location 204, and / or by the monitoring agent 212 to visually identify a specific individual when evaluating such video or other images.
[0026] Advantageously, in some embodiments, application 228 can be configured to indicate (e.g., by highlighting, labeling, or otherwise marking) the corresponding detected UI element 108 on screen 102 when the video being played back arrives at a frame of the recorded video in which features of those detected UI elements 108 (e.g., motion, people, faces, etc.) are detected, thereby allowing the user to easily associate the corresponding detected UI element 108 with a specific portion of the video being played back. Figure 1A In the exemplary screen 102 shown, for example, the detection UI element 108C is highlighted to indicate that the video being presented in the video playback window 104 has reached the frame of the detected person represented by the thumbnail image 122A. However, it should be understood that, in addition to or instead of highlighting, any of a variety of other types of identifiers may be used, such as adding circles, squares, check marks, etc., near the detection UI element 108 being indicated, or adding annotations, such as rectangles in red or other highlighted colors, to the entire detection UI element 108 or some portions of the detection UI element 108 (e.g., the face image 120 or the thumbnail image 122 being indicated).
[0027] refer to Figure 1AIn the exemplary screen 102 shown, when application 228 determines that the video being played back in video playback window 104 has reached the frame corresponding to thumbnail image 122B, application 228 can then cause detection UI element 108C to stop being highlighted or otherwise indicated, and instead cause detection UI element 108D to be highlighted or otherwise indicated, thereby notifying the user 216 viewing the recorded video in video playback window 104 that the video has reached the frame in which the feature represented by detection UI element 108D is detected. This state of detection UI elements 108C and 108D (i.e., where detection UI element 108D is highlighted or otherwise indicated and detection UI element 108C has stopped being highlighted or otherwise indicated) is reflected in Figure 1B On the exemplary screen 134 shown. It can also be noted that, Figure 1B In the middle, the position of the progress bar 106 on screen 134 indicates that the playback of the video in the video playback window 104 has progressed to more than [a certain point]. Figure 1A The time and location.
[0028] Subsequently, when application 228 determines that the video being played in video playback window 104 has reached the detected UI element (e.g., below) after the detected UI element 108D, see... Figure 1C When the frame corresponding to the detected element 180E is reached, the application 228 can then (A) move the list of detected UI elements 108 up to reveal the next detected UI element 108E in the list and (B) stop the detected UI element 108D from being highlighted or otherwise indicated and highlight or otherwise indicate the newly revealed detected UI element 108E, thereby notifying the user 216 viewing the recorded video in the video playback window 104 that the video has reached a frame in which the feature of the next detected UI element 108E in the list has been detected. It can also be noted that in Figure 1C In the middle, the position of the progress bar 106 on screen 136 indicates that the playback of the video in the video playback window 104 has progressed to or even exceeded the target. Figure 1A The time and location.
[0029] In some implementations, application 228 may be configured to prevent automatic scrolling to reveal another detected UI element 108, as described above, while video playback continues, in response to determining that user 216 has engaged in one or more specific interactions with UI elements on screen 102. For example, in response to determining that user 216 has manually scrolled the list of detected UI elements 108, e.g., by dragging a finger up or down on the list, application 228 may disable automatic scrolling for at least a short period, thereby allowing user 216 exclusive control of the scrolling function to achieve a goal, such as manually scrolling to reveal the detected UI element 108 including facial image 120 and selecting UI element 124 to associate or unassociate that facial image 120 with a visitor profile.
[0030] As another feature, such as Figure 1A As shown, in some implementations, application 228 may cause screen 102 to display UI element 128, which, when selected by user 216, enables application 228 to receive real-time or near real-time streaming video from camera 202 at monitoring location 204 and display it in video playback window 104, such as by using WebRTC functionality to establish a peer-to-peer connection between camera 202 and application 228.
[0031] Figure 2 Exemplary components of a security system 200 configured according to some embodiments of the present disclosure are illustrated, as well as exemplary interactions or data flows that may occur between such components. As shown in the figures, in addition to those described above... Figure 1A In addition to the application 228 and user device 214 that can be used to display screen 102, security system 200 may also include one or more cameras 202 located at monitoring location 204 (e.g., residence, business, parking lot, etc.) and located at locations away from camera 202 (e.g., in places such as those described below). Figure 8 The monitoring service 206 (e.g., including one or more servers 208) within the cloud-based service of the monitoring center environment 826. (See the following text in conjunction with...) Figure 8 The above refers to one or more monitoring devices 210 operated by the corresponding monitoring agent 212. The following is in conjunction with... Figure 9 The description can be used to implement an exemplary computing system 900 of any of the computer-based components disclosed herein, such as camera 202, server 208, monitoring device 210, and / or user device 214. Although Figure 2 Not shown, but it should be understood that the various components shown can communicate with each other via one or more networks (e.g., the Internet).
[0032] like Figure 2As shown, among other components, camera 202 may also include motion sensor 230, image sensor 218, and edge image processing component 220. In some implementations, camera 202 may include one or more processors and one or more computer-readable media, and the one or more computer-readable media may be encoded with instructions that, when executed by one or more processors, cause camera 202 to perform some or all of the functions of the edge image processing component 220 described herein. Furthermore... Figure 2 As shown, in addition to other components, monitoring service 206 may also include remote image processing component 222. In some embodiments, server 208 of monitoring service 206 may include one or more processors and one or more computer-readable media, and the one or more computer-readable media may be encoded with instructions that, when executed by one or more processors, cause server 208 to perform some or all of the functions of the remote image processing component 222 described herein.
[0033] like Figure 2 As shown by middle arrows 232, 234, 236, 238, 240, and 242, the edge image processing unit 220, the remote image processing unit 222, the monitoring application 226, and the application 228 can, for example, be connected via one or more networks (such as those combined below). Figure 8 The network 820 communicates with the event / video data repository 224. In some embodiments, another component within the monitoring service 206 or the monitoring center environment 826 (see [link to documentation]) communicates with the event / video data repository 224. Figure 8 It can provide one or more application programming interfaces (APIs) that can be used by the edge image processing unit 220, the remote image processing unit 222, the monitoring application 226, and the application 228 to write data to the event / video data repository 224 and / or retrieve data from the event / video data repository 224 as needed.
[0034] like Figure 2As shown, image sensor 218 can acquire image 244 (e.g., digital data representing one or more frames of acquired pixel values) from monitored location 204 and pass such image 244 to edge image processing component 220 for processing. In some implementations, for example, motion sensor 230 can detect motion at monitored location 204 and provide a signal to image sensor 218. Motion sensor 230 may be, for example, a passive infrared (PIR) sensor. In response to receiving a signal from motion sensor 230, image sensor 218 can begin acquiring frames of image 244 of the scene within the camera's field of view. In some implementations, image sensor 218 can continuously collect frames of image 244 until motion sensor 230 does not detect motion for a threshold time period (e.g., twenty seconds). Therefore, image 244 acquired by image sensor 218 may be a video clip of the scene within the camera's field of view, which begins when motion is first detected and ends after the threshold time period in which motion has stopped.
[0035] In some implementations, instead of relying on the motion sensor 230 (e.g., a PIR sensor) to trigger the collection of frames of image 244, camera 202 may instead collect frames of image 244 continuously, and the camera may rely on one or more image processors of edge image processing component 220 (e.g., machine learning (ML) models and / or other computer vision (CV) processing components) to process the collected frames to detect motion within the field of view of camera 202. Therefore, in such implementations, instead of relying on motion indications provided by the motion sensor 230 to determine the start and end of video segments for further processing, camera 202 may instead rely on motion indications provided by such image processors to achieve this purpose.
[0036] In some implementations, the edge image processing unit 220 may include one or more image processors (e.g., ML models and / or other computer vision processing components) for recognizing features (e.g., motion, people, objects, etc.) within the image 244, and the remote image processing unit 222 may include one or more different image processors (e.g., ML models and / or other CV processing components) for recognizing features within the image 244. The image processors may process the image 244, for example, to detect motion, recognize people, recognize faces, recognize objects, perform facial recognition, etc. In some implementations, the processing power of the server 208 used by the monitoring service 206 may be significantly greater than the processing power of the processors included in the edge image processing unit 220, thereby allowing the monitoring service 206 to employ more complex image processors and / or execute a greater number of such image processors in parallel.
[0037] like Figure 2As shown, the edge image processing unit 220 can generate an edge processing result 246 corresponding to one or more identified features of the image 244 (and optionally, the image 244 itself), and can send the edge processing result 246 to the event / video data repository 224 so that the event / video data repository 224 generates a new record for a specific event (e.g., by combining the following). Figure 4A (Create a new row in the described table 402) and store the data for that event in that record. Although Figure 2 As not shown, it should be understood that in some embodiments, the edge image processing unit 220 may additionally or alternatively send the edge processing result 246 directly to the remote image processing unit 222 for processing, so that the remote image processing unit 222 does not need to wait to create a new record in the event / video data repository 224 before starting to analyze the edge processing result 246. In some embodiments, the edge processing result 246 may include metadata for the event, such as an identifier for the event, a timestamp indicating when the event occurred, an identifier for a user residing in or otherwise authorized to access the monitoring location 204, an identifier for the monitoring location 204, an identifier for the camera 202 that captured the image 244, etc.
[0038] As described above, in some implementations, the event / video data repository 224 may include table 402 (see Table 402). Figure 4A This table includes data rows representing records of corresponding detected events. The individual columns of Table 402 may represent items or fragments of data or metadata associated with the record represented in the corresponding row (e.g., a unique identifier for the event, a timestamp for the event, an image of the event and / or a pointer to a location where the image of the event is stored, an identifier of the monitoring location 204 associated with the record, an identifier of a user residing in or otherwise authorized to access the monitoring location 204, an identifier of the camera 202 that captured the image of the event, etc.). In some embodiments, Table 402 may represent a compilation of records of a large number of events detected by the security system 200, including records that need to be assigned to the monitoring agent 212 for review, records that have already been assigned to the monitoring agent 212 for review, and records that have been processed / cancelled by the monitoring agent 212 or are the result of automated processing performed by the security system 200. Additional details regarding the exemplary Table 402 are provided in conjunction with... Figure 4A As described below.
[0039] Figure 3 This is a flowchart of an exemplary routine 300 according to some aspects of this disclosure, which may be provided by Figure 2 The application 228 shown is executed to achieve the above combination. Figure 1A The functions of the screen 102.
[0040] At step 302 of routine 300, the application of the computing device may provide a user interface (UI), such as screen 102, to (i) play back a video in a first area of the screen (e.g., video playback window 104) and (ii) display a plurality of interactive elements (e.g., detection UI element 108) corresponding to features detected in the video, which are displayed in a second area of the screen different from the first area.
[0041] At step 304 of routine 300, application 228 can determine that the video playback has reached the first time position in the video (e.g., with time offset 436 (see below)). Figure 4B The corresponding position), the first time position corresponds to the first interactive element among the multiple interactive elements displayed in the second area.
[0042] At step 306 of routine 300, application 228 may invoke the first interactive element (e.g., Figure 1A The change in appearance of the detected UI element 108C is used to visually distinguish the first interactive element from other elements among multiple interactive elements (e.g., Figure 1A The change (shown as detection UI elements 108A, 108B, and 108D) is temporary, such that when the video playback progresses beyond a first time position, for example, when the playback reaches a time position corresponding to another detection UI element 108, the appearance of the first interactive element is restored to the appearance it displayed before reaching that first time position in the video.
[0043] Figure 4A An exemplary table or data structure for event 402 is shown, which can be used to store records for various events detected by security system 200. As shown, for a single event, table 402 can be populated with data specifically representing the following items: event identifier (ID) 404, timestamp 406, user ID 408, location ID 410, camera ID 412, image 414, first frame time 416, feature identifier 418, event state 420, and event handling 422. The nature of these entries, as well as other components of application 228 and security system 200, can use these entries to implement the above-described combination. Figure 1A The way the functions are described in the overview is described in more detail below.
[0044] Event ID 404 can identify different events detected by security system 200, and data in the same row as a given event ID 404 can correspond to the same event.
[0045] The event timestamp 406 can indicate the time when the corresponding event was detected. In some implementations, application 228 can use the event timestamp 406 to populate the date and time identifier 130 on the screen 102 of user device 214, such as... Figure 1A As shown.
[0046] User ID 408 may represent a user associated with the detected event (e.g., a user residing in or otherwise authorized to access the monitoring location 204 where the event was detected). In some implementations, application 228 may use user ID 408 to identify one or more event records in event table 402 that are available for review by user 216. For example, in some implementations, application 228 may present screen 102 for a specific event in response to user 216 selecting a UI element representing that event record from a set of UI elements representing various event records available for review by user 216.
[0047] Location ID 410 can identify the monitored location where the event was detected (e.g., monitored location 204). In some implementations, application 228 can use location ID 410 to populate a location identifier (not shown) on screen 102 of user device 214, such as by indicating that the event in question was detected at user 216’s main residence, user 216’s vacation home, etc.
[0048] Camera ID 412 may represent a camera (e.g., camera 202) that has recorded one or more images of the detected event. In some implementations, application 228 may use camera ID 412 in the event log to identify the camera with which application 228 wants to establish a connection (e.g., a peer-to-peer connection) to receive real-time or near-real-time video feeds in response to selection of UI element 128 on screen 102, as described above.
[0049] Image 414 may represent the camera identified by camera ID 412 when an event is detected (e.g., by...). Figure 2 The images 244 acquired by the camera 202 shown are one or more images (e.g., snapshots or video streams) and / or may represent one or more images created using such acquired images, such as generating a face image 120 by cropping an acquired image that includes a detected face and / or annotating the acquired image (e.g., by overlaying the acquired image with a rectangle of red or other prominent coloring around an element) to identify specific features (e.g., face, person, weapon, etc.). In some embodiments, the image 414 entry in Table 402 may include an object containing links or pointers to such images.
[0050] The first frame time 416 may be a timestamp indicating the time of day for the first frame of a video recording for an event. In some embodiments, the first frame time 416 may be slightly offset from the event timestamp 406, such as when there is a slight delay between event detection and the start of video recording. In some embodiments, the remote image processing unit 222 may use the first frame time 416 to calculate a time offset 436 between the time of day for the video frame recording which includes detected features (e.g., motion, people, faces, weapons, etc.) and the first frame time 416 (see [link to documentation]). Figure 4B As described in more detail below, in some implementations, application 228 may use such a time offset 436 (e.g., seconds) to determine whether and when a particular detection UI element 108 is highlighted or otherwise indicated on screen 102 during playback of the recorded video within video playback window 104 and / or to determine the location to jump to within such recorded video in response to user 216 selecting one of the detection UI elements 108.
[0051] Feature identifier 418 may include information about one or more features identified in the recorded image 414 (e.g., features identified by edge image processing unit 220 and / or remote image processing unit 222) and / or one or more features identified by monitoring agent 212 during the review of the event log via monitoring application 226. Such information may include, for example, identifiers of motion detected in image 414, identifiers of people detected in image 414, identifiers of faces detected in image 414, identifiers of weapons detected in image 414, etc. An exemplary data structure including feature identifier 418 of the event log is described below. Figure 4B To describe.
[0052] Event status 420 can indicate the processing status of each record in security system 1. For example, event status 420 for a record can indicate that the record is active and requires further processing (e.g., "new"), is awaiting review by monitoring agent 212 (e.g., "assigned"), is being actively reviewed by monitoring agent 212 (e.g., "under review"), has been marked as "cancelled" or "processed" (e.g., by monitoring agent 212 or automatically by a component of security system 200), has "expired", has caused an emergency "dispatch" service and / or is in a "suspended" state (e.g., has been grouped with similar related records currently being reviewed by monitoring agent 212).
[0053] In some implementations, event status 420 may additionally or alternatively indicate whether the corresponding event is actively monitored by monitoring agent 212, such as by auditing event data 252 and taking one or more actions 254 related to the event, as referenced below. Figure 2 More detailed description. In such embodiments, application 228 may use event state 420 to populate monitoring status identifier 132 on screen 102 of user device 214, such as by displaying a “monitoring” status, for example, as Figure 1A As shown, this indicates that the event in question is actively monitored by monitoring agent 212.
[0054] Incident handling 422 can refer to the handling of an incident in question after review by one or more monitoring agents 212 and / or user 216, such as an "emergency" situation (e.g., a life-threatening or violent situation) or an "urgent" situation (e.g., package theft, property damage, or vandalism), the incident being "handled" by monitoring agent 212, police or fire departments being "dispatched" to handle the incident, the review of the incident being "cancelled" after personnel accurately provide a security word or other identifying information, or the review of the incident being "cancelled" by user 216 (e.g., via application 228), etc. In some implementations, the marked incident handling 422 can be used, for example, to determine whether to send a notification to the user (e.g., push notification, SMS message, email, etc.), whether to flag a record for user review, or whether to include a record in a list of records to be reviewed in response to a user query specifying one or more filtering criteria, etc.
[0055] although Figure 4A Not shown, but it should be understood that Table 402 may also include additional data that can be used for various purposes, such as indications of the geographic location / coordinates of monitoring location 204, descriptions of the records (e.g., “motion detected by the backyard camera”), actions taken by monitoring agent 212 in reviewing information corresponding to the records, one or more recorded audio tracks for the records, status changes of one or more sensors (e.g., door lock sensors) at monitoring location 204, etc.
[0056] Refer again Figure 2 Similar to edge image processing unit 220, remote image processing unit 222 can process image 244 (or portions of image 244, such as one or more frames identified by edge image processing unit 220) acquired by camera 202 to identify one or more features. In some implementations, the processing performed by one or more image processors of edge image processing unit 220 can be used to inform and / or enhance the processing performed by one or more image processors of remote image processing unit 222.
[0057] As an example, one or more image processors of the edge image processing unit 220 can perform initial processing to identify keyframes in the image that may represent motion, people, faces, etc., and one or more image processors of the remote image processing unit 222 can perform additional processing only on the keyframes identified by the one or more image processors of the edge image processing unit 220. As another example, one or more image processors of the edge image processing unit 220 can process the image to identify specific frames including motion, and one or more image processors of the remote image processing unit 222 can process the image to detect people only on the specific frames identified by the one or more image processors of the edge image processing unit 220. As yet another example, one or more image processors of the edge image processing unit 220 can process the image to identify specific frames of the image including people, and one or more image processors of the remote image processing unit 222 can process the image to detect and / or identify faces only on the specific frames identified by the one or more image processors of the edge image processing unit 220. As another example, one or more image processors of the edge image processing unit 220 can process images to identify specific frames including facial images, and one or more image processors of the remote image processing unit 222 can process images to perform enhanced facial recognition and / or recognize faces only on specific frames identified by one or more image processors of the edge image processing unit 220. Furthermore, in some implementations, the remote image processing unit 222 itself can process images using multiple different image processing models, where some image processors rely on results obtained by one or more other image processors.
[0058] In some implementations, the remote image processing component 222 may be a software application executed by one or more processors of the monitoring service 206. For example, as described above, in some implementations, the server 208 of the monitoring service 206 (see...) Figure 2 The device may include one or more computer-readable media encoded with instructions that, when executed by one or more processors of server 208, enable server 208 to perform the functions of the remote image processing unit 222 described herein.
[0059] like Figure 2As shown, the remote image processing component 222 may receive content 248 (e.g., some or all of the data from a row of table 402) of a record stored in the event / video data repository 224. Content 248 may include, for example, one or more images (e.g., still images and / or videos) or pointers to one or more locations where such images are stored, and other data that may come from the record, such as identifiers for the record, identifiers for features identified within the recorded images, timestamps indicating when an event was detected, identifiers for those residing in or otherwise authorized to access monitoring location 204, identifiers for monitoring location 204, identifiers for the camera 202 that captured the image, etc. As discussed above, in some embodiments, the remote image processing component 222 may retrieve content 248 in response to receiving an indication or otherwise determining that a record stored in the event / video data repository 224 has been added to or modified. For example, the remote image processing component 222 may receive such an indication (e.g., from the event / video data repository 224, the event handler, or the edge image processing component 220) whenever one or more images 414 are added to or modified for a record.
[0060] In some implementations, the remote image processing unit 222 may further receive context data from one or more context data repositories (not shown). Such context data may include, for example, information from one or more profiles corresponding to the monitoring location 204 and / or the user, and this information may be used to enhance or improve the processing performed by the remote image processing unit 222. As an example, the context data may include one or more biometric embedding vectors for a known individual that can be used, for example, for facial recognition processing (e.g., corresponding to a visitor profile created for such an individual).
[0061] The remote image processing unit 222 can process images (and possibly other data) included in or pointed to by content 248 received from the event / video data repository 224 (and optionally, data received from the context data repository) to detect and / or confirm the presence of one or more features (e.g., motion, people, faces, identified faces, etc.) within such images. The remote image processing unit 222 can generate one or more identifiers 250 corresponding to the identified features, and such identifiers 250 can be added to the record for the event, for example, by writing the identifiers to the row corresponding to the event in table 402 (e.g., as feature identifier 418).
[0062] For example, based on identifier 250 received from remote image processing unit 222 or otherwise, exemplary information may be included within the characteristic identifier 418 for personal records in Table 402. Figure 4B Presented in tabular form, as a data object or table 430. In some implementations, Figure 4B The information shown can be stored in table 402 as a data object, for example, as... Figure 4A The entry "FI1" of feature identifier 418 shown. For example... Figure 4B As shown, such data objects can describe one or more features detected in a corresponding video frame for a detected event, and may include, for example, feature type 432, feature image pointer 434, time offset 436, and feature metadata 438. In some implementations, application 228 can use the information in such data objects to generate a scrollable list of detected UI elements 108, such as Figure 1A Those shown. For example, Figure 4B The rows of Table 430 shown may include information corresponding to the respective detected UI element 108 to be included in such a scrollable list. About Figure 1A The exemplary screen 102 shown may include eight rows of information in a table / data object 430 for detected events reviewed via this screen 102, since UI element 110 indicates a total of “8” detection UI elements 108 available for review by user 216.
[0063] Feature type 432 can indicate the type of feature (e.g., “motion,” “person,” “identified face,” “unidentified face,” “weapon,” etc.) detected by the edge image processing unit 220, the remote image processing unit 222, and / or the monitoring agent 212 via the monitoring application 226. In some embodiments, the application 228 can use feature type 432 to determine which features to include. Figure 1A The information in the detection type identifier 116 shown.
[0064] The feature image pointer 434 can identify the location of the image storing the detected features (e.g., face image 120, thumbnail image 122, etc.). In some implementations, the application 228 can use the feature image pointer 434 to identify and retrieve images (e.g., face image 120, thumbnail image 122, etc.) to be included in the corresponding detection UI element 108.
[0065] The time offset 436 can represent a calculated amount of time (e.g., seconds) between the first frame time 416 (e.g., indicating the time of day for the first frame of a video recording for a detected event) and the time of day for a video frame recording the feature in question. As described above, in some embodiments, the remote image processing unit 222 can determine the time offset 436 of the detected feature by calculating the difference between a timestamp indicating the time of day for a video frame recording the feature including the detected feature (e.g., motion, person, face, weapon, etc.) and the first frame time 416. In some embodiments, the application 228 can use such a time offset 436 to determine the value of a corresponding timestamp 118 represented on the detection UI element 108, for example, by adding the time offset 436 of the detected feature to the first frame time 416 to determine the approximate time of day for a video frame recording the feature including the detected feature.
[0066] In some implementations, application 228 may additionally or alternatively use time offset 436 to determine whether and when a particular detection UI element 108 is highlighted or otherwise indicated on screen 102 during playback of a video recorded within video playback window 104 and / or to determine the jump position within such recorded video in response to user 216 selecting a corresponding detection UI element among detection UI elements 108. For example, while application 228 is playing back a recorded video within video playback window 104, application 228 may track the relative time position of the currently displayed video frame relative to the first frame of the recorded video (e.g., by using the same playback counter value used to update progress bar 106) to identify when that time position matches time offset 436, and in response to identifying such a match, application 228 may cause the corresponding detection UI element 108 to be highlighted or otherwise indicated on screen 102, and also cause another detection UI element 108 that was previously highlighted or otherwise indicated (if another detection UI element was previously indicated) to be de-highlighted or otherwise indicated.
[0067] Additionally or alternatively, in response to application 228 determining that user 216 has selected one of the detection UI elements 108 on screen 102, application 228 may cause the video being played back in video playback window 104 to jump to a video frame located at a time position matching the time offset 436 of the selected detection UI element 108 (e.g., determined using the same playback counter value used to update progress bar 106). In some embodiments, time offset 436 may be set to a value slightly lower than the actual time difference calculated as described above (e.g., 2-3 seconds smaller than the actual time difference) so that the detection UI element 108 of the detected feature is highlighted or otherwise indicated shortly before reaching the video frame including the detected feature during playback of the video recorded within video playback window 104 and / or, in response to user input selection of the detection UI element 108 of the detected feature, the video being played back in video playback window 104 jumps to a position shortly before (e.g., 2-3 seconds earlier) the video frame including that feature.
[0068] Feature metadata 438 may represent additional information about the detected features, such as the detection type identifier 116 for the identified face (e.g., Figure 1A The name of the person shown in the detection type identifier 116A).
[0069] Figure 5 This is a flowchart illustrating an exemplary process 505 that can be employed by a remote image processing unit 222 to perform relevant image processing according to some embodiments of the present disclosure. Figure 5 As shown, process 505 may begin at step 510, where the remote image processing unit 222 may receive content 248 from records (e.g., activity records) within the event / video data repository 224, and may also optionally receive data (e.g., context data) from a context data repository (not shown). In some embodiments, a record in table 402 may be considered "active" if it has an event state 420 of "new," "assigned," "under review," or "pending." The remote image processing unit 222 may identify the active record to be processed in any of a variety of ways, and may retrieve content 248 and / or context data, for example, in response to receiving a notification or otherwise determining that content 248 and / or context data has changed in a potentially relevant manner.
[0070] In step 515, the remote image processing component 222 may determine the next frame of the recorded video included in or pointed to by the content 248 received from the event / video data repository 224. In some implementations, for example, the content 248 may include or point to a sequence of video frames, and the remote image processing component 222 may process these frames sequentially, or may process a subset of frames (e.g., every ten frames), wherein the “next frame” determined in step 515 corresponds to the next unprocessed frame in the frame sequence.
[0071] In step 520 of process 505, the remote image processing component 222 may, for example, cause one or more first image processors to process the frame (and possibly one or more adjacent or nearby frames) to determine whether the frame corresponds to a moving object. In some implementations, motion may be detected, for example, by using one or more functions of the OpenCV library (accessible at the Uniform Resource Locator (URL) "opencv.org") to detect differences between frames that indicate that the object represented in the frame is in motion. When, in step 520, the remote image processing component 222 determines that the frame includes an object that was in motion at the time the frame was acquired, the remote image processing component 222 may generate a feature indicator 418 indicating the detected motion and have the feature indicator 418 added to a record for the event.
[0072] Based on decision 525, if the remote image processing component 222 determines that the frame does not correspond to a moving object, process 505 may terminate. Alternatively, if the remote image processing component 222 determines (at decision 525) that the frame does correspond to a moving object, process 505 may proceed to step 530, where the remote image processing component 222 may cause one or more second image processors to process the frame to determine whether the frame includes a person. An example of an ML model that can be used for person detection is YOLO (accessible via the URL "github.com"). When the remote image processing component 222 determines in step 530 that the frame includes a person, the remote image processing component 222 may generate a feature indicator 418 indicating the detected person and add the feature indicator 418 to the record for the event.
[0073] According to decision 535, if the remote image processing component 222 determines that the frame does not contain a person, process 505 may terminate. Alternatively, if the remote image processing component 222 determines (at decision 535) that the frame contains a person, process 505 may instead proceed to step 540, in which the remote image processing component 222 may cause one or more third image processors to process the frame to determine whether the frame contains a face. An example of an ML model that can be used for face detection is RetinaFace (accessible via the URL "github.com"). When the remote image processing component 222 determines in step 540 that the frame contains a face, the remote image processing component 222 may generate a feature indicator 418 indicating the detected face and have the feature indicator 418 added to the record for the event.
[0074] According to decision 545, if the remote image processing component 222 determines that the frame does not contain a face, process 505 may terminate. Alternatively, if the remote image processing component 222 determines (at decision 545) that the frame contains a face, process 505 may instead proceed to step 550, in which the remote image processing component 222 may cause one or more fourth image processors to perform an enhanced face recognition process to more accurately identify and locate faces in the frame. An example of an ML model that can be used for enhanced face detection is MTCNN_face_detection_alignment (accessible via the URL "github.com"). The remote image processing component 222 may then generate a new feature indicator 418 indicating the result of the enhanced face detection and add the feature indicator 418 to the record for the event, and / or modify the feature indicator generated in step 540 to include this result.
[0075] Finally, process 505 may proceed to step 555, where the remote image processing component 222 may perform face recognition on the detected face in the frame, such as by generating biometric embedding vectors of the detected face and comparing these embedding vectors with a known face database, in an attempt to identify the person based on the identified face. An example of an ML model that can be used for face recognition is AdaFace (accessible via the URL "github.com"). When, in step 555, the remote image processing component 222 determines that the frame represents a known face, the remote image processing component 222 may generate a feature identifier 418 indicating the identified face and have the feature identifier 418 added to a record for the event. As described above, in some embodiments, such a feature identifier 418 for the identified face may include feature metadata 438 indicating the name of the identified person.
[0076] It should be understood that in some embodiments, the edge image processing unit 220 and / or the remote image processing unit 222 do not perform image processing (e.g., as...). Figure 5 Instead of using the methods shown above, image processing of the aforementioned type or possibly other types can be performed in parallel or partially parallel using one or more ML models and / or other computer vision (CV) processing components to identify one or more other feature types. In such implementations, edge image processing component 220 and / or remote image processing component 222 can generate feature indicators 418 indicating features detected by the respective components, and add feature indicators 418 to the record as soon as they are generated by the respective ML model and / or other computer vision (CV) processing component. Additionally, as described above, in some implementations, edge image processing results received from edge image processing component 220 can be used to enhance the image processing performed by remote image processing component 222, for example, by identifying one or more keyframes to be further processed by remote image processing component 222.
[0077] In some implementations, notifications regarding "actionable" events represented in event table 402 (e.g., an event where the remote image processing unit 222 identifies one or more features of interest) can be dispatched to the appropriate monitoring application 226 for review by the monitoring agent 212. In some implementations, the monitoring service 206 can use the contents of event table 402 to assign individual events to the various monitoring agents 212 currently online with the monitoring application 226. Figure 2As shown, in some embodiments, a monitoring application 226, operated by a monitoring agent 212 to which an event log has been assigned for review, can receive event data 252 of the event log to be reviewed and can cause the monitoring device 210 to present various user interface screens based on the event data. This event data allows the monitoring agent 212 to determine whether the event represents an actual security problem rather than a harmless situation, such as by reviewing recorded video of the event, evaluating the accuracy of one or more feature detections performed by the edge image processing unit 220 and / or the remote image processing unit 222, reviewing real-time or near-real-time video from the monitoring location, and possibly communicating with personnel at the monitoring location 204 (e.g., via the speaker and microphone of the camera 202). During and / or after such review, the monitoring application 226 can transmit event actions 254 to the event data / video data repository 224 based on one or more inputs provided by the monitoring agent 212 to such user interface screens, thereby updating certain information in table 402 (e.g., event status 420 and / or event handling 422). As described above, in some embodiments, application 228 may use user device 214 to present a list of event records available for review by user 216, and application 228 may cause user device 214 to display screen 102 in response to the user's selection of one of those event records. In some embodiments, the event records presented on such a list of "reviewable" event records may be determined at least in part based on the values of the event status 420 and / or event handling 422 entries in table 402.
[0078] like Figure 2 As shown, in some implementations, application 228 may receive event details 256 from event data / video data repository 224, in particular to present Figure 1A Screen 102 is shown. Now refer to... Figure 6 The description includes an exemplary routine 600 performed by application 228 using such event details 256.
[0079] like Figure 6 As shown, routine 600 may begin at step 602, where application 228 may identify one or more event records in table 402 that are available for review by user 216. In some implementations, for example, certain event records in the event logs may be marked as available for review by user 216, for example, based on user preferences indicating the type of event and / or type of event handling that user 216 wishes to review.
[0080] At step 604 of routine 600, application 228 may display a UI screen (not shown) that allows the user to select specific event records to be audited, such as by presenting multiple selectable UI elements for the corresponding event record. In some implementations, application 228 may provide one or more additional UI elements that allow user 216 to filter and / or sort event records available for auditing (e.g., based on property location, such as if user 216 has multiple properties monitored by security system 200, based on date and / or time, based on event type, based on event handling, based on features identified in video acquired for the event, and / or based on any other criteria using entries for event records in table 402).
[0081] At decision 606 in routine 600, application 228 can determine (e.g., by monitoring touch input provided to the touchscreen of user device 214) whether an event log has been selected via the UI screen presented according to step 604. As indicated, application 228 can continue (according to steps 602 and 604) to identify and enable the selection of new event logs available for review by user 216 until application 228 determines (at decision 606) that user 216 has selected a specific event log to review.
[0082] When application 228 determines at decision 606 that an event record has been selected for review, routine 600 may proceed to step 608, where application 228 retrieves event details 256 of the selected record from event data / video data repository 224 (see [link to relevant documentation]). Figure 2 ), thereby enabling the application 228 to, for example, via Figure 1A The screen 102 shown presents the detected UI element 108 and the recorded video in response to the event, allowing the user 216 to quickly navigate to the part of the recorded video of particular interest. (The above is combined with...) Figure 4A and Figure 4B Information from Table 402 describes the various elements that can be used to fill and render on screen 102.
[0083] In some implementations, the event details 256 received by application 228 may include metadata describing the detection of multiple different features (e.g., face, motion, and person) within the recorded video of the selected event, wherein the corresponding feature types are described in a separate list of detection metadata. For example, as Figure 7 As shown on the left, the event details 256 received by application 228 may include a first list 702 for face detection and a second list 704 for motion detection (in... Figure 7 The third list of people detected (marked as "moving"), and the third list of people detected (706) Figure 7 (Indicated as "track"). In some implementations, the individual entries on lists 702, 704, and 706 can be associated with... Figure 4B The data in the corresponding rows of Table 430 shown corresponds to this. For example... Figure 7 As shown, the items in lists 702, 704, and 706 may include the above-mentioned items. Figure 4B The time offset of the type mentioned is 436 (in) Figure 7 The data is indicated as “pts_seconds” and other metadata.
[0084] Refer again Figure 6 At step 610 of routine 600, application 228 can use the retrieved event details 256 to render the screen, for example, Figure 1A Screen 102 is shown in the diagram. In some embodiments, such as... Figure 7 As shown on the right, application 228 can merge lists 702, 704, and 706 of detections received from event / video data repository 224 to generate a combined list 708 of detections, and can sort the combined list 708 chronologically using a time offset of 436. Application 228 can thus create a list of individual features detected during video recording, organized chronologically according to the order in which they appear in the video. After merging and sorting lists 702, 704, and 706 in this way, application 228 can use the combined list 708 to present the detected UI elements 108 in the area of screen 102 below video playback window 104 as a visible and scannable vertical list of detected UI elements 108 ordered chronologically, such as... Figure 1A As shown.
[0085] When application 228 first renders screen 102, application 228 can cause the recorded video for the event to begin playback in video playback window 104 starting from the first (temporally) recorded frame of the video, and application 228 can also cause the first few detected UI elements 108 for the event (e.g., Figure 1AThe detection UI elements 108A-108D are displayed on the screen in Figure 102, indicating, along with UI element 110, the total number of detection UI elements (e.g., "8" detection UI elements) created for the event log. The selection of the detection UI element 108 initially displayed on screen 102 when video playback begins can be based on a time offset 436 within a combination list 708, such as by selecting the four entries in combination list 708 with the lowest time offset 436. In some implementations, if no features are identified in the first few video frames (e.g., by edge image processing unit 220 and / or remote image processing unit 222), the displayed detection UI element 108 will not be highlighted or otherwise indicated when the recorded video first begins playback. The progress bar 106 also indicates that the recorded video has just begun playing when video playback begins.
[0086] At decision 612 in routine 600, application 228 can determine (e.g., by monitoring touch input provided to the touchscreen of user device 214) whether one of the displayed detection UI elements 108 has been selected by user 216 (e.g., by touching it with a finger).
[0087] When application 228 determines (at decision 612) that one of the UI elements 108 has been selected for detection, routine 600 may proceed to step 614, where application 228 may cause the recorded video being played back in video playback window 104 to jump to a position corresponding to the time offset 436 for the detected UI element, such as by instructing a video playback application on user device 214 that is processing video playback within video playback window 104 to jump to such a position. As described above, the time offset 436 for detecting UI element 108 may represent the amount of time (e.g., seconds) between the first frame of the video and a frame of the recorded video in which features of the detected UI element 108 (e.g., motion, people, identified faces, unidentified faces, weapons, etc.) are detected. When application 228 causes the played-back video to jump in this manner, application 228 may also cause the progress bar 106 to be updated to indicate the relative position of the newly displayed video frame (e.g., the video frame in which features of the selected detected UI element 108 are detected) relative to the entire sequence of video frames of the recorded video for the event being reviewed. After step 614, the routine can return to decision 612.
[0088] When application 228 determines at decision 612 that user 216 has not yet selected detection UI element 108, routine 600 may proceed to decision 616, where application 228 may determine (e.g., based on data received from a video player application on user device 214 that is processing video playback within video playback window 104) whether the progress of video playback (e.g., the time position of the current frame relative to a first video frame, such as indicated by progress bar 106) has reached the time offset 436 indicated in the combination list 708 of newly detected UI element 108. As described above, in some embodiments, application 228 may selectively pause or resume playback of recorded video in video playback window 104 in response to input from user 216, such as by toggling between a "play" state and a "pause" state in response to user 216 touching video playback window 104.
[0089] When application 228 determines at decision 616 that the progress of video playback has not yet reached the time offset 436 of the newly detected UI element 108, routine 600 can return to decision 612. On the other hand, when application 228 determines that the progress of video playback has reached the time offset 436 of the newly detected UI element 108, routine 600 can proceed to decision 618, where application 228 can determine (e.g., by evaluating which detected UI elements 108 are currently displayed on screen 102) whether the newly detected UI element in question is "hidden," for example, currently invisible on screen 102. For example, in Figure 1A In the exemplary screen 102 shown, only four of the eight detection UI elements 108 available for review are visible on screen 102 at any given time. However, as described above, the list of detection UI elements 108 can be scrolled manually or automatically to reveal other "hidden" detection UI elements 108.
[0090] When at decision 618 the application 228 determines (e.g., by evaluating which detected UI elements 108 are currently displayed on screen 102) that there are no newly detected UI elements currently hidden (identified according to decision 616), routine 600 can proceed to step 620, where the application 228 can cause the identified detected UI elements 108 to be highlighted or otherwise indicated. Figure 1A In the example shown, UI element 108C has been highlighted. Also, at step 620, if step 620 has been performed previously for a different UI element 108, the application 228 can remove the highlight or other indication from another UI element 108.
[0091] On the other hand, when application 228 (based on decision 618) determines that the newly detected UI element is currently hidden, routine 600 can proceed to decision 622, where application 228 can determine (e.g., by tracking user interactions with screen 102 over time) whether user 216 recently (e.g., within the previous five seconds) manually scrolled the list of detected UI elements 108. Application 228 can make such a determination, for example, such that if user 216 controls the scrolling operation for some purpose, the application can avoid automatically scrolling the list of detected UI elements 108 (as described below), for example, to determine whether to take action on facial image 120, such as adjusting a visitor profile by selecting UI element 124.
[0092] When application 228 determines at decision 622 (e.g., by tracking user interactions with screen 102 over time) that user 216 has not recently manually scrolled the list of detected UI elements 108, application 228 can proceed to step 624, where application 228 can scroll the list of detected UI elements 108 to reveal newly detected UI elements (identified according to decision recognition 616). For example, refer to... Figure 1A If UI element 108D is highlighted when step 624 is reached, application 228 can scroll the list of detected UI elements 108 to reveal the "hidden" detected UI element 108 located directly below detected UI element 108D. After scrolling the list of detected UI elements 108 (according to step 624) to reveal the new detected UI element (identified according to decision 616), the routine can proceed to step 620, where the new detected UI element can be highlighted or otherwise indicated as described above.
[0093] When at decision 622, application 228 determines that user 216 has recently (e.g., within the last five seconds) manually scrolled through the list of detected UI elements 108, application 228 can proceed directly to step 620, where a new detected UI element (even if it is currently hidden) can be highlighted or otherwise indicated, thereby ensuring that the correct detected UI element 108 is highlighted or otherwise indicated (based on video playback progress) if user 216 continues to manually scroll through the list of detected UI elements 108 to reveal the newly highlighted detected UI element 108.
[0094] Figure 8 This is a schematic diagram of an exemplary security system 800 that can be adopted by various aspects of this disclosure. As shown, in some embodiments, the security system 800 may include a plurality of monitoring locations 204 ( Figure 8The diagram only shows one monitoring location), monitoring center environment 822, monitoring center environment 826, one or more client devices 214, and one or more communication networks 820. Monitoring location 204, monitoring center environment 822, monitoring center environment 826, one or more client devices 214, and communication network 820 may each include one or more computing devices (e.g., as shown in the reference below). Figure 9 (As described). User device 214 may include one or more applications 228, for example, as applications hosted on user device 214 or otherwise accessible to the user device. In some embodiments, application 228 may be embodied as a web application accessible via a browser on user device 214. Monitoring center environment 822 may include one or more monitoring applications 226, for example, as applications hosted on a computing device within monitoring center environment 822 or otherwise accessible to the computing device. In some implementations, monitoring application 226 may be embodied as a web application accessible via a browser on a computing device operated by monitoring agent 212 within monitoring center environment 822. Monitoring center environment 826 may include monitoring service 830 and one or more transport services 828.
[0095] like Figure 8 As shown, the monitored location 204 may include one or more image capture devices (e.g., cameras 202A and 202B), one or more contact sensor assemblies (e.g., contact sensor assembly 806), one or more keypads (e.g., keypad 808), one or more motion sensor assemblies (e.g., motion sensor assembly 810), a base station 812, and a router 814. As illustrated, the base station 812 may host a monitoring client 816.
[0096] In some implementations, router 814 may be a wireless router configured to communicate with devices (e.g., devices 202A, 202B, 806, 808, 810, and 812) located at monitoring position 204 via communication conforming to communication standards (such as any of the various Institute of Electrical and Electronics Engineers (IEEE) 308.11 standards). Figure 8As shown, router 814 can also be configured to communicate with network 820. In some implementations, router 814 can establish a local area network (LAN) within and near the monitored location 204. In other implementations, other types of networking technologies may be used additionally or alternatively within the monitored location 204. For example, in some implementations, base station 812 may receive and forward communication packets sent by one or both of cameras 202A and 202B via a point-to-point personal area network (PAN) protocol such as BLUETOOTH. Other suitable wired, wireless, and mesh networking technologies and topologies will become apparent from this disclosure and are intended to fall within the scope of the examples disclosed herein.
[0097] Network 820 may include one or more public and / or private networks supporting, for example, Internet Protocol (IP) communication. Network 820 may include, for example, one or more LANs, one or more PANs, and / or one or more wide area networks (WANs). Possible LANs include wired or wireless networks supporting various LAN standards, such as versions of IEEE 308.11. Possible PANs include wired or wireless networks supporting various PAN standards, such as BlueTooth, ZigBee, etc. Possible WANs include wired or wireless networks supporting various WAN standards, such as Code Division Multiple Access (CMDA), Global System for Mobile Communications (GSM), etc. Regardless of the specific network technology used, network 820 can connect components within monitoring location 204, monitoring center environment 822, monitoring center environment 826, and client device 214, and enable data communication between these components. In at least some implementations, both monitoring center environment 822 and monitoring center environment 826 may include network components (e.g., similar to router 814) configured to communicate with network 820 and various computing devices within these environments.
[0098] The monitoring center environment 826 may include physical space, communications, cooling, and power infrastructure to support the networked operation of a large number of computing devices. For example, the infrastructure of the monitoring center environment 826 may include rack space where computing devices can be installed, uninterruptible power supply, cooling pressurization chambers and equipment, and network devices. The monitoring center environment 826 may be dedicated to the security system 800, may be a non-dedicated, commercially available cloud computing service (e.g., Microsoft Azure, Amazon Web Services, Google Cloud, etc.), or may include a hybrid configuration of both dedicated and non-dedicated resources. Figure 8 As shown, regardless of its physical or logical configuration, the monitoring center environment 826 can be configured to host monitoring services 830 and transport services 828.
[0099] The monitoring center environment 822 may include multiple computing devices (e.g., desktop computers) and network devices (e.g., one or more routers) that enable communication between the computing devices and network 820. Client devices 214 may each include personal computing devices (e.g., desktop computers, laptops, tablets, smartphones, etc.) and network devices (e.g., routers, cellular modems, cellular radio components, etc.). Figure 8 As shown, the monitoring center environment 822 can be configured to host the monitoring application 226, and the client device 214 can be configured to host the client application 228.
[0100] Devices 202A, 202B, 806, and 810 can be configured to acquire analog signals via sensors incorporated in the device, generate digital sensor data based on the acquired signals, and transmit the sensor data (e.g., via a wireless link with router 814) to one or more components (e.g., the remote image processing component 222 described above) within base station 812 and / or monitoring center environment 826. The type of sensor data generated and transmitted by these devices can vary depending on the characteristics of the sensors included in these devices. For example, image capture devices or cameras 202A and 202B can acquire ambient light, generate one or more frames of image data based on the acquired light, and transmit the frames to one or more components within base station 812 and / or monitoring center environment 826, but the pixel resolution and frame rate can vary depending on the capabilities of the device. In some implementations, cameras 202A and 202B may also receive and store filtering region configuration data, and filter the frames using one or more filtering regions (e.g., regions within the camera's FOV from which image data will be edited for various reasons, such as excluding trees that are likely to generate false positive motion detection results in windy weather) before transmitting the frames to base station 812 and / or monitoring center environment 826. Figure 8 In the example shown, camera 202A has a field of view (FOV) originating from the vicinity of the front door of monitored location 204, and can acquire images of the sidewalk 836, road 838, and the space between monitored location 204 and road 838. On the other hand, camera 202B has an FOV originating from the vicinity of the bathroom of monitored location 204, and can acquire images of the living room and dining room of monitored location 204. Camera 202B can further acquire images of outdoor areas outside monitored location 204, for example, through windows 818A and 818B to the right of monitored location 204.
[0101] Individual sensor assemblies deployed at the monitored locations 204 (e.g., Figure 8The contact sensor assembly 806 shown may include, for example, a sensor that can detect the presence of a magnetic field generated by a magnet when the magnet is near the sensor. When a magnetic field is present, the contact sensor assembly 806 can generate Boolean sensor data indicating a closed state of a window, door, etc. When no magnetic field is present, the contact sensor assembly 806 can alternatively generate Boolean sensor data indicating an open state of a window, door, etc. In either case, Figure 8 The contact sensor assembly 806 shown can transmit sensor data to the base station 812 indicating whether the front door at monitoring location 204 is open or closed.
[0102] Individual motion sensor assemblies deployed at the monitored locations 204 (e.g., Figure 8 The motion sensor assembly 810 shown may include, for example, a component capable of emitting high-frequency pressure waves (e.g., ultrasound) and a sensor capable of acquiring the reflection of the emitted waves. When the sensor detects a change in the reflected pressure waves (e.g., due to movement of one or more objects within the space monitored by the sensor), the motion sensor assembly 810 may generate Boolean sensor data indicating a warning state. When the sensor does not detect a change in the reflected pressure waves (e.g., due to no movement of objects within the monitored space), the motion sensor assembly 810 may instead generate Boolean sensor data indicating a stationary state. In either case, the motion sensor assembly 810 may transmit the sensor data to the base station 812. It should be noted that the specific sensing methods described above are not limited to this disclosure. For example, as just one example of an alternative implementation, the motion sensor assembly 810 may alternatively (or additionally) operate based on the detection of changes in reflected electromagnetic waves.
[0103] While specific types of sensors have been described above, it should be understood that other types of sensors may be additionally or alternatively employed within the monitored location 204 to detect the presence and / or movement of humans, or other conditions of interest (such as smoke, elevated carbon dioxide levels, standing water, etc.), and data indicating such conditions may be transmitted to base station 812. For example, although Figure 8 Not shown, but in some embodiments, one or more sensors may be employed to detect sudden changes in the measured temperature, sudden changes in incident infrared radiation, sudden changes in incident pressure waves (e.g., sound waves), etc. Furthermore, in some embodiments, some of these sensors and / or base station 812 may be additionally or alternatively configured to identify specific signal profiles indicating specific conditions, such as sound profiles indicating broken glass, footsteps, coughs, etc.
[0104] Figure 8The illustrated keypad 808 can be configured to interact with a user and interoperate with other devices located in the monitored location 204 in response to such interaction. For example, in some instances, the keypad 808 can be configured to receive input from a user specifying one or more commands and to communicate the specified commands to one or more addressing devices and / or processes, such as one or more devices in the monitored location 204, the monitoring application 226, and / or the monitoring service 830. The communicated commands may include, for example, codes authenticating the user as a resident of the monitored location 204 and / or codes requesting activation or deactivation of one or more devices in the monitored location 204. In some implementations, the keypad 808 may include a user interface (e.g., a haptic interface, such as a set of physical buttons or a set of “soft” buttons on a touchscreen) configured to interact with a user (e.g., receive input from the user and / or present output to the user). Further, in some implementations, the keypad 808 may receive responses to the communicated commands and present such responses as visual or audio output via the user interface.
[0105] Figure 8 The base station 812 shown can be configured to interact with other security system devices located at monitoring location 204 to provide local command and control functions and / or store-and-forward functions via executing monitoring client 816. To implement local command and control functions, base station 812 can perform various programmed operations by executing monitoring client 816 in response to various events. Examples of such events include receiving commands from keypad 808, receiving commands via network 820 from either monitoring application 226 or client application 228, and detecting the occurrence of scheduled events. Programmed operations performed by base station 812 via executing monitoring client 816 in response to events may include, for example, activating or deactivating one or more of devices 202A, 202B, 806, 808, and 810; sounding an alarm; reporting an event to monitoring service 830; and / or transmitting "location data" to one or more transmission services 828. Such location data may include, for example, data specifying sensor readings (sensor data), image data acquired by one or more cameras 202, configuration data of one or more devices in the apparatus set at monitoring location 204, commands input and received from the user (e.g., via keypad 808 or client application 228), or data derived from one or more of the above data types (e.g., filtered sensor data, filtered image data, sensor data summary, data specifying events detected at monitoring location 204 via sensor data, etc.).
[0106] In some implementations, to achieve store-and-forward functionality, base station 812 can receive sensor data by executing monitoring client 816, encapsulate the data for transmission, and store the encapsulated sensor data in local memory for subsequent transmission. This transmission of the encapsulated sensor data may include, for example, transmitting the encapsulated sensor data as a message payload to one or more transmission services in transmission service 828 when a communication link via network 820 to transmission service 828 is in operation. In some implementations, this encapsulation of sensor data may include filtering the sensor data using one or more filtering regions and / or generating one or more summaries of multiple sensor readings (maximum value, average value, changes in value since previous transmissions, etc.).
[0107] The transport service 828 of the monitoring center environment 826 can be configured to receive messages from a monitored location (e.g., monitored location 204), parse the messages to extract the payloads contained therein, and store the payloads and / or data derived from the payloads in one or more data repositories hosted in the monitoring center environment 826. (This is followed by a description of a process called "in conjunction with...") Figure 9 Describe an example of such a data repository. In some implementations, transport service 828 may expose and implement one or more application programming interfaces (APIs) configured to receive, process, and respond to calls from a base station (e.g., base station 812) via network 820. Individual cases of transport service 828 may be associated with and specific to certain manufacturers and / or models of location-based monitoring equipment (e.g., SIMPLI ISAFE devices, RING devices, etc.).
[0108] The Transport Service 828 API can be implemented using various architectural styles and interoperability standards. For example, in some implementations, one or more such APIs may include web service interfaces implemented using the Representational State Transfer (REST) architectural style. In such implementations, API calls may be encoded using Hypertext Transfer Protocol (HTTP) and JavaScript Object Notation (JSON) and / or Extensible Markup Language. Such API calls may be addressed to one or more Uniform Resource Locators (URLs) corresponding to API endpoints monitored by Transport Service 828. In some implementations, portions of HTTP communication may be encrypted for enhanced security. Alternatively (or additionally), in some implementations, one or more APIs of Transport Service 828 may be implemented as a .NET web API responding to HTTP posts pointing to a specific URL. Alternatively (or additionally), in some implementations, one or more APIs of Transport Service 828 may be implemented using simple file transfer protocol commands. Therefore, the Transport Service 828 API is not limited to any particular implementation.
[0109] The monitoring service 830 within the monitoring center environment 826 can be configured to control the overall logical settings and operation of the security system 800. Therefore, the monitoring service 830 can communicate and interact with the transmission service 828, the monitoring application 226, and various devices located at the monitoring location 204 via the network 820. In some embodiments, the monitoring service 830 can be configured to monitor data from multiple sources in response to events (e.g., intrusion events) and notify one or more of the monitoring applications 226 and / or 228 of the event when it is detected.
[0110] In some implementations, the monitoring service 830 may be additionally configured to maintain status information about the monitored location 204. This status information may indicate, for example, whether the monitored location 204 is secure or under threat. In some implementations, the monitoring service 830 may be configured to change the status information to indicate that the monitored location 204 is secure only upon receiving communication indicating a specific event (e.g., rather than simply due to the absence of any additional detected event). This feature prevents a "crash and smash" robbery (e.g., an intruder immediately destroying or disabling the monitoring equipment) from being successfully executed. Additionally, in some implementations, the monitoring service 830 may be configured to monitor one or more specific areas within the monitored location 204, such as one or more specific rooms or other different areas within and / or around the monitored location 204, and / or corresponding image capture devices deployed at the monitored location (e.g., Figure 8 One or more defined areas within the FOV of the cameras 202A and 202B shown.
[0111] Individual monitoring applications 226 in the monitoring center environment 822 can be configured to enable monitoring personnel to interact with corresponding computing devices to provide monitoring services for a given location (e.g., monitored location 204) and to perform various procedural operations in response to such interactions. For example, in some implementations, monitoring application 226 may control its host computing device to provide information to personnel operating the computing device regarding events detected at the monitored location (such as monitored location 204). Such events may include, for example, detected movement within a specific area of monitored location 204. In some embodiments, monitoring application 226 may enable monitoring device 210 to present video of events within individual event windows on the screen, and may further establish streaming connections with one or more cameras 202 at the monitored location, and enable monitoring device 210 to provide streaming video from such cameras 202 within a main viewer window and / or secondary viewer window on the screen, as well as allow audio communication between monitoring device 210 and cameras 202. Such streaming connections can be established, for example, using the Web Real-Time Communication (WebRTC) function of a browser on the monitoring device 210.
[0112] Application 228 of user device 214 may be configured to enable users to interact with their computing devices (e.g., their smartphones or personal computers) to access various services provided by security system 800 for their residences or other locations (e.g., monitoring location 204), and to perform various programmed operations in response to such interactions. For example, in some embodiments, application 228 may control user device 214 (e.g., smartphones or personal computers) to provide information to the user operating user device 214 regarding events detected at the monitoring location (such as monitoring location 204). Such events may include, for example, detected movement within a specific area of the monitored location 204. In some embodiments, application 228 may additionally or alternatively be configured to process input received from the user to activate or deactivate one or more devices located within monitoring location 204. Furthermore, application 228 may be additionally or alternatively configured to establish a streaming connection with one or more cameras 202 at the monitoring location, and to enable user device 214 to display streaming video from such cameras 202, as well as to allow audio communication between user device 214 and cameras 202. Such a streaming connection may be established, for example, using the Web Real-Time Communication (WebRTC) functionality of a browser on user device 214.
[0113] Now go to Figure 9 This schematically illustrates the computing system 900. (For example...) Figure 9As shown, the computing system 900 may include at least one processor 902, volatile memory 904, one or more interfaces 906, non-volatile memory 908, and interconnection mechanism 914. The non-volatile memory 908 may include executable code 910 and, as shown, may additionally include at least one data storage repository 912.
[0114] In some implementations, the non-volatile (non-transitory) memory 908 may include one or more read-only memory (ROM) chips; one or more hard disk drives or other magnetic or optical storage media; one or more solid-state drives (SSDs), such as flash drives or other solid-state storage media; and / or one or more hybrid magnetic SSDs. Further, in some implementations, code 910 stored in the non-volatile memory may include an operating system and one or more applications or programs configured to execute under the control of the operating system. In some implementations, code 910 may additionally or alternatively include dedicated firmware and embedded software executable independently of a commercially available operating system. Regardless of its configuration, execution of code 910 may produce manipulated data that can be stored as one or more data structures in a data repository 912. The data structures may have fields associated by their location within the data structure. Such associations can also be implemented by allocating storage for the fields in locations within memory that communicate the associations between the fields. However, other mechanisms may be used to establish associations between information in the fields of a data structure, including by using pointers, tags, or other mechanisms.
[0115] The processor 902 of the computing system 900 may be embodied by one or more processors configured to execute one or more executable instructions, such as a computer program specified by code 910, to control the operation of the computing system 900. Functions, operations, or sequences of operations may be hard-decoded into circuitry or soft-decoded by instructions stored in a memory device (e.g., volatile memory 904) and executed by circuitry. In some implementations, the processor 902 may be embodied as one or more application-specific integrated circuits (ASICs), microprocessors, digital signal processors (DSPs), graphics processing units (GPUs), neural processing units (NPUs), microcontrollers, field-programmable gate arrays (FPGAs), programmable logic arrays (PLAs), or multi-core processors.
[0116] Before executing code 910, processor 902 may copy code 910 from non-volatile memory 908 to volatile memory 904. In some implementations, volatile memory 904 may include one or more static or dynamic random access memory (RAM) chips and / or cache memory (e.g., memory disposed on a silicon die of processor 902). Volatile memory 904 may provide a faster response time than main memory such as non-volatile memory 908.
[0117] By executing code 910, processor 902 can control the operation of interface 906. Interface 906 may include a network interface. Such a network interface may include one or more physical interfaces (e.g., radio, Ethernet port, USB port, etc.), and a software stack including drivers and / or other code 910 configured to communicate with one or more physical interfaces to support one or more LAN, PAN, and / or WAN standard communication protocols. Such communication protocols may include, for example, TCP and UDP. Therefore, the network interface enables computing system 900 to access and communicate with other computing devices via a computer network.
[0118] Interface 906 may include one or more user interfaces. For example, in some implementations, user interface 906 may include user input and / or output devices (e.g., keypad, mouse, touchscreen, display, speaker, camera, accelerometer, biometric scanner, environmental sensor, etc.) and a software stack including drivers and / or other code 910 configured to communicate with the user input and / or output devices. Thus, user interface 906 enables computing system 900 to interact with a user to receive input and / or present output. The presented output may include, for example, one or more GUIs, which include one or more controls configured to display output and / or receive input. The received input may specify a value to be stored in data repository 912. The displayed output may indicate the value stored in data repository 912.
[0119] The various features of the computing system 900 described above can communicate with each other via interconnection mechanism 914. In some implementations, interconnection mechanism 914 may include a communication bus.
[0120] The following clauses describe exemplary methods, systems, and computer-readable media that embody various aspects of this disclosure.
[0121] Clause 1. A method comprising: providing a user interface via an application of a computing device to (i) play back a video in a first area of a screen and (ii) display a plurality of interactive elements corresponding to features detected in the video, the plurality of interactive elements being displayed in a second area of the screen different from the first area; determining via the application that the playback of the video has reached a first time position in the video corresponding to a first interactive element among the plurality of interactive elements displayed in the second area; and causing via the application to cause a change in the appearance of the first interactive element to visually distinguish the first interactive element from the other interactive elements among the plurality of interactive elements, the change being temporary such that, as the playback of the video progresses beyond the first time position, the appearance of the first interactive element reverts to the appearance it was displayed before reaching the first time position in the video.
[0122] Clause 2. The method described in Clause 1 further includes: determining, through the application, that the first interactive element has been selected; and, through the application and based at least in part on the selection of the first interactive element, causing the playback of the video to jump to the first time position.
[0123] Clause 3. The method according to Clause 1 further includes: determining, via the application, that the playback of the video has reached a second time position in the video corresponding to a second interactive element among a plurality of interactive elements displayed in the second area; and, in response to the playback of the video having reached the second time position, restoring via the application the appearance of the first interactive element to the appearance it displayed before reaching the first time position in the video.
[0124] Clause 4. The method according to Clause 3 further includes: determining, by the application, that the second interactive element is not currently displayed on the screen; and, at least in part based on the fact that the second interactive element is not currently displayed on the screen, causing a list of interactive elements, including at least the first interactive element and the second interactive element, to scroll to reveal the second interactive element.
[0125] Clause 5. The method according to Clause 4 further comprises: adjusting the relative position of the first interactive element in the second area by determining through the application that the user has not recently provided input, wherein scrolling of the list of interactive elements is based at least in part on the user not recently providing input.
[0126] Clause 6. The method according to Clause 3 further comprises: determining, via the application, that the second interactive element is not currently displayed in the second area; determining, via the application, that the user has provided input to adjust the relative position of the first interactive element in the second area; and, via the application and at least in part based on the user having provided input, preventing the list of interactive elements including at least the first and second interactive elements from scrolling to reveal the second interactive element.
[0127] Clause 7. The method according to Clause 1 further includes: after determining that the playback of the video has reached the first time position, determining by the application that the first interactive element is not currently displayed on the screen; and at least in part based on the fact that the first interactive element is not currently displayed on the screen, causing a list of interactive elements including at least the first interactive element to scroll to reveal the first interactive element.
[0128] Clause 8. The method according to Clause 7 further comprises: adjusting the relative position of the first interactive element in the second area by determining through the application that the user has not recently provided input; wherein scrolling of the list of interactive elements is based at least in part on the user not recently providing input.
[0129] Clause 9. A system comprising: one or more processors; and one or more computer-readable media encoded with instructions that, when executed by the one or more processors, cause the system to perform the method according to any one of Clauses 1 to 8.
[0130] Clause 10. One or more non-transitory computer-readable media encoded with instructions that, when executed by one or more processors of a system, cause the system to perform the method according to any one of Clauses 1 to 8.
[0131] Clause 11. A method comprising: receiving, via an application, first data representing a video of an event detected by a camera, second data representing at least a first feature detected in the video, and third data indicating a first time position in the video where the first feature was detected; via the application and using the first data, causing a device to play back at least a portion of the video in a first area of a screen; via the application and using the second data, causing the device to display a first user interface (UI) element indicating the first feature in a second area of the screen; via the application, determining that the playback of the video has reached the first time position; and via the application and at least in part based on the third data and that the playback of the video has reached the first time position, causing a change in the appearance of the first UI element to visually distinguish the first UI element from at least a second UI element displayed on the screen, the second UI element indicating the second feature detected in the video.
[0132] Clause 12. The method described in Clause 11 further includes: determining, via the application, that the first UI element has been selected; and via the application and based at least in part on the selection of the first UI element, causing the playback of the video to jump to the first time position.
[0133] Clause 13. The method according to Clause 11 further comprises: receiving, via the application, fourth data indicating the detection of the second feature in the video, and fifth data indicating a second time position in the video where the second feature was detected; via the application and using the fourth data, causing the device to display the second UI element together with the first UI element; via the application, determining that the playback of the video has reached the second time position; and via the application and at least in part based on the fifth data and that the playback of the video has reached the second time position, causing the device to change the appearance of the second UI element to visually distinguish the second UI element from at least the first UI element displayed on the screen.
[0134] Clause 14. The method according to Clause 13 further includes: determining, via the application, that the second UI element is not currently displayed on the screen; wherein causing the device to display the second UI element includes scrolling a list of UI elements including at least the first UI element and the second UI element to reveal the second UI element.
[0135] Clause 15. The method according to Clause 14 further comprises: adjusting the relative position of the first UI element in the second area by determining through the application that the user has not recently provided input; wherein scrolling of the list of UI elements is based at least in part on the user not recently providing input.
[0136] Clause 16. The method according to Clause 11 further includes: determining, via the application, that the first UI element is not currently displayed on the screen; wherein displaying the first UI element by the device includes scrolling a list of UI elements including at least the first UI element and the second UI element to reveal the first UI element.
[0137] Clause 17. The method according to Clause 16 further comprises: adjusting the relative position of the first UI element in the second area by determining through the application that the user has not recently provided input; wherein scrolling of the list of UI elements is based at least in part on the user not recently providing input.
[0138] Clause 18. The method according to Clause 11 further comprises: receiving, via the application, fourth data indicating the detection of the second feature in the video, and fifth data indicating a second time position within the video where the second feature was detected; determining via the application that playback of the video has reached the second time position; determining via the application that the second UI element associated with the second feature is not currently displayed in the second area; determining via the application that a user has provided input to adjust the relative position of the first UI element in the second area; and via the application and at least in part based on the user having provided the input, preventing a list of UI elements including at least the first UI element and the second UI element from scrolling to reveal the second UI element.
[0139] Clause 19. A system comprising: one or more processors; and one or more computer-readable media encoded with instructions that, when executed by the one or more processors, cause the system to perform the method according to any one of Clauses 11 to 18.
[0140] Clause 20. One or more non-transitory computer-readable media encoded with instructions that, when executed by one or more processors of a system, cause the system to perform the method pursuant to any one of Clauses 11 to 18.
[0141] Various inventive concepts may be embodied in one or more methods, and examples of such methods have been provided. Actions performed as part of a method can be ordered in any suitable manner. Therefore, examples can be constructed in which actions are performed in an order different from the order shown, and these examples may include actions performed simultaneously, even if these actions are shown as sequential in the illustrative examples.
[0142] The use of ordinal terms such as "first," "second," "third," etc., in claims to modify claim elements does not in itself imply any priority, precedence, or order of a claim element over another, or the chronological order of actions of a method. Such terms are merely labels to distinguish one claim element with a certain name from another element with the same name (but using ordinal terms).
[0143] The examples of methods and systems discussed herein are not limited to the details of the construction and arrangement of components described in the following description or illustrated in the figures. Methods and systems can be implemented in other examples and can be practiced or implemented in various ways. Examples of specific implementations provided herein are for illustrative purposes only and are not intended to be limiting. Specifically, actions, components, elements, and features discussed in relation to any one or more examples are not intended to exclude similar roles in any other example.
[0144] Furthermore, the wording and terminology used herein are for descriptive purposes and should not be considered restrictive. Any reference to an instance, component, element, or action of a system or method mentioned herein in the singular may cover multiple instances, and any reference to any instance, component, element, or action mentioned herein in the plural may cover only the singular instances. References in either the singular or plural form are not intended to limit the currently disclosed system or method, its components, actions, or elements.
[0145] As used herein, the terms “including / comprising,” “having,” “containing,” “involving,” and variations thereof are intended to cover the items listed thereafter and their equivalents, as well as additional items. References to “or” are to be interpreted as inclusive, such that any term described using “or” can refer to any one, more than one, or all of the terms described. Furthermore, in the event of any inconsistency in terminology between this document and any document incorporated herein by reference, the terminology used in the incorporated reference shall supplement the terminology used in this document; in the case of irreconcilable inconsistencies, the terminology used in this document shall prevail.
[0146] Several examples have been described in detail, and various modifications and improvements will readily occur to those skilled in the art. These modifications and improvements are intended to remain within the scope of this disclosure. Therefore, the above description is by way of example and is not intended to be limiting.
Claims
1. A method comprising: The application receives first data representing a video of an event detected by the camera, second data representing at least a first feature detected in the video, and third data indicating a first time location within the video where the first feature was detected. The device is made to play back at least a portion of the video in a first area of the screen using the application and the first data. The application uses the second data to cause the device to display a first user interface (UI) element indicating the first feature in a second area of the screen; The application determines that the video playback has reached the first time position; as well as The application, based at least in part on the third data and the fact that the playback of the video has reached the first time position, causes a change in the appearance of the first UI element to visually distinguish the first UI element from at least a second UI element displayed on the screen, the second UI element indicating a second feature detected in the video.
2. The method according to claim 1, further comprising: The application determines that the first UI element has been selected. as well as The video playback is redirected to the first time position via the application and at least in part based on the selection of the first UI element.
3. The method according to claim 1 or claim 2, further comprising: The application receives fourth data representing the second feature detected in the video, and fifth data indicating the second time position in the video where the second feature was detected. The device displays the second UI element and the first UI element together using the application and the fourth data; The application determines that the video playback has reached the second time position; as well as The device alters the appearance of the second UI element by means of the application and at least in part based on the fifth data and the fact that the playback of the video has reached the second time position, so as to visually distinguish the second UI element from at least the first UI element displayed on the screen.
4. The method of claim 3, further comprising: The application determines that the second UI element is not currently displayed on the screen; Displaying the second UI element by the device includes scrolling a list of UI elements, including at least the first UI element and the second UI element, to reveal the second UI element.
5. The method of claim 4, further comprising: The application determines the relative position of the first UI element within the second area based on whether the user has recently provided input. The scrolling of the list of UI elements is based at least in part on the fact that the user has not recently provided the input.
6. The method according to any one of claims 1 to 5, further comprising: The application determines that the first UI element is not currently displayed on the screen; Displaying the first UI element by the device includes scrolling a list of UI elements, including at least the first UI element and a second UI element, to reveal the first UI element.
7. The method of claim 6, further comprising: The application determines the relative position of the first UI element within the second area based on whether the user has recently provided input. The scrolling of the list of UI elements is based at least in part on the fact that the user has not recently provided the input.
8. The method according to any one of claims 1 to 7, further comprising: The application receives fourth data representing the second feature detected in the video, and fifth data indicating the second time position in the video where the second feature was detected. The application determines that the video playback has reached the second time position; The application determines that the second UI element associated with the second feature is not currently displayed in the second area; The application determines the user-provided input to adjust the relative position of the first UI element within the second area; as well as By means of the application and at least in part based on the input already provided by the user, the scrolling of the UI element list, which includes at least the first UI element and the second UI element, is avoided to reveal the second UI element.
9. A system comprising: One or more processors; as well as One or more computer-readable media encoded with instructions that, when executed by the one or more processors, cause the system to: The application receives first data representing a video of an event detected by the camera, second data representing at least a first feature detected in the video, and third data indicating a first time location within the video where the first feature was detected. The device is made to play back at least a portion of the video in a first area of the screen using the application and the first data. The application uses the second data to cause the device to display a first user interface (UI) element indicating the first feature in a second area of the screen; The application determines that the video playback has reached the first time position; as well as The application, based at least in part on the third data and the fact that the playback of the video has reached the first time position, causes a change in the appearance of the first UI element to visually distinguish the first UI element from at least a second UI element displayed on the screen, the second UI element indicating a second feature detected in the video.
10. The system of claim 9, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the first UI element has been selected; and The video playback is redirected to the first time position via the application and at least in part based on the selection of the first UI element.
11. The system of claim 9 or claim 10, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application receives fourth data representing the second feature detected in the video, and fifth data indicating the second time position in the video where the second feature was detected. The device displays the second UI element and the first UI element together using the application and the fourth data; The application determines that the video playback has reached the second time position; as well as The device alters the appearance of the second UI element by means of the application and at least in part based on the fifth data and the fact that the playback of the video has reached the second time position, so as to visually distinguish the second UI element from at least the first UI element displayed on the screen.
12. The system of claim 11, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the second UI element is not currently displayed on the screen; and The device displays the second UI element, at least in part, by scrolling a list of UI elements, including at least the first UI element and the second UI element, to reveal the second UI element.
13. The system of claim 12, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the user has not recently provided input in order to adjust the relative position of the first UI element within the second area; and This is at least in part based on the fact that the user has not recently provided the input to scroll the list of UI elements.
14. The system according to any one of claims 9 to 13, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the first UI element is not currently displayed on the screen; and The device displays the first UI element, at least in part, by scrolling a list of UI elements, including at least the first UI element and the second UI element, to reveal the first UI element.
15. The system of claim 14, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the user has not recently provided input in order to adjust the relative position of the first UI element within the second area; and This is at least in part based on the fact that the user has not recently provided the input to scroll the list of UI elements.
16. The system according to any one of claims 9 to 15, wherein the one or more computer-readable media are further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application receives fourth data representing the second feature detected in the video, and fifth data indicating the second time position in the video where the second feature was detected. The application determines that the video playback has reached the second time position; The application determines that the second UI element associated with the second feature is not currently displayed in the second area; The application determines the user-provided input to adjust the relative position of the first UI element within the second area; as well as By means of the application and at least in part based on the input already provided by the user, the scrolling of the UI element list, which includes at least the first UI element and the second UI element, is avoided to reveal the second UI element.
17. One or more non-transitory computer-readable media encoded with instructions that, when executed by one or more processors of a system, cause the system to: The application receives first data representing a video of an event detected by the camera, second data representing at least a first feature detected in the video, and third data indicating a first time location within the video where the first feature was detected. The device is made to play back at least a portion of the video in a first area of the screen using the application and the first data. The application uses the second data to cause the device to display a first user interface (UI) element indicating the first feature in a second area of the screen; The application determines that the video playback has reached the first time position; as well as The application, based at least in part on the third data and the fact that the playback of the video has reached the first time position, causes a change in the appearance of the first UI element to visually distinguish the first UI element from at least a second UI element displayed on the screen, the second UI element indicating a second feature detected in the video.
18. The one or more non-transitory computer-readable media of claim 17, further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the first UI element has been selected; and The video playback is redirected to the first time position via the application and at least in part based on the selection of the first UI element.
19. One or more non-transitory computer-readable media according to claim 17 or claim 18, further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application receives fourth data representing the second feature detected in the video, and fifth data indicating the second time position in the video where the second feature was detected. The device displays the second UI element and the first UI element together using the application and the fourth data; The application determines that the video playback has reached the second time position; as well as The device alters the appearance of the second UI element by means of the application and at least in part based on the fifth data and the fact that the playback of the video has reached the second time position, so as to visually distinguish the second UI element from at least the first UI element displayed on the screen.
20. The one or more non-transitory computer-readable media of claim 19, further encoded with additional instructions, which, when executed by the one or more processors, further cause the system to: The application determines that the second UI element is not currently displayed on the screen; and The device displays the second UI element, at least in part, by scrolling a list of UI elements, including at least the first UI element and the second UI element, to reveal the second UI element.
Citation Information
Patent Citations
Information processing device and information processing method
CN104135694A
UI interface interaction method, display device, server and vehicle
CN118131964A
Social interaction user interface for videos
US20190141402A1
User interface with metadata content elements for video navigation
US20220308742A1