Identify regions of interest in the imaging field of view
Through automatic image analysis and user selection to identify areas of interest, and only the image parts of these areas are processed and transmitted, the problems of low processing efficiency and network time delay in traditional systems are solved, and more efficient image reconstruction and motion detection are achieved.
Patent Information
- Application Number
- CN202280083631.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-03
- Filing Date
- 2022-11-16
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-11-16
AI Technical Summary
Traditional home or building surveillance systems need to process the entire image during motion detection, resulting in low processing efficiency, extended network time and difficulty in accurately specifying the region of interest.
Multiple areas of the scene are identified through automatic image analysis technology, and users are allowed to select or recommend areas of interest as motion detection areas, transmit and process only the image parts of these areas, and combine metadata for image reconstruction and event detection.
It improves processing efficiency and network delay performance, reduces unnecessary computing resource consumption, and achieves more efficient image reconstruction and motion detection.
Smart Images

Figure CN118525309B_ABST
Abstract
Description
Technical Field
[0001] Aspects of the technology described herein relate to image segmentation systems and methods, and more particularly, to systems and methods for identifying regions of interest. Background Art
[0002] Conventional home or building surveillance systems often use one or more image capture devices to capture images of the scene surrounding the home or building. Such surveillance systems can use the images to perform motion detection (e.g., by processing the images locally at the home or building and / or transmitting the captured images to a server). If motion is detected, the system can send an alert to the user and / or the user's device. Summary of the Invention
[0003] The present disclosure relates to techniques for identifying one or more regions of a scene for motion detection. In some embodiments, these techniques provide computerized methods, systems, and / or non-transitory computer-readable media that perform: determining a plurality of regions from an image of a scene using automated image analysis techniques; displaying the plurality of regions; receiving a user selection of one or more of the plurality of regions, the user selection indicating that the one or more regions are designated as motion detection zones; and storing the one or more designated motion detection zones for use in performing motion detection on one or more subsequent images of the scene.
[0004] In some embodiments, the techniques provide computerized methods, systems, and / or non-transitory computer-readable media that perform: determining a plurality of regions from an image of a scene using automated image analysis techniques; determining designations for the plurality of regions, wherein the designations indicate whether each of the plurality of regions is associated with triggering / non-triggering of motion detection; determining one or more of the plurality of regions as designated motion detection zones based on the designations for the plurality of regions; and storing the one or more designated motion detection zones for use in performing motion detection on one or more subsequent images of the scene.
[0005] In some embodiments, the techniques provide a computerized method, system, and / or non-transitory computer-readable medium that performs: receiving multiple regions of one or more images of a scene from a communication network, wherein the multiple regions are designated as image analysis zones; performing image analysis on the multiple regions to detect the presence of one or more events; and in response to detecting the presence of at least one event in one of the multiple regions designated as image analysis zones, sending an alert to the communication network, wherein the alert indicates the presence of an event in one of the image analysis zones.
[0006] In some embodiments, the techniques provide computerized methods, systems, and / or non-transitory computer-readable media that perform: receiving multiple regions of one or more images of a scene from a communication network, receiving metadata containing information associated with the multiple regions of the one or more images of the scene from the communication network; and reconstructing at least one of the one or more images of the scene using the multiple regions of the one or more images of the scene and the metadata.
[0007] The various embodiments described herein can provide advantages over conventional systems in terms of improved processing efficiency, processing speed, and / or network latency. For example, for image reconstruction, the techniques described herein enable a camera and / or system to transmit only the portion of an image within a region of interest to a processing device (e.g., a server or cloud) for image reconstruction. This results in significant improvements in processing efficiency, network latency, and processing speed compared to processing the entire image. Other advantages of the various embodiments described herein include easy-to-use tools, such as user interfaces, that allow users to easily define regions using automatic image segmentation techniques and / or recommended motion detection zone recommendations. These and other techniques for specifying regions of interest are further described in this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Further embodiments of the present disclosure and their features and advantages will become more apparent by reference to the description herein taken in conjunction with the accompanying drawings.The components in the figures are not necessarily to scale.
[0009] Figure 1 is a diagram of an example system for identifying motion detection zones, according to some embodiments.
[0010] Figure 2 is a flow chart describing an exemplary computerized method for identifying motion detection zones using automatic image analysis and user selection, in accordance with some embodiments.
[0011] Figure 3 is a flow chart describing an exemplary computerized method for automatically identifying motion detection zones in accordance with some embodiments.
[0012] Figure 4 is a flow chart describing an exemplary computerized method for detecting the presence of event(s) in one or more designated image analysis regions of a scene, according to some embodiments.
[0013] Figure 5 is an example image of a scene captured from an image capture device in accordance with some embodiments.
[0014] Figure 6 An example of multiple areas in a scene according to some embodiments is illustrated.
[0015] Figure 7 An example of multiple regions in a scene is illustrated, wherein each region is labeled with a class identification, according to some embodiments.
[0016] Figure 8 An illustrative implementation of a computer system that can be used to perform any aspects of the techniques and embodiments disclosed herein, according to some embodiments, is shown. DETAILED DESCRIPTION
[0017] For the purposes of promoting an understanding of the principles of the present disclosure, reference will now be made to the embodiments illustrated in the drawings, and specific language will be used to describe these embodiments. It will be understood, however, that no limitation of the invention's scope is intended thereby.
[0018] Traditional security systems typically use one or more image capture devices mounted on a property to capture images of the scene. The system can then transmit the captured images to a server for motion detection. If motion is detected, the server can then send an alert to the user's device.
[0019] The techniques and systems described herein provide an easy-to-use tool for allowing a user to define one or more regions of interest. The region(s) of interest can be used to reconstruct an image of a scene. One or more regions of interest can also or alternatively be designated for further image analysis, such as motion detection in the reconstructed image. Rather than using fixed or predetermined geometric shapes (e.g., rectangles), regions of interest can be designated based on their natural shapes, which can significantly reduce the number of pixels in the image that need to be processed by the system.
[0020] The techniques described herein can offer advantages over conventional systems in terms of improved processing efficiency, processing speed, and / or network latency. For example, for image reconstruction, the techniques described herein enable a camera and / or system to transmit only the portion of an image within a region of interest to a processing device (e.g., a server or cloud) for image reconstruction. Thus, in this configuration, only a portion of the captured image is used to reconstruct the region of interest, rather than the entire image. As another example, when the region(s) of interest are used for image analysis (e.g., motion detection), the system can transmit only portions of the image to a remote location (e.g., a server or cloud) to perform image analysis on those portions. This results in significant improvements in processing efficiency, network latency, and processing speed compared to processing the entire image.
[0021] Other advantages of the systems and methods described herein include easy-to-use tools, such as user interfaces, that allow users to easily define regions using automated image segmentation techniques and / or techniques for recommending motion detection zones. Because the shape of a region can be any shape determined by segmentation (e.g., and not limited to a rectangular bounding box), users can accurately specify regions of interest. Such techniques can also easily allow users to select regions of interest (e.g., without having to manually draw a free-form region using a mouse, as in some systems). These and other techniques for specifying regions of interest are further described in this disclosure.
[0022] In some embodiments, various techniques are described herein, including systems, computerized methods, and non-transient instructions, that allow a user to select regions of interest as designated image analysis zones (or zones), such as motion detection zones. In some embodiments, the system may allow the user to select designated image analysis zones at the semantic region level. For example, the system may determine multiple semantic regions from an image of a scene and display the multiple regions for the user to select / deselect as designated image analysis zones. A scene may include any surrounding environment of a house or building, or any structure to be monitored by the surveillance system. A scene may include outdoor or indoor areas, or a combination thereof. For example, a scene may include a street view in front of a house or building. A scene may also include a view of a house interior, such as a living room, bedroom, and / or other areas within the house. A scene may also include any area within a commercial building, such as a reception area, conference rooms, secure areas (e.g., vaults, control rooms), and the like.
[0023] Semantic regions in an image can include areas where pixels belong to semantically related objects. For example, semantic regions for a scene around a home can include a front porch, a path, a lawn, trees, decorative items around the house (e.g., planting boxes, flowers, etc.), a swimming pool, shrubs, garden furniture, etc. In some examples, the system can use automated image analysis techniques to determine multiple regions. For example, automated image analysis techniques can include performing semantic segmentation, which is configured to segment an image into multiple semantic regions. Once multiple segmented regions are displayed, the user can select / deselect (e.g., by clicking) these regions to designate / de-designate regions as image analysis areas.
[0024] In some embodiments, the system may allow the user to select designated image analysis zones at the sub-region level, where sub-regions may represent instances of objects in the image. For example, a semantic region may be a tree region, where the semantic region may include multiple sub-regions, each representing an instance of a tree (i.e., an individual tree). Similarly, a patio furniture region may include multiple sub-regions (instances) of patio furniture, and so on. In some embodiments, the system may perform instance segmentation on multiple segmented regions, associating each region with a corresponding category (e.g., tree, patio, furniture, front porch, swimming pool, etc.), and identifying one or more sub-regions (instances) for each region. Once determined, the system may display the sub-regions (instances) of the region, and the user may select or deselect each sub-region as a designated motion detection zone.
[0025] In some embodiments, the system can recommend image analysis regions to the user. For example, the system can display a score associated with one or more regions in the scene, where the score indicates the likelihood of the region being used as part of the ultimately designated image analysis region. The user can then select (or not select) (one or more) system-recommended regions to determine the (one or more) final image analysis regions. In some embodiments, the system can automatically designate regions as image analysis regions. In some embodiments, the system can designate a region as an image analysis region based on a category associated with the region (e.g., a sidewalk can be automatically designated as a possible region of interest for an image analysis region). In some embodiments, the system can automatically designate a region as an image analysis region based on previous activity in the region (e.g., if there is a lot of movement in a region, it can be designated as a region of interest for an image analysis region).
[0026] In some embodiments, a designated image analysis zone can be associated with a zone type. For example, a designated zone can be a delivery zone where packages can be delivered, a swimming pool area that includes a swimming pool (e.g., detecting motion to prevent children from entering the swimming pool area without adult supervision), an intruder zone (e.g., a window or front porch area), a pet zone (e.g., an area in the backyard), etc. Each zone type can be associated with a set of one or more monitoring parameters. For example, for an intruder zone, one or more monitoring parameters can include events to be detected for image analysis, such as motion. One or more monitoring parameters can also include the time of day for detecting events. For example, for an intruder zone, the time for detecting motion events can be 24 hours / 7 days a week, evening hours / 7 days a week, etc. For a delivery zone, the time for detecting motion events can be normal working hours. Therefore, outside of normal working hours, the system can be configured to not detect any events in the delivery zone, resulting in further reduction in network bandwidth usage and computing power.
[0027] In some embodiments, once a user has selected an area as a designated motion detection zone, the user can also designate the zone as having a zone type as described above. In some embodiments, the system can determine the zone type based on how the user reacts to an alert for that zone. For example, if the system is configured to provide an alert via a call or text message to the user's device upon detection of motion in the zone, and in response to the alert, the user dispatches the police from the user's device (e.g., via a call to the emergency number), the system can designate the designated motion detection zone as an intruder zone. In another example, if the system is configured to provide an alert via a call to the user's device upon detection of motion in the designated motion detection zone, but the user does not answer the call, the system can designate the motion detection zone as non-emergency. Thus, the techniques described herein also allow the zone type of a given designated motion detection zone to be initially determined and / or updated over time based on future user responses.
[0028] In some embodiments, each zone type can be associated with one or more monitoring parameters. The monitoring parameter(s) associated with a zone type can be predetermined. For example, for an intruder zone, one or more monitoring parameters may include motion detection during all hours on a 24 / 7 basis, whereas one or more monitoring parameters for a delivery zone may include motion detection only during that day. In some embodiments, the system can determine / update the monitoring parameter(s) for different zones based on previous activity in those zones. For example, if most motion triggers in a delivery zone are detected during that day, the system can determine the monitoring parameters for the delivery zone to include motion detection only during that day. In another example, a swimming pool zone may be mostly active in the afternoons during the summer (with frequent motion). Thus, the system can determine the monitoring parameters for the swimming pool zone to include motion detection during the mornings and evenings.
[0029] In some embodiments, the system may capture one or more subsequent images of a scene using an image capture device and transmit a portion of each image to a server for processing, rather than transmitting the entire image. The system may determine the portion of the image to transmit based on a designated image analysis region. When transmitting a portion of the image(s), the system may only transmit the pixels of the image(s) within the designated region. Additionally, the system may transmit metadata describing the designated region. For example, the metadata may include information about the pixels within the designated region, such as their relative position within the image. In some embodiments, the metadata may include any of the type of designated region, one or more monitoring parameters associated with the designated region, or a combination thereof. Additionally or alternatively, the metadata may include extracted features that define the designated region. For example, the metadata may include motion flow, color histogram, image pixel density, and / or the direction of motion of a group of pixels. At the server, the system may reconstruct the image using the transmitted portion of the captured image(s) and the metadata.
[0030] The system can also detect one or more events in a reconstructed image of one or more designated areas based on the monitoring parameter(s) associated with each designated area. In the embodiments described herein, the reconstructed image is much more compressed than the entire captured image because the reconstructed image only includes pixels in the designated area. In some embodiments, the system can send an alert to the user device in response to detecting an event. The alert can also include a type, and the system can send the alert based on the alert type. In some examples, the alert type can include a communication method (e.g., a call, text message, or other notification method) and / or a time for delivery (e.g., immediately; when the user is available; at a fixed time of day; or on certain days).
[0031] As described above, the techniques described herein can provide advantages over conventional systems in terms of improved processing speed and network latency. For example, the techniques described herein for designating image analysis zones (e.g., motion detection zones) enable the system to transmit only portions of an image within the designated image analysis zones and perform image analysis on only those portions. This results in significant improvements in network latency and processing speed compared to transmitting and processing the entire image. Furthermore, assigning different types of designated zones for image analysis, allowing different monitoring parameters (or parameters) to be associated with different zones, and having different types of alerts to be transmitted enables the system to transmit and process captured image data in the most efficient manner, conserving additional network bandwidth and computing resources. For example, allowing users to define delivery zones that detect motion only during business hours reduces wasted network and computing resource utilization during non-business hours. Additionally, the system provides an easy-to-use tool for designating image analysis zones in a semi-automatic or fully automated manner.
[0032] The techniques described herein can generally be used to identify regions of interest, with exemplary applications being the use of these regions in image reconstruction and motion detection from captured images. In motion detection applications, the system can identify designated motion detection zones so that the system can perform motion detection in these designated zones. Thus, without limiting the scope of the present disclosure, various embodiments are further described using examples of designated motion detection zones. Motion detection can include monitoring changes in the state of any object in a scene. Non-limiting examples of motion detection include the swaying of a tree, the walking of a person, the flight of a bird, the passing of a car, the movement of leaves and / or any other object into / out of a scene, and the like.
[0033] Although various embodiments have been described, it will be apparent to those skilled in the art that many more embodiments and implementations are possible. Therefore, the embodiments described herein are examples, not the only possible embodiments and implementations. Furthermore, the advantages described above are not necessarily the only advantages, and it is not necessarily expected that all described advantages will be achieved by every embodiment.
[0034] Figure 1 is a diagram of an example system for identifying regions of interest (e.g., motion detection zones) according to some embodiments. System 100 may include an image capture device (e.g., 102) configured to monitor a scene. Figure 1 As shown, camera 102 may be installed to monitor the front of a house. System 100 may also include one or more processing devices, such as server 108, user device 110, and / or one or more processors on the camera. System 100 may also include one or more communication networks (e.g., 104, 106) that provide communication between the one or more processing devices and image capture devices. For example, in a home surveillance scenario, system 100 may include home network 104 that establishes communication between user device 110 and camera 102. Home network 104 may also be configured to establish communication between camera 102 and the Internet outside the home. In some embodiments, system 100 may also include communication network 106 that establishes communication between camera 102 and a server 108 on the Internet via communication network 104. Additionally or alternatively, user device 110 may also communicate with server 108 on the Internet via communication network 106. In some embodiments, user device 110 may also communicate with image capture devices (e.g., 102) via home network 104. Each communication network 104, 106 may be any suitable network, such as a wired, wireless, mesh network, or any other suitable network.While several network configurations are shown, it should be understood that variations of these configurations are possible.
[0035] Further references Figure 1, system 100 can perform various operations to identify one or more motion detection zones. For example, operations can be performed within camera 102 or on user device 110 or any other device to analyze the image of the scene and identify one or more designated motion detection zones. Once one or more designated motion detection zones are identified, the system can store these zones for later use. In some embodiments, the system can store the designated motion detection zones on storage device 120 on camera 102. Alternatively and / or in addition, the designated motion detection zones can be stored on user device (e.g., 110) or on a server (e.g., 108).
[0036] In some techniques described herein, the system 100 may allow a user to select a designated motion detection zone via a user interface. In some embodiments, the system may allow a user to select a designated motion detection zone at the semantic region level. For example, the system may determine multiple semantic regions from an image of a scene and display the multiple regions in a user interface for the user to select / unselect as designated motion detection zones. Various techniques may be implemented in the user interface to allow the user to select / unselect regions of interest (e.g., designate / undesignate as motion detection zones). For example, multiple regions may be displayed as user-selectable tiles that may be distinguished by graphical features, such as color, texture, or other graphical representations. Each tile may be toggled between "selected" and "unselected" by the user clicking on the tile. Other implementations of user selection / unselection of regions are possible.
[0037] In some examples, the system can use automated image analysis techniques to determine the multiple regions. For example, the automated image analysis techniques can include performing semantic segmentation, which is configured to segment the image into multiple semantic regions. Examples of semantic regions for a scene around a home can include a front porch, a road, a lawn, trees, decorative items around the house (e.g., planting boxes, flowers, etc.), a swimming pool, shrubs, and patio furniture. Once the multiple segmented regions are displayed, the user can select / deselect (e.g., by clicking) these regions to designate / de-designate them as motion detection zones.
[0038] In some embodiments, the system may include a user interface that allows a user to select designated motion detection zones at the sub-region level, where sub-regions can represent instances of objects in an image. For example, a tree region may include multiple instances of a tree, each representing an individual tree. A patio furniture region may include multiple instances of furniture. The system may perform instance segmentation on the multiple regions to associate each region with a corresponding category (e.g., tree, patio, furniture, front porch, swimming pool, etc.), and identify one or more sub-regions (instances) for each semantic region. Once the instances are identified, they may be displayed as sub-regions in the user interface, and the user may select / deselect each sub-region as a designated motion detection zone in a manner similar to selecting / deselecting regions as described above.
[0039] In some embodiments, the system can recommend to the user the designation of a motion detection zone for a region or sub-region (instance). For example, the system can display a score associated with a region, where the score indicates the likelihood of the region being used as part of a ultimately designated motion detection zone. The user can select a system-recommended region by clicking on it in the user interface. In some embodiments, the system can automatically designate a region as a motion detection zone based on a category associated with the region. In some embodiments, the system can automatically designate a region as a motion detection zone based on previous activity in the region.
[0040] In some embodiments, the system may capture one or more subsequent images of a scene using an image capture device (e.g., 102) and transmit a portion of each image to a server (e.g., 108) for processing, rather than transmitting the entire image. The system may determine the portion of the image to transmit based on the designated motion detection zone. When transmitting the portion of the image(s), the system may only transmit the pixels of the image(s) that are within the designated motion detection zone(s). Additionally, the system may transmit metadata describing the designated motion detection zone(s). For example, the metadata may include the location of the pixels in the designated motion detection zone(s) relative to the image(s) of the scene. Additionally, the metadata may include a zone type assigned to each designated motion detection zone, one or more monitoring parameters associated with each zone type, and / or an alert type associated with each designated motion detection zone, or a combination thereof.
[0041] At the server (e.g., 108), the system can reconstruct an image using the transmitted portion(s) of the captured image(s) and the metadata. The system can also detect one or more events in the reconstructed image in one or more designated motion detection zones. In some embodiments, the system can send an alert to a user device (e.g., 110) in response to detecting an event. The system can send an alert based on the type of alert. For example, the type of alert that can be included in the metadata sent to the server can include a communication means (e.g., call, text, or other notification means) and a time for delivery (e.g., immediately, when the user is available, or at a fixed time of day, or on certain days). Reference Figures 2 to 7 Various embodiments that may be implemented in system 100 are described in further detail.
[0042] Figure 2 is a flow chart describing an exemplary computerized method for identifying motion detection zones using automatic image analysis and user selection according to some embodiments. In some embodiments, the method 200 for identifying motion detection zones may be performed in the system 100 ( Figure 1 ), for example, on the user device 110. In other embodiments, the method 200 may be implemented in a camera (e.g., Figure 1 In some embodiments, the system may be implemented in a processor on a camera (e.g., 102). In this case, the camera may have a processor. The camera may additionally have a built-in display. Alternatively, the system may have an external display communicatively coupled to the camera. In some embodiments, the display may be a touch screen configured to display multiple regions and receive user selections. In some examples used to describe the technology herein, a method 200 for identifying one or more regions of a scene for motion detection may include determining multiple regions from an image of the scene using automatic image analysis techniques (act 202). Each region may represent an area of the scene that a user may be interested in tracking motion, reconstructing images, and / or performing other image analysis. The image of the scene may be captured from an image capture device (e.g., ( Figure 1 The camera 102) can be mounted to track the motion in the scene of interest to the user. Figure 1 In the example of FIG, camera 102 may be mounted in front of a house and configured to capture the scene in front of the house. Figure 5 An example image of a scene in front of a house is shown in .
[0043] In some embodiments, the image of the scene may be transmitted from the image capture device (e.g., camera 102) via the communication network 104 to the user device 110 for processing. In this case, the user device may generate the plurality of regions in act 202. In other embodiments, the image of the scene may be transmitted via the communication network 106 to the server 108 for processing. In this case, the server (e.g., Figure 1 The server 108 may receive the image scene and generate a plurality of regions from the image scene. The server 108 may transmit the generated plurality of regions to the user device 110 via the communication network 106, wherein the user device may receive the plurality of regions generated by the server in action 202. In some embodiments, the camera (e.g., Figure 1 The user device may receive the plurality of regions generated by the camera in act 202 ) by the camera.
[0044] Therefore, the method for identifying a motion detection zone may include: Figure 1 )) displays multiple areas (action 204). Figure 6 An example of multiple regions in a scene obtained from act 202 according to some embodiments is illustrated. For example, a user device (e.g., Figure 1 110) can be displayed from action 202 ( Figure 2 ) to obtain multiple regions. These multiple regions can be displayed in various ways. For example, Figure 6 As shown, region 602 may be represented by a plurality of pixels including a region of a first color (e.g., green) and a region 606 of a second color (e.g., yellow). It should be understood that the colors of each region may be varied so that a user can visually distinguish between the multiple regions. Alternatively, other colors or grayscales may be used to represent the multiple regions. In some embodiments, other representations such as bounding boxes, textures, graphical symbols, and / or any combination thereof may also represent the multiple regions. Figure 6 As shown, multiple areas (e.g., 602-610) can each be from Figure 5 For example, regions 602, 604 may represent a scene (e.g., Figure 5 ) in the image; area 606 may represent the path in front of the house; and areas 608, 610 may represent vases in the front porch area.
[0045] In some embodiments, determining the multiple regions (e.g., in act 202) may utilize semantic segmentation techniques. Semantic segmentation may involve segmenting an image into multiple semantic regions. Semantic regions may include semantically related pixels or portions of an image. For example, semantic regions may include the background of a scene, the foreground of a scene, a person, a path, a tree, a porch, a swimming pool, a street, a house, and so on. In semantic segmentation, each pixel in an image may be classified into a corresponding region, and pixels in semantically related regions of the image may be classified into the same region. In some embodiments, various image semantic segmentation techniques may be used to generate multiple regions from an image of a scene. For example, a deep machine learning model, such as a neural network model, may be pre-trained and utilized. An input image may be provided to the pre-trained machine learning model, which is configured to output segmentation results using the input image. A training set comprising multiple training images may be used as training data. The training data may also include gold standard data, which includes gold standard semantic regions for each training image in the training data. The machine learning model may be obtained using an appropriate training method (e.g., gradient descent). In some examples, the methods that can be used are described in S. Ren et al., “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” arXiv:1506.01497, 2015 (https: / / arxiv.org / abs / 1506.01497), which is incorporated herein by reference in its entirety.
[0046] return Figure 2 , the method 200 for identifying a motion detection zone may also include receiving one or more user selections indicating a designation of a motion detection zone (act 206). Figure 6 , the user can click / unclick each segmented area (e.g., 602-610) to indicate whether one or more of the displayed multiple areas should be designated as motion detection zones. In some embodiments, the segmented areas can be overlaid on the image of the scene so that the user can identify these areas (e.g., sidewalk, front porch, tree, etc.). Figure 6 As shown, each region generally has an irregular shape that represents or corresponds to the natural shape of the semantic region, rather than a rectangular shape (e.g., a bounding box). This allows the region to be accurately represented in the user interface, allowing the user to identify and accurately determine whether the region should be designated as a motion detection zone. As will be further described in this disclosure, representing the semantic region in its natural shape allows the system to process only pixels within that region, compared to processing the region in a bounding box, which results in saving the processing of pixels in other regions that are not of interest to the user.
[0047] exist Figure 6 In the example of Figure 1 In the user interface shown in FIG110 , a user can select an area, such as 606 (front porch), as a designated motion detection zone. Conversely, the user can skip (or not select) other areas where false positives might be generated, such as 602 and 604 (trees). For example, wind might cause a tree to move, and motion events detected based on the tree's movement might be false positives. In some examples, the user can also select areas 608 and 610 (sides of the porch) as non-designated motion detection zones because they contain the structure or furnishings of the house, which are unlikely to be target areas for motion detection.
[0048] return Figure 2 Once the user selection is received (at act 206), method 200 may proceed to storing the user-specified motion detection zones (act 208). In some embodiments, each user-specified motion detection zone may be stored in any suitable representation, such as a bounding box, an outline, a plurality of pixels comprising the zone, or a combination thereof. In some embodiments, the specified motion detection zones may be stored on an image capture device (e.g., Figure 1 In a storage medium (eg, 120) of 102) in a user device (eg, Figure 1 In an example of performing actions 202-206 of 110), the user device may transmit the user-specified motion detection zone to the image capture device (e.g., Figure 1 102 ) to be stored in, for example, a storage device (eg, 120 ) installed in a camera (eg, 102 ).
[0049] In various embodiments, each of the multiple regions obtained from act 202 may be associated with one of multiple categories. In some embodiments, each semantic region may also include multiple subregions, each representing an instance of an object in the image. For example, a tree region may include multiple instances, each of which is a subregion and represents a corresponding tree. Therefore, method 200 may perform both semantic segmentation and instance segmentation in act 202. Semantic segmentation has been described previously. Additionally, method 200 may perform instance segmentation on the multiple semantic regions to associate each region with a corresponding category. Instance segmentation is the process by which the system understands each semantic region and assigns a unique identifier (ID) to each region (such as a tree, front porch, swimming pool, sidewalk, etc.). Instance segmentation may also determine one or more subregions (instances of an object) for each semantic region. In the above example, instance segmentation may associate a TREE_ID with the tree region and determine one or more subregions whose TREE_ID may have different values, for example, TREE_ID = 1 for the first tree and TREE_ID = 2 for the second tree. Various instance segmentation techniques may be used. For example, methods that can be used are described in Romera-Paredes B., Torr PHS, “Recurrent Instance Segmentation,” In: Leibe B., Matas J., Sebe N., Welling M. (eds) Computer Vision – ECCV 2016. Lecture Notes in Computer Science, vol9910, pp. 312-319. Springer, Cham., which is incorporated herein by reference in its entirety.
[0050] Figure 7 An example of multiple regions or sub-regions (instances) in a scene according to some embodiments is illustrated, wherein each region or sub-region is labeled with a category. Figure 7 As shown, Figure 6 Region 602 in can be segmented into sub-regions 702-1 and 702-2, each sub-region representing a corresponding instance of the detected object. In this example, 702-1 and 702-2 are instances of potted plants. Region 704 is classified as a potted plant from 604; and regions 708, 710 are classified as vases. In the user interface, each segmented region can also be displayed and selected / unselected individually at the instance level. For example, the system can display one or more instances of each of the multiple regions, and the user's selection of one or more regions from the multiple regions can also include the user's selection of one or more instances of the corresponding region, wherein the user's selection of the instance indicates that the instance of the corresponding region is designated as a motion detection zone.
[0051] In some embodiments, method 200 may additionally generate information associated with each region (e.g., in act 202). This additional information associated with each region can serve as a guide or recommendation for the user to designate a motion detection zone. For example, the system can recommend a region to the user as a designated motion detection zone by displaying (e.g., in act 204) additional information associated with the region, where this information may indicate the likelihood of the region being used as part of the ultimately designated motion detection zone. For example, a region associated with a low score value may indicate that the region is unlikely to be a motion detection zone. A region associated with a higher score value may indicate that the region is likely to be a motion detection zone. In some embodiments, the information indicating likelihood may be presented in other forms, such as a graphical representation. For example, a graphical symbol (e.g., a circle, bar, star, etc.) may be used to indicate the likelihood of a region being a motion detection zone.
[0052] In some embodiments, the determination of information indicating the likelihood of the corresponding region being used as part of the ultimately designated motion detection zone can be based on the corresponding category associated with the corresponding region. In a non-limiting example, in response to the region being classified as a tree, a low score can be assigned to indicate that the tree region is unlikely to be a motion detection zone. In another non-limiting example, in response to the region being classified as a front porch, a higher score can be assigned to indicate that the corresponding region is likely a motion detection zone. In some embodiments, in act 204, method 200 can display a score value or graphical representation associated with each region along with the plurality of regions. For example, the score value or graphical representation can be overlaid on each associated region.
[0053] In some embodiments, if at least one of the plurality of regions is associated with the first category, method 200 may determine that the at least one of the plurality of regions should not be designated as a motion detection zone, and not display the at least one of the plurality of regions. For example, method 200 may classify the region as a tree region (e.g., in act 202). Method 200 may determine that the tree region should not be designated as a motion detection zone, thereby not displaying the tree region (in act 204). In other words, the determined tree region may be automatically removed from the display to the user. Thus, the tree region will not be designated as a motion detection zone. The methods described herein may also be used to avoid other areas that may result in false positives, such as bushes (which may move with the wind), roads (with traffic movement), and the sky (with birds flying), by not displaying them to the user.
[0054] Additionally and / or alternatively, the method 200 may be based on referring to Figure 3The various techniques described herein automatically select designated motion detection zones and display the recommended selection on the display. The system can allow the user to override the system's recommended designated motion detection zones by allowing the user to deselect them. Similarly, the system can designate a region as a non-motion zone and display the non-motion zone as unselected. The system can allow the user to override the designated motion detection zone by selecting the zone as a non-designated motion zone. The various techniques described herein for user selection can be implemented at the region (e.g., semantic region) and sub-region (e.g., object instance) levels.
[0055] Further references Figure 2 , method 200 may further include using the stored region of interest (eg, the designated motion detection region obtained from act 208) in various applications. For example, method 200 may obtain a stored region of interest from an image capture device (eg, Figure 1 102 ) obtains one or more subsequent images of the scene (act 210 ). Method 200 may then transmit the captured portion of the image to a server (eg, Figure 1 108 ) (act 212 ). The portion that is transmitted may include a user-specified motion detection region in the image, rather than the entire image. For example, the method 200 may transmit only pixels in the region of interest (e.g., the designated motion detection region) to the communication network 106 , while leaving out pixels in the image in other regions (e.g., non-designated motion detection regions). In a non-limiting example, an image mask may be used that includes pixels in the designated motion detection region, with zero padding in other regions. In another non-limiting example, a plurality of pixels comprising the designated motion detection region may be transmitted to the communication network along with metadata describing the region of interest (e.g., the designated motion detection region). For example, the metadata may include the position of the pixel in each region of interest relative to the captured image.
[0056] Additionally and / or alternatively, the metadata may also include image dimensions or other information about the region of interest. In some embodiments, the metadata may also include one or more monitoring parameters for a given designated motion detection zone. In some embodiments, the metadata may additionally include an alert type for a given designated motion detection zone. Figure 4 Describes the details of the detection event in the specified motion detection zone.
[0057] Further references Figure 2 The method may further include transmitting a signal from a communication network (e.g., from a mobile station) to a designated motion detection zone if one or more events are detected in the designated motion detection zone in one or more subsequent images. Figure 1 The server 108 receives an alert (action 214). Thus, the alert may indicate an event detected in at least one designated motion detection zone. Figure 2 .
[0058] In some embodiments, method 200 may include receiving a first monitoring parameter for a first area of the plurality of areas; transmitting the first monitoring parameter to a communication network; and receiving a first alert from the communication network, the first alert indicating detection of a first event in the first area in one or more subsequent images of the scene based on the first monitoring parameter. In some examples, each designated motion detection zone may be associated with a corresponding type of zone. For example, the zone types for the designated motion detection zones may include a delivery zone, an intruder zone, a sidewalk, a street, a swimming pool zone, and the like. In some embodiments, when a user selects a designated motion detection zone (e.g., in act 206), the user may also assign a zone type to the selected zone via a user interface.
[0059] In some embodiments, a zone type can be associated with one or more monitoring parameters that indicate what events are to be detected for that zone and how they are to be detected. For example, the one or more monitoring parameters may include the events to be detected and / or the times to be used to detect the events. For example, a delivery zone may be associated with motion detection during normal operating hours for a delivery service. A swimming pool zone may be associated with motion detection during non-play time, which can be any time of day except the afternoon. Thus, for a delivery zone, method 200 may receive a first alert (e.g., in act 214) indicating that a delivery was detected during delivery operating hours. Similarly, method 200 may receive a second alert (e.g., in act 214) indicating that motion was detected in the swimming pool zone during non-play time.
[0060] In some embodiments, different types of alerts can be associated with corresponding ones of the multiple zones based on the type of motion detection zone designated. In other words, for a given designated motion detection zone, different types of alerts can be received in response to an event detected in that zone. In some embodiments, the alert type can include the form and / or time of delivery. For example, the alert can be received via a call, text message, or other electronic means to the user device. Alerts can also be delivered to the user device at different times. For example, if the motion detection zone is an intruder zone (e.g., a front porch area), an event detected in that zone can be considered urgent in nature. After detecting motion in the intruder zone, the user device can immediately receive a call. In another example, if the motion detection zone is a delivery zone, such as a garage door, an event detected in that zone can be considered non-urgent. Thus, after detecting an event in the delivery zone, the user device can receive a text message at or after the delivery is detected.
[0061] Figure 3is a flow chart illustrating an exemplary computerized method for automatically identifying motion detection zones according to some embodiments. The method 300 for identifying motion detection zones may be implemented in the system 100 ( Figure 1 ), for example, on user device 110. In other embodiments, method 300 may be implemented in a camera (e.g., Figure 1 Method 300 may be similar to method 200, except that method 200 may automatically determine the designated motion detection zone without receiving a user selection. In some embodiments, method 300 may be implemented in the same manner as method 200 ( Figure 2 ), determining a plurality of regions from the image scene (act 302) in a manner similar to act 202 in [ 300 ]. For example, method 300 may include performing semantic segmentation to segment the image of the scene into a plurality of regions; and performing instance segmentation on the plurality of regions to associate each region with a corresponding category from a plurality of categories. Method 300 may also determine one or more subregions of each of the plurality of regions, where the subregions may represent instances of objects in the image.
[0062] In some embodiments, method 300 may determine the designation of a plurality of regions (Act 304), wherein the designation indicates whether each of the plurality of regions is associated with triggering / non-triggering of motion detection. Furthermore, method 300 may determine one or more of the plurality of regions as designated motion detection zones based on the designation of the plurality of regions (Act 306). Methods for determining the designation of the plurality of regions and the designated motion detection zones in Acts 304 and 306 are described in further detail herein.
[0063] In some embodiments, determining the designation of multiple regions in act 304 may include determining the designation of each of the multiple regions based on an association of the region's category with an indication that motion detection is triggered / not triggered. For example, each category from instance segmentation may be associated with an indication that motion detection is triggered / not triggered. In a non-limiting example, a region classified as a tree may be associated with an indication that motion detection is not triggered because trees can easily trigger motion detection when windy, wherein such detected motion may be a false positive. In another example, a region classified as a front porch may be associated with an indication that motion detection is triggered because detected motion in this region may indicate the arrival or delivery of an intruder or visitor. Subsequently, in act 306, if the region is associated with an indication that motion detection is triggered, method 300 may determine the region as a designated motion detection zone. Conversely, if the region is associated with an indication that motion detection is not triggered, method 300 may determine the region as a non-designated motion detection zone.
[0064] In some embodiments, the system may designate one or more motion detection zones based on previous activity around a house. For example, in the example of action 304 above, the designation of zones to indicate whether a zone is associated with triggering / not triggering motion detection may be implemented based on previous motion detection results around the scene, where these previous motion detection results can be monitored and used in a number of ways. In some embodiments, the system may monitor motion detection results for specific areas within the scene. For example, the system may monitor the frequency of alarms being triggered in a specific area (e.g., the front porch). If multiple alarms have been triggered in that specific area in the past, the system may associate that area with triggering motion detection. In other examples, if no alarms have been received in a specific area (e.g., no motion has been detected) for an extended period of time, the system may associate that area with not triggering motion detection. This could be a scenario where a front porch, typically associated with a delivery service, has not received any deliveries because all deliveries are typically made to the house's side door. In this case, the system can automatically learn past activity in the area around the house and use this information to associate that area with triggering / not triggering motion detection.
[0065] In some embodiments, the system can also determine the zone type associated with a designated motion detection zone based on previous activity in the zone. For example, when an alert associated with the designated motion detection zone includes a call or text message to a user's device, where the alert also causes the user to dispatch the police from the user's device (e.g., via a call to the emergency number), the system can determine that the designated motion detection zone is an intruder zone. In another example, when an alert associated with the designated motion detection zone includes a call to the user's device, where the user never answers the call, the system can determine that the designated motion detection zone is a non-emergency zone. Thus, the techniques described herein also allow the zone type of a given designated motion detection zone to be updated over time.
[0066] Additionally and / or alternatively, the system may also automatically determine / update the monitoring parameter(s) associated with the zone using information about previous activity around the premises. For example, for an intruder zone, the associated monitoring parameters may include motion detection during all time periods on a 24 / 7 basis, wherein the monitoring parameters for a delivery zone may include motion detection only during that day. These monitoring parameters for different zones may be determined based on previous activity for the zones. For example, if most motion triggers in a delivery zone are detected during that day, the system may determine the associated monitoring parameters for the delivery zone to include motion detection only during that day. In another example, a swimming pool zone may be mostly active in the afternoons during the summer (there is frequent motion). Thereby, the system may determine the monitoring parameters for the swimming pool zone to include motion detection during the mornings and evenings. Further Reference Figure 3, other actions (eg, 308 , 310 , 312 , 314 ) may be performed in a similar manner to actions 208 , 210 , 212 , 214 , respectively.
[0067] Figure 4 is a flow chart illustrating an exemplary computerized method for detecting event(s) in one or more designated image analysis zones of a scene according to some embodiments. The method 400 for detecting one or more events in one or more subsequent images of a scene and sending an alert may be performed on a server (e.g., Figure 1 Method 400 may include receiving a plurality of regions of one or more images of a scene (act 402). The images may be captured from an image capture device (e.g., Figure 1 102) capture, wherein a plurality of areas can be designated as motion detection areas. For example, referring to Figure 2 and Figure 3 Various embodiments of specifying motion detection zones are described. These specified motion detection zones may be stored in an image capture device (e.g., Figure 1 When the image capture device captures one or more subsequent images, the plurality of regions designated as motion detection zones in the one or more images may be transmitted to a server (e.g., Figure 1 108) for processing.
[0068] In some embodiments, the multiple regions received in act 402 may be portions of an image, wherein the portions of the image represent one or more designated motion detection zones. In some embodiments, method 400 may further include performing image analysis to detect event(s) in the designated motion detection zones (act 408). As previously described in this disclosure, one or more monitoring parameters may be associated with each designated motion detection zone. For example, the monitoring parameters may include the type of event to be detected, such as motion. The monitoring parameters may also include a time or duration for which the event needs to be detected. In non-limiting examples, various methods may be used to detect motion in the designated motion detection zones. For example, a method may be used to detect the presence of motion based on the difference between two or more consecutive images, wherein changes in pixels in the designated motion detection zones of the two or more consecutive images may indicate the presence of motion. In other embodiments, a machine learning model may be trained and used to detect the presence of motion in an image. For example, methods that can be used are described in R. Cutler and LSDavis, "Robust real-time periodic motion detection, analysis, and applications," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 22, no. 8, pp. 781-796, Aug. 2000, which is incorporated herein by reference in its entirety.
[0069] In some embodiments, method 400 may optionally include act 404 of receiving metadata associated with each of the plurality of zones. For example, as previously described, the metadata associated with the zones may include the location of pixels relative to the captured image in each designated motion detection zone. Additionally, the metadata may also include image dimensions or other information about the designated motion detection zone. As previously described, the metadata may also include the type of designated motion detection zone, such as an intruder zone, a swimming pool zone, a delivery zone, etc. Thus, method 400 may detect events in a given designated motion detection zone based on the type of the motion detection zone (e.g., in act 408), wherein the detection of the event may be performed based on one or more monitoring parameters associated with the zone as previously described in the present disclosure. In some embodiments, the one or more monitoring parameters for each designated detection zone may include monitoring parameters transmitted over a communication network (e.g., Figure 1In other embodiments, one or more monitoring parameters may be stored in a storage device of the system and associated with the type of motion detection zone. In this case, given the type of each designated motion detection zone obtained in the metadata, the system can obtain the associated monitoring parameter(s), for example, via a lookup table. It should be understood that other variations of storing and accessing monitoring parameters for detection zones are also possible.
[0070] Further references Figure 4 , method 400 may also include sending one or more alerts (act 410) in response to detecting the presence of the event(s) in the plurality of zones in act 408. For example, method 400 may send an alert to a communication network (e.g., Figure 1 106) sends an alert, wherein the alert indicates that an event exists in a designated motion detection zone. In return, the user device (e.g., Figure 1 The alert may be received via a communication network (e.g., 110). As previously described in this disclosure, the alert may be of a type associated with the type of motion detection zone in which the event triggering the alert was detected. In some examples, the type of alert associated with the motion detection zone may also be included in metadata transmitted to the server.
[0071] In the examples previously described in this disclosure, method 400 can send an alert to a user device via a call, text message, or other electronic means, depending on the type of alert. For example, if the motion detection zone is an intruder zone (e.g., a front porch area), an event detected in that zone can be considered urgent in nature. In response to detecting a motion event in the intruder zone, the system can immediately initiate a call to the user device. In another example, if the motion detection zone is a delivery zone, such as a garage door, an event detected in that zone can be considered non-urgent. Thus, in response to detecting an event in the delivery zone, the system can send a text message to the user device upon or after the delivery is detected.
[0072] Further references Figure 4, method 400 may optionally use the image portions related to the region of interest (e.g., the designated motion detection zone) and the metadata associated with each zone to reconstruct one or more images of the scene (in act 406). For example, the portion of the image received in act 402 may include multiple pixels of the region of interest (e.g., the designated motion detection zone), wherein the metadata received in act 404 may include the positions of the multiple pixels. In this case, method 400 may use information about the position of each of the multiple pixels to reconstruct an image of the scene with an accurately located region of interest (e.g., the designated motion zone). Other areas in the image (e.g., non-designated areas) may be filled with zero (NULL) valued pixels. This allows subsequent processing (e.g., act 408) to be performed on the reconstructed image. It will be understood that Figure 4 Variations of the above embodiments are possible. For example, constructing an image (act 406) can be implemented as part of a motion detection system (e.g., including acts 408 and 410). Alternatively or additionally, constructing an image (act 406) can be implemented in a standalone application. For example, method 400 can include acts 402, 404, and 406 to reconstruct an image captured by a camera at a scene.
[0073] Additionally or alternatively, the system can use information associated with multiple regions to derive data that provides the system with a meaningful understanding of the real-world scene. For example, the segmentation technique described above may associate each of the multiple regions with a corresponding category (e.g., tree, sky, road, courtyard, furniture, etc.). This information regarding the association of the categories of the multiple regions may be included in metadata. In some embodiments, the server may receive the metadata (e.g., action 404) and use the metadata to infer the type of scene. For example, if multiple regions are associated with trees, sky, road, etc., the system may use this information to determine or infer that the scene is outdoor. If multiple regions are associated with chairs, couches, tables, etc., the system may use this information to determine or infer that the scene is indoor. In some embodiments, the system may use this inference of the scene to determine the appropriate model for subsequent image analysis (e.g., motion detection). For example, when detecting motion, the system may use different machine learning models for outdoor and indoor scenes, depending on the scene type. Thus, by using different models for different scene types, the system can perform motion detection or other image analysis operations with low latency, high accuracy, and / or high efficiency.
[0074] Figure 8 An illustrative embodiment of a computer system 800 that can be used to perform any aspect of the techniques and embodiments disclosed herein is shown in FIG. For example, the computer system 800 can be installed on a server 110, such as a server 110. Figure 1In another example, the computer system 800 may be installed on an image capture device (e.g., camera 102). In another example, the computer system 800 may be installed on a user device (e.g., Figure 1 110). The computer system 800 may be configured to execute Figures 2 to 4 The various methods and actions described herein. The computer system 800 may include one or more processors 810 and one or more non-transitory computer-readable storage media (e.g., memory 820 and one or more non-volatile storage media 830), as well as a display 840. The processor 810 may control the writing of data to and reading of data from the memory 820 and non-volatile storage device 830 in any suitable manner, as the aspects of the invention described herein are not limited in this respect. To perform the functions and / or techniques described herein, the processor 810 may execute one or more instructions stored in one or more computer-readable storage media (e.g., memory 820, storage media, etc.), which may serve as non-transitory computer-readable storage media storing instructions for execution by the processor 810.
[0075] In conjunction with the techniques described herein, code for, for example, detecting anomalies in an image / video can be stored on one or more computer-readable storage media of the computer system 800. The processor 810 can execute any such code to provide any of the techniques for detecting anomalies as described herein. Any other software, program, or instruction described herein can also be stored and executed by the computer system 800. It will be understood that the computer code can be applied to any aspect of the methods and techniques described herein. For example, the computer code can be applied to interact with an operating system to detect anomalies through traditional operating system processes.
[0076] The various methods or processes outlined herein may be encoded as software that can be executed on one or more processors employing any of a variety of operating systems or platforms. Additionally, such software may be written using any of a variety of suitable programming languages and / or programming or scripting tools, and may also be compiled into executable machine language code or intermediate code that is executed on a virtual machine or suitable framework.
[0077] In this regard, the various inventive concepts may be embodied as at least one non-transitory computer-readable storage medium (e.g., computer memory, one or more floppy disks, compact disks, optical disks, magnetic tapes, flash memory, circuit configurations in field programmable gate arrays or other semiconductor devices, etc.) encoded with one or more programs that, when executed on one or more computers or other processors, implement various embodiments of the present invention. The one or more non-transitory computer-readable media may be transportable, such that the one or more programs stored thereon can be loaded onto any computer resource to implement each of the aspects of the present invention as described above.
[0078] The terms "program", "software" and / or "application" are used herein in a generic sense to refer to any type of computer code or computer-executable instruction set that can be used to program a computer or other processor to implement various aspects of the embodiments described above. In addition, it should be understood that, according to one aspect, one or more computer programs that, when executed, perform the methods of the present invention need not reside on a single computer or processor, but can be distributed in a modular manner among different computers or processors to implement various aspects of the present invention.
[0079] Computer-executable instructions can take many forms, such as program modules, that are executed by one or more computers or other devices. Typically, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. Typically, the functionality of program modules can be combined or distributed as desired in various embodiments.
[0080] Furthermore, the data structure may be stored in any suitable form in a non-transitory computer-readable storage medium. A data structure may have fields that are related by location within the data structure. This relationship may also be achieved by allocating storage for the fields at locations in the non-transitory computer-readable medium that convey the relationship between the fields. However, any suitable mechanism may be used to establish a relationship between information in the fields of a data structure, including through the use of pointers, tags, or other mechanisms that establish relationships between data elements.
[0081] Various inventive concepts can be embodied as one or more methods, examples of which have been provided. The actions performed as part of a method can be ordered in any suitable manner. Thus, embodiments can be constructed in which the actions are performed in an order different from that illustrated, which can include performing some actions simultaneously, even though shown as sequential actions in an illustrative embodiment.
[0082] Unless expressly indicated to the contrary, the indefinite articles "a" and "an" as used in this specification and claims should be understood to mean "at least one". As used in the specification and claims, the phrase "at least one" with respect to a list of one or more elements should be understood to mean at least one element selected from any one or more elements in the list of elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding any combination of elements in the list of elements. This allows for the optional presence of elements other than the elements specifically identified within the list of elements to which the phrase "at least one" refers, whether related or unrelated to those elements specifically identified.
[0083] The phrase "and / or" as used in the specification and claims should be understood to mean "either or both" of the elements so combined, i.e., elements that are present in combination in some cases and separately in other cases. Multiple elements listed with "and / or" should be interpreted in the same manner, i.e., "one or more" of the elements so combined. Other elements besides the elements specifically identified by the "and / or" clause may optionally be present, whether related or unrelated to those specifically identified elements. Thus, as a non-limiting example, a reference to "A and / or B," when used in conjunction with open language such as "comprising," may refer to only A (optionally including elements other than B) in one embodiment; to only B (optionally including elements other than A) in another embodiment; to both A and B (optionally including other elements) in yet another embodiment; and so on.
[0084] As used in the specification and claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" should be interpreted as inclusive, i.e., including at least one of a plurality of elements or a list of elements, but also including more than one, and optionally including additional unlisted items. Only terms that clearly indicate the contrary (such as "only one" or "exactly one") or, when used in a claim, "consisting of..." will refer to including exactly one element of a plurality of elements or a list of elements. Generally, when the term "or" is preceded by an exclusive term (such as "any one," "one of," "only one of," or "exactly one of"), the term "or" as used herein should only be interpreted to indicate an exclusive alternative (i.e., "one or the other, but not both"). When used in a claim, "consisting essentially of..." should have its ordinary meaning as used in the field of patent law.
[0085] The use of ordinal terms such as "first," "second," "third," etc. in the claims to modify claim elements does not, by itself, imply any priority, precedence, or sequence of one claim element with respect to another, or a temporal order in which the acts of the method are performed. These terms serve merely as labels to distinguish one claim element having a particular name from another element having the same name (but using an ordinal term).
[0086] The phraseology and terminology used herein is for the purpose of description and should not be regarded as limiting. The use of "includes," "comprising," "having," "containing," "involving," and variations thereof, is meant to encompass the items listed thereafter and additional items.
[0087] Having described several embodiments of the present invention in detail, those skilled in the art will readily appreciate various modifications and improvements. These modifications and improvements are within the spirit and scope of the present invention. Therefore, the foregoing description is intended to be exemplary only and not restrictive.
[0088] Each aspect is described in this disclosure, including but not limited to the following aspects.
Claims
1. A method comprising: receiving first pixel data corresponding to a first area within the camera's imaging field of view and second pixel data corresponding to a second area within the imaging field of view; determining that the first pixel data represents a first event; determining first metadata associating the first region with a first delivery form of the alert; causing a first alert to be sent to a first application of the terminal device using the first transmission form based at least in part on first pixel data representing the first event and first metadata associating the first region with the first transmission form; determining that the second pixel data represents a second event; determining second metadata associating the second region with a second delivery form of the alert, wherein the second delivery form is different from the first delivery form; as well as Based at least in part on second pixel data representing a second event and second metadata associating a second region with a second transmission form, a second alert is transmitted to a second application of the terminal device using the second transmission form, the second application being different from the first application.
2. The method according to claim 1, wherein: Causing the first alert to be sent to the first application includes configuring the first alert to be output via the first application; and Causing the second alert to be sent to the second application includes configuring the second alert to be output by the second application.
3. The method according to claim 1 or 2, wherein: The first application is a mobile phone application; and The second application is a text messaging application.
4. The method according to claim 1, further comprising: determining that the first metadata further associates the first region with a first sending time of an alert following event detection; as well as determining the second metadata further associating the second region with a second time for sending an alert after event detection; wherein causing the first alert to be sent to the terminal device further comprises, based at least in part on first metadata associating the first area with the first sending time, immediately causing the first alert to be sent to the terminal device upon determining that the first pixel data represents the first event; and The sending of the second alert to the terminal device further comprises, based at least in part on second metadata associating the second area with the second sending time, delaying sending of the second alert to the terminal device after determining that the second pixel data represents a second event.
5. The method according to claim 1, further comprising: determining the first metadata further associates the first region with one or more first monitoring parameters; processing the first pixel data using the one or more first monitoring parameters to detect a first event based at least in part on first metadata associating the first area with the one or more first monitoring parameters; determining the second metadata further associating the second region with one or more second monitoring parameters, wherein the one or more second monitoring parameters are different from the one or more first monitoring parameters; as well as The second pixel data is processed using the one or more second monitoring parameters to detect a second event based at least in part on second metadata associating the second area with the one or more second monitoring parameters.
6. The method according to claim 5, wherein: One or more first monitoring parameters identify a first time window during which the first pixel data is processed to detect an event; as well as The one or more second monitoring parameters identify a second time window during which the second pixel data is processed to detect the event, wherein the second time window is different from the first time window.
7. The method according to claim 5, wherein: One or more first monitoring parameters identify a first event type to be detected based on the first pixel data; as well as The one or more second monitoring parameters identify a second event type to be detected based on the second pixel data, wherein the second event type is different from the first event type.
8. The method according to claim 5, wherein: Processing the first pixel data using the one or more first monitoring parameters further includes sending the first pixel data to a remote computing system for processing; as well as Processing the second pixel data using the one or more second monitoring parameters also includes sending the second pixel data to a remote computing system for processing.
9. The method according to claim 8, further comprising: Sending third pixel data included in the image with the first pixel data or the second pixel data to the remote computing system is avoided.
10. The method according to claim 9, further comprising: Third metadata identifying a relative location of the first pixel data or the second pixel data within the image is sent to the remote computing system.
11. The method according to claim 8, further comprising: sending one or more first monitoring parameters to a remote computing system so that the remote computing system processes the first pixel data according to the one or more first monitoring parameters; as well as The one or more second monitoring parameters are sent to the remote computing system so that the remote computing system processes the second pixel data according to the one or more second monitoring parameters.
12. The method according to claim 8, further comprising: sending a first indication that the first region is associated with one or more first monitoring parameters to a remote computing system to cause the remote computing system to process the first pixel data according to the one or more first monitoring parameters; as well as A second indication that the second region is associated with the one or more second monitoring parameters is sent to the remote computing system to cause the remote computing system to process the second pixel data according to the one or more second monitoring parameters.
13. The method of claim 1, wherein: The first event corresponds to motion detected within the first area; and The second event corresponds to motion detected within the second area.
14. The method according to claim 1, further comprising: determining, using automated image analysis techniques, a plurality of regions from an image of a scene, the plurality of regions comprising at least a first region and a second region; displaying the plurality of regions; receiving a user selection of a first area and a second area from the plurality of areas, wherein the user selection indicates that the first area and the second area are designated as motion detection areas; as well as First data designating the first area and the second area as motion detection areas is stored for use in performing motion detection on one or more subsequent images acquired by the camera within the field of view.
15. The method according to claim 14, further comprising: Information associated with each of the plurality of areas is displayed, indicating a likelihood that each of the plurality of areas will be used as a motion detection area.
16. The method according to claim 15, wherein The displayed information associated with each of the plurality of areas includes a respective score value or a graphical representation associated with each of the plurality of areas.
17. The method according to claim 14, wherein: Determining the plurality of regions further comprises: Semantic segmentation is performed on the image of the scene to segment the image into the plurality of regions.
18. A system comprising: one or more processors; and one or more non-transitory computer-readable media encoded with instructions that, when executed by the one or more processors, cause the system to: receiving first pixel data corresponding to a first area within the camera's imaging field of view and second pixel data corresponding to a second area within the imaging field of view; determining that the first pixel data represents a first event; determining first metadata associating a first region with a first time of sending an alert following event detection; causing a first alert to be sent to the terminal device a first amount of time after detecting the first event according to the first sending time based at least in part on first pixel data representing the first event and first metadata associating the first region with the first sending time; determining that the second pixel data represents a second event; determining second metadata associating the second region with a second sending time of the alert following event detection, wherein the second sending time is different from the first sending time; as well as Based at least in part on second pixel data representing a second event and second metadata associating a second area with a second transmission time, cause a second alert to be sent to the terminal device a second amount of time after detecting the second event according to the second transmission time, the second amount of time being greater than the first amount of time.
19. The system according to claim 18, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: causing the first alert to be sent to the terminal device a first amount of time after detecting the first event by causing the first alert to be sent to the terminal device immediately upon determining that the first pixel data represents the first event; as well as The second alert is caused to be sent to the terminal device a second amount of time after detecting the second event at least in part by delaying sending the second alert to the terminal device after determining that the second pixel data represents the second event.
20. The system of claim 18, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: determining the first metadata further associates the first region with one or more first monitoring parameters; processing the first pixel data using the one or more first monitoring parameters to detect a first event based at least in part on first metadata associating the first area with the one or more first monitoring parameters; determining the second metadata further associating the second region with one or more second monitoring parameters, wherein the one or more second monitoring parameters are different from the one or more first monitoring parameters; as well as The second pixel data is processed using the one or more second monitoring parameters to detect a second event based at least in part on second metadata associating the second area with the one or more second monitoring parameters.
21. The system of claim 20, wherein: One or more first monitoring parameters identify a first time window during which the first pixel data is processed to detect an event; as well as The one or more second monitoring parameters identify a second time window during which the second pixel data is processed to detect the event, wherein the second time window is different from the first time window.
22. The system of claim 20, wherein: One or more first monitoring parameters identify a first event type to be detected based on the first pixel data; and The one or more second monitoring parameters identify a second event type to be detected based on the second pixel data, wherein the second event type is different from the first event type.
23. The system of claim 20, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: processing the first pixel data using the one or more first monitoring parameters at least in part by sending the first pixel data to a remote computing system for processing; as well as The second pixel data is processed using the one or more second monitoring parameters at least in part by sending the second pixel data to a remote computing system for processing.
24. The system of claim 23, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: Sending third pixel data included in the image with the first pixel data or the second pixel data to the remote computing system is avoided.
25. The system of claim 24, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: Third metadata identifying a relative position of the first pixel data or the second pixel data within the image is sent to the remote computing system.
26. The system of claim 23, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: sending one or more first monitoring parameters to a remote computing system so that the remote computing system processes the first pixel data according to the one or more first monitoring parameters; as well as The one or more second monitoring parameters are sent to the remote computing system so that the remote computing system processes the second pixel data according to the one or more second monitoring parameters.
27. The system of claim 23, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: sending a first indication that the first region is associated with one or more first monitoring parameters to a remote computing system to cause the remote computing system to process the first pixel data according to the one or more first monitoring parameters; as well as A second indication that the second region is associated with the one or more second monitoring parameters is sent to the remote computing system to cause the remote computing system to process the second pixel data according to the one or more second monitoring parameters.
28. The system of claim 18, wherein: The first event corresponds to motion detected within the first area; and The second event corresponds to motion detected within the second area.
29. The system of claim 18, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: determining, using automated image analysis techniques, a plurality of regions from an image of a scene, the plurality of regions comprising at least a first region and a second region; displaying the plurality of regions; receiving a user selection of a first area and a second area from the plurality of areas, wherein the user selection indicates that the first area and the second area are designated as motion detection areas; as well as First data designating the first area and the second area as motion detection areas is stored for use in performing motion detection on one or more subsequent images acquired by the camera within the field of view.
30. The system of claim 29, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: Information associated with each of the plurality of areas is displayed, indicating a likelihood that each of the plurality of areas will be used as a motion detection area.
31. The system of claim 30, wherein: The displayed information associated with each of the plurality of areas includes a respective score value or a graphical representation associated with each of the plurality of areas.
32. The system of claim 18, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: determining that the first metadata further associates the first region with the first delivery form of the alert; and determining that the second metadata further associates the second region with a second delivery form of the alert, wherein the second delivery form is different from the first delivery form; wherein causing the first alert to be sent to the terminal device further comprises causing the first alert to be sent to the first application of the terminal device using the first sending form based at least in part on first metadata associating the first area with the first sending form; and Sending the second alert to the terminal device further includes, based at least in part on second metadata associating the second area with the second sending form, sending the second alert to a second application of the terminal device using the second sending form, the second application being different from the first application.
33. The system of claim 32, wherein: The one or more non-transitory computer-readable media are further encoded with additional instructions that, when executed by the one or more processors, further cause the system to: causing the first alert to be sent to the first application at least in part by configuring the first alert to be output via the first application; and The second alert is caused to be sent to the second application at least in part by configuring the second alert to be output by the second application.
34. A system according to claim 32 or 33, wherein: The first application is a mobile phone application; and The second application is a text messaging application.
35. A method comprising: receiving pixel data corresponding to different regions within a field of view of a camera, the different regions including at least a first region and a second region; identifying a plurality of events based on the pixel data, the plurality of events including at least a first event based on first pixel data corresponding to the first region and a second event based on second pixel data corresponding to the second region; Determining, based on different areas, a plurality of alarm sending forms and / or alarm sending times after event monitoring, the plurality of sending forms and / or sending times including at least a first sending form and / or first sending time based on a first area and a second sending form and / or second sending time based on a second area; as well as At least partially based on the multiple events and the multiple sending forms and / or sending times, multiple alarms are sent to the terminal device, the multiple alarms including at least a first alarm for a first event and a second alarm for a second event, the first alarm (A) being sent to a first application of the terminal device using a first sending form and / or (B) being sent a first amount of time after detecting the first event according to the first sending time, the second alarm (A) being sent to a second application of the terminal device different from the first application using a second sending form and / or (B) being sent a second amount of time after detecting the second event according to the second sending time, the second amount of time being greater than the first amount of time.
Citation Information
Patent Citations
Doorbell system and security method thereof
US20190066470A1
Methods and systems for determining object activity within a region of interest
US20190318171A1