Systems and methods for remote monitoring in internet of things (IOT) environments
Patent Information
- Application Number
- CN202480088749.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-02-28
- Filing Date
- 2024-12-10
- Publication Date
- 2026-09-25
AI Technical Summary
然而,用户不能确定警报是否是假警报
Smart Images

Figure CN122826809A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to surveillance systems, and more specifically, to systems and methods for remote surveillance in Internet of Things (IoT) environments. Background Technology
[0002] In recent years, with the increase in connected devices and smart homes, home monitoring systems have also proliferated. These systems aim to enhance user safety and well-being. For example, parents can set up a home monitoring system in their home to remotely monitor their children when they are not home. Thus, users are empowered to participate in real-time observation and supervision.
[0003] Home surveillance systems utilize a combination of sensors. In some cases, cameras can be used to monitor various locations within an area (e.g., a house). However, it is difficult to install cameras in every location within that area. Furthermore, it is difficult for users to monitor locations when no cameras are installed.
[0004] For example, elderly people, children, or pets may be present in locations where no cameras are installed. In related technologies, motion alerts can be generated, and users can be notified about these alerts. However, users cannot determine whether an alert is a false alarm. For example, an alert might be generated due to the movement of a pet or an elderly person in a location where no cameras are installed. Furthermore, users may not be able to visualize the location or why an alert was generated. That is, in related technologies, text alerts are generated without clear visuals, and users do not receive accurate visualization. In addition, there is an increased number of false alarms in related technologies. Therefore, cognitive load increases and user engagement decreases significantly. Summary of the Invention
[0005] [Technical Issues]
[0006] According to one embodiment of this disclosure, a method for remote monitoring in an Internet of Things (IoT) environment is disclosed. The method includes: detecting one or more triggering conditions associated with the occurrence of an event at a location within the IoT environment, wherein the one or more triggering conditions are associated with an event criterion; detecting one or more entities present at the location associated with the event using one or more non-imaging sensors; determining the operational state of one or more IoT devices at the location when the one or more entities are detected; determining whether the one or more triggering conditions are false triggering conditions based on the correlation between the one or more entities and the operational state of the one or more IoT devices; and modifying the event criterion based on determining that the one or more triggering conditions are false triggering conditions.
[0007] According to another embodiment of this disclosure, a system for remote monitoring in an Internet of Things (IoT) environment is disclosed. The remote monitoring system includes a memory and at least one processor communicatively coupled to the memory. The at least one processor is configured to: detect one or more triggering conditions associated with the occurrence of an event at a location within the IoT environment, wherein the one or more triggering conditions are associated with an event criterion; use one or more non-imaging sensors to detect one or more entities present at the location associated with the event; determine the operational state of one or more IoT devices at the location upon detection of the one or more entities; determine whether the one or more triggering conditions are false triggering conditions based on the correlation between the one or more entities and the operational state of the one or more IoT devices; and modify the event criterion based on the determination that the one or more triggering conditions are false triggering conditions.
[0008] According to another embodiment of this disclosure, a method for remote monitoring in an Internet of Things (IoT) environment is disclosed. The method includes: detecting one or more triggering conditions associated with the occurrence of an event at a location within the IoT environment, wherein the one or more triggering conditions are associated with an event criterion; detecting one or more entities present at the location associated with the event using one or more non-imaging sensors; determining the operational state of one or more IoT devices at the location when the one or more entities are detected; determining whether the one or more triggering conditions are false or true triggering conditions based on the correlation between the one or more entities and the operational state of the one or more IoT devices; and generating a virtual visual representation of the location within the IoT environment based on the determination that the one or more triggering conditions are true triggering conditions, the virtual visual representation graphically indicating the occurrence of the event at the location.
[0009] According to another embodiment of this disclosure, a non-transitory computer-readable medium storing instructions is disclosed. The instructions cause at least one processor to: detect one or more triggering conditions associated with the occurrence of an event at a location within an IoT environment, wherein the one or more triggering conditions are associated with an event criterion; use one or more non-imaging sensors to detect one or more entities present at the location associated with the event; determine the operational state of one or more IoT devices at the location upon detection of the one or more entities; determine whether the one or more triggering conditions are false or true triggering conditions based on the correlation between the one or more entities and the operational state of the one or more IoT devices; and based on the determination that the one or more triggering conditions are true triggering conditions, generate a virtual visual representation of the location within the IoT environment, the virtual visual representation graphically indicating the occurrence of the event at the location. Attached Figure Description
[0010] These and other features, aspects, and advantages of the invention will become better understood when the following detailed description is read with reference to the accompanying drawings, wherein the same characters denote the same parts throughout the drawings, wherein:
[0011] Figure 1 The home monitoring process is shown;
[0012] Figure 2 An exemplary representation of an Internet of Things (IoT) environment according to embodiments of the present disclosure is shown;
[0013] Figure 3 A block diagram of a system for monitoring an IoT environment according to an embodiment of the present disclosure is shown;
[0014] Figure 4 A process flow depicting the operation of an entity detection module according to an embodiment of the present disclosure is shown;
[0015] Figure 5 An exemplary visual representation generated by the system according to an embodiment of this disclosure is shown;
[0016] Figure 6 An exemplary flowchart illustrating remote monitoring of an IoT environment according to embodiments of the present disclosure is shown; and
[0017] Figure 7-8 An exemplary flowchart of a process for remote monitoring in an IoT environment according to various embodiments of the present disclosure is shown.
[0018] Furthermore, those skilled in the art will understand that the elements in the accompanying drawings are shown for simplicity and may not necessarily be drawn to scale. For example, flowcharts illustrate the method according to the most prominent steps involved to aid in understanding various aspects of the invention. Additionally, regarding the construction of the device, one or more components of the device may be indicated by symbols in the drawings, and the drawings may only show specific details relevant to understanding embodiments of the invention, so as not to obscure the drawings with details readily understood by those of ordinary skill in the art who benefit from the description herein. Detailed Implementation
[0019] To facilitate understanding of the principles of the invention, reference will now be made to various embodiments, which will be described using specific language. However, it will be understood that this is not intended to limit the scope of the invention, and such variations and further modifications in the illustrated systems, as well as such further applications of the principles of the invention as shown therein, are contemplated as would normally occur to those skilled in the art to which this invention pertains.
[0020] Those skilled in the art will understand that the foregoing general description and the following detailed description are explanations of the invention and are not intended to limit it.
[0021] Throughout this specification, references to "one aspect," "another aspect," or similar language mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this disclosure. Therefore, the phrases "in an embodiment," "in another embodiment," and similar language throughout this specification may, but do not necessarily, refer to the same embodiment.
[0022] The terms “comprise,” “comprising,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a process or method that includes a list of steps may include not only those steps but may also include other steps not expressly listed or inherent to such process or method. Similarly, without further limitation, a list of one or more devices, subsystems, elements, structures, or components beginning with “comprising…” does not exclude the presence of other devices or subsystems or elements or structures or components, or additional devices or subsystems or elements or structures or components.
[0023] refer to Figure 1 It describes the home surveillance process. For example... Figure 1 As seen, region 100 can be an area where motion and audio can be detected at location 102. For example, motion could be caused by a cat moving, and audio could be caused by a cat meowing at location 102. As a result of motion and audio detection, an alarm is generated and displayed to a remote user on user device 104. As an example, a notification could be sent to the user. However, user 106 may not have a detailed context and visualization of location 102. User 106 may not know why an alarm was generated.
[0024] As another example, an elderly person might be drinking water from a glass in a specific location. The glass might break, causing a sound. A fall detection warning could be sent to the user; however, the user might not be sure whether the warning is due to an accident or some other reason. The user lacks context and visualization.
[0025] Therefore, systems and methods are needed to overcome at least some of the aforementioned limitations.
[0026] Figure 2An exemplary representation of an Internet of Things (IoT) environment 200 is shown. Environment 200 may be associated with a system 210 configured for remote monitoring of the IoT environment 200. In embodiments, system 210 may be provided within the IoT environment 200, such as integrated with one or more components within the IoT environment 200. In embodiments, system 210 may be provided remotely from the IoT environment 200, such as as a cloud-based unit. In embodiments, system 210 may be provided in a distributed manner, wherein one or more components of system 210 are provided within the IoT environment 200 and remotely from the IoT environment 200.
[0027] The IoT environment 200 can be associated with an area to be monitored by a remote user. The IoT environment 200 can be associated with a house, office, etc., within which monitoring is required. The IoT environment 200 can include multiple locations 202 in which one or more entities 204 may be present. The multiple locations 202 can include, for example, a living room, bedroom, kitchen area, garage, etc. It should be understood that the details of the invention can be explained with reference to the locations among the multiple locations 202; however, these details also apply to all locations 202 within the IoT environment 200.
[0028] The IoT environment 200 may further include one or more non-imaging sensors 206 and one or more IoT devices 208. As a non-limiting example, the one or more non-imaging sensors 206 may include motion sensors, audio sensors, light sensors, positioning sensors, etc. As a non-limiting example, the one or more IoT devices 208 may include televisions, air conditioners, lighting units, refrigerators, etc. In embodiments, the camera may not be part of the system 210.
[0029] System 210 can be communicatively coupled to user equipment 220. User equipment 220 can be a user device for a remote user, enabling the remote user to perform remote monitoring in conjunction with system 210. User equipment 220 may include a user interface 222 that allows the user to view a visual representation of the monitored location 202. User equipment 220 can be a mobile phone, laptop computer, computer, tablet computer, or any suitable electronic device that can communicate with system 210. It should be understood that, although Figure 2 A single user equipment 220 is shown, but the system 210 can also be connected to multiple user equipments without departing from the scope of the invention.
[0030] refer to Figure 3 The diagram illustrates a detailed block diagram of a system 210 according to an embodiment of the present disclosure. System 210 may include multiple modules 301, a processor 302, an input / output (I / O) interface 303, a memory 304, and a transceiver 305.
[0031] In an exemplary embodiment, processor 302 may be operatively coupled to each of I / O interface 303, multiple modules 301, transceiver 305, and memory 304. In one embodiment, processor 302 may include a graphics processing unit (GPU) and / or an artificial intelligence engine (AIE). In one embodiment, processor 302 may include at least one data processor for performing processes in a virtual memory area network. Processor 302 may include dedicated processing units such as an integrated system (bus) controller, a memory management control unit, a floating-point unit, a graphics processing unit, a digital signal processing unit, etc. In one embodiment, processor 302 may include a central processing unit (CPU), a graphics processing unit (GPU), or both. Processor 302 may be one or more general-purpose processors, digital signal processors, application-specific integrated circuits, field-programmable gate arrays, servers, networks, digital circuits, analog circuits, combinations thereof, or other devices now known or hereafter developed for analyzing and processing data. Processor 302 may execute software programs, such as manually generated (i.e., programmed) code, to perform desired operations.
[0032] The processor 302 can be configured to communicate with one or more input / output (I / O) devices via the I / O interface 303. In some embodiments, the processor 302 can communicate with the user equipment 220 using the I / O interface 303. The I / O interface 303 can employ Near Field Communication (NFC), Bluetooth, Code Division Multiple Access (CDMA), High Speed Packet Access (HSPA+), Global System for Mobile Communications (GSM), Long Term Evolution (LTE), WiMax, etc.
[0033] Using I / O interface 303, system 210 can communicate with one or more I / O devices. For example, input devices may be antennas, microphones, touchscreens, touchpads, storage devices, transceivers, video devices / sources, etc. Output devices may be video displays (e.g., cathode ray tube (CRT), liquid crystal displays (LCD), light-emitting diodes (LEDs), plasma displays, plasma display panels (PDP), organic light-emitting diode displays (OLEDs), etc.), audio speakers, etc.
[0034] Processor 302 can be configured to communicate with a communication network via a network interface. In an embodiment, the network interface may be I / O interface 303. The network interface can be connected to the communication network to enable system 210 to connect to user equipment 220 and / or the external environment. The network interface may employ connection protocols, including but not limited to direct connection, Ethernet (e.g., twisted pair 10 / 100 / 1000BaseT), Transmission Control Protocol / Internet Protocol (TCP / IP), Token Ring, IEEE 802.11a / b / g / n / x, etc. The communication network may include, but is not limited to, direct interconnect, local area network (LAN), wide area network (WAN), wireless network (e.g., using Wireless Application Protocol), Internet, etc. Using the network interface and the communication network, system 210 can communicate with other devices.
[0035] In some embodiments, memory 304 may be communicatively coupled to processor 302. Memory 304 may be configured to store data and instructions executable by processor 302. In another embodiment, memory 304 may be provided via a cloud-based unit. In yet another embodiment, memory 304 may communicate with processor 302 via a bus within system 210. In yet another embodiment, memory 304 may be located remotely from processor 302 and may communicate with processor 302 via a network. Memory 304 may include, but is not limited to, non-transitory computer-readable storage media, such as various types of volatile and non-volatile storage media, including but not limited to random access memory, read-only memory, programmable read-only memory, electrically programmable read-only memory, electrically erasable read-only memory, flash memory, magnetic tape or disk, optical media, etc. In one example, memory 304 may include cache or random access memory for processor 302. In alternative examples, memory 304 is decoupled from processor 302, such as processor cache memory, system memory, or other memory. Memory 304 may be an external storage device or database for storing data. Memory 304 is operable to store instructions executable by processor 302. The functions, actions, or tasks shown or described in the figures can be executed by the programmed processor 302 to execute the instructions stored in memory 304. Functions, actions, or tasks are independent of a particular type of instruction set, storage medium, processor, or processing strategy, and can be executed by software, hardware, integrated circuits, firmware, microcode, etc., operating individually or in combination. Similarly, processing strategies can include multiprocessing, multitasking, parallel processing, etc.
[0036] In some embodiments, multiple modules 301 may be included within memory 304. Memory 304 may also include a database for storing data. Multiple modules 301 may include an instruction set that can be executed to cause system 210 (particularly the processor 302 of system 210) to perform any one or more of the methods / processes disclosed herein. Multiple modules 301 may be configured to perform steps of this disclosure using data stored in the database. For example, multiple modules 301 may be configured to perform… Figure 7-8 The steps disclosed herein. In an embodiment, each of the plurality of modules 301 may be a hardware unit that may be external to memory 304. Furthermore, memory 304 may include an operating system for performing one or more tasks of system 210, such as those performed by a general-purpose operating system.
[0037] Transceiver 305 can be configured to receive signals from and / or transmit signals to user equipment 220. In one embodiment, the database can be configured to store information required by the plurality of modules 301 and processor 302 to perform one or more functions. Figure 7-8 Exemplary embodiments are described in detail below.
[0038] The multiple modules 301 may include, but are not limited to, an event detection module 310, an entity detection module 312, a device status module 314, a trigger module 316, a semantic module 318, and a vision module 320. The multiple modules 301 can be implemented through suitable hardware and / or software applications.
[0039] In some embodiments, at least one of the plurality of modules 301 may use an AI model. AI-related functions may be performed via non-volatile memory, volatile memory, and processor 302.
[0040] Processor 302 may include one or more processors. At this time, the one or more processors may be general-purpose processors (such as central processing unit (CPU), application processor (AP), etc.), pure graphics processing units (such as graphics processing unit (GPU), vision processing unit (VPU)) and / or AI-specific processors (such as neural processing unit (NPU)).
[0041] One or more processors control the processing of input data according to predefined operating rules stored in non-volatile memory, or may employ a suitable artificial intelligence (AI) model executed from a server or local memory module.
[0042] AI models can consist of multiple neural network layers. Each layer has multiple weight values, and layer operations are performed by computing the previous layer and operating on the multiple weights. Examples of neural networks include, but are not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted Boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), and deep Q-networks.
[0043] Learning techniques are methods used to train a predetermined target device (e.g., a robot) using multiple learning data sets to enable, allow, or control the target device to make determinations or predictions. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, active learning, and reinforcement learning. Processor 302 can perform preprocessing operations on the data to transform it into a form suitable for use as input to an artificial intelligence (AI) model.
[0044] Reasoning and prediction are techniques for making logical inferences and predictions based on information, and include, for example, knowledge-based reasoning, optimization prediction, preference-based planning or recommendation.
[0045] Furthermore, this disclosure envisions a computer-readable medium that includes instructions or receives and executes instructions in response to a propagated signal. Additionally, instructions can be sent or received via a network through a communication port or interface or using a bus (not shown). The communication port or interface can be part of processor 302 or can be a separate component. The communication port can be created in software or can be a physical connection in hardware. The communication port can be configured to connect to a network, external media, a display, or any other component or combination thereof in the system. The connection to the network can be a physical connection, such as a wired Ethernet connection, or can be established wirelessly. Similarly, additional connections to other components of system 210 can be physical or can be established wirelessly. The network can alternatively be directly connected to the bus. For brevity, the architecture and standard operation of memory 304, processor 302, transceiver 305, and I / O interface 303 are not discussed in detail.
[0046] It should be understood that the details described with reference to multiple modules 301 can be executed in conjunction with processor 302 and memory 304. This will now be achieved through common reference. Figure 2-3 The details of the present invention will now be described.
[0047] Initially, one or more entities 204 may exist at location 202. One or more entities 204 may include, for example, an elderly person, a child, a pet, etc. As an example, a child (entity) may exist in the living room (location). As another example, a cat (entity) may be sitting in the living room (location). As yet another example, an elderly person (entity) may exist in the kitchen (location).
[0048] The event detection module 310 can be configured to detect the occurrence of an event at location 202. The event detection module 310 can be configured to detect one or more triggering conditions associated with the occurrence of an event at location 202 within the IoT environment 200. As an example, the event could include an elderly person (entity) falling in the kitchen (location). As another example, the event could include a cat (entity) playing with toys in the living room (location). In embodiments, the event can be detected based on readings from one or more non-imaging sensors 206.
[0049] One or more triggering conditions can be associated with event criteria. Event criteria can refer to predefined criteria that help distinguish between real and false events to avoid triggering false alarms. In an embodiment, event criteria can be a set of dynamic criteria that can be modified based on the occurrence of events and associated feedback over a period of time.
[0050] Entity detection module 312 can be configured to detect one or more entities 204 associated with an event at location 202. That is, entity detection module 312 can be configured to detect which entity is linked to an event already detected by event detection module 310. In an embodiment, entity detection module 312 can be configured to determine the location of one or more entities within location 202.
[0051] The entity detection module 312 can be configured to detect one or more entities based on readings from one or more non-imaging sensors 206. The entity detection module 312 can be configured to monitor readings from one or more non-imaging sensors 206 associated with one or more entities 204 present at location 202. The one or more non-imaging sensors 206 may include motion sensors, audio sensors, light sensors, positioning sensors (ultra-wideband), etc.
[0052] The entity detection module 312 can also be configured to determine the action being performed by one or more entities 204 based on the processing of readings monitored by one or more non-imaging sensors 206. Considering the example of an elderly person, the action being performed by one or more entities 204 could include a sudden movement of the elderly person and / or an elderly person screaming in the kitchen.
[0053] refer to Figure 4 The process flow depicting the operation of the entity detection module 312 is illustrated according to an embodiment of the present invention. For example... Figure 4As seen in operation S402, initially, the occurrence of an event is detected at a specific location (e.g., location 202 within an institution (e.g., a building)) based on one or more sensors (e.g., one or more non-imaging sensors 206). In operation S404, multiple motion detection points can be determined based on readings from one or more sensors (e.g., motion sensors from one or more non-imaging sensors 206). Furthermore, in operation S406, a region of interest corresponding to location 202 is determined. In operation S408, the motion and changes in motion of one or more entities 204 can be continuously monitored. Additionally, motion tracking can be performed based on object size, where the object corresponds to one or more entities.
[0054] In the same or another embodiment, at operation S410, audio data detected by an audio sensor from one or more non-imaging sensors (e.g., one or more non-imaging sensors 206) can be received. At operation S412, the audio data can be processed to determine audio frequencies and generate a spectrogram of the audio data. In an embodiment, a variable window technique can be used to perform continuous audio tracks on the received audio data to generate the spectrogram. At operation S414, the spectrogram output can be passed to a recognition model for processing. The recognition model can be an artificial intelligence (AI) model, such as a multi-class audio recognition model.
[0055] In some embodiments, operations S404-S408 and operations S410-S414 can be executed simultaneously. At operation S416, based on the motion tracking at operation S408 and the output of the AI model at operation S414, one or more entities 204 and their positions can be detected.
[0056] Refer again Figure 2-3 The device status module 314 can be configured to determine the corresponding operational state of one or more IoT devices 208 present at location 202. The corresponding operational state refers to the state of one or more IoT devices 208 when one or more entities 204 are detected and / or when an event occurs. For example, the device status module 314 can determine that at location 202 where one or more entities 204 are detected, a television is turned on, an air conditioner is turned on, or lights are switched to dim mode, etc.
[0057] Furthermore, the triggering module 316 can be configured to determine whether one or more triggering conditions are false or true triggering conditions. Classifying one or more triggering conditions as false or true triggering conditions indicates whether a detected event is a false event or a true event. A false event may refer to an event for which an alarm does not need to be provided to a remote user on user equipment 220. A true event may refer to an event for which an alarm needs to be generated on the user equipment 220 of a remote user.
[0058] Triggering module 316 can be configured to determine whether one or more triggering conditions are false triggering conditions based on the correlation between the detected one or more entities 204 and the corresponding operational states of the determined one or more IoT devices 208. To determine the correlation, semantic module 318 can be configured to determine semantic information indicating the current scene at location 202 when one or more entities 204 are detected. The current scene refers to the scene at location 202 when the event has occurred, defined by the actions of one or more entities, the operational states of one or more IoT devices 208, and the surrounding environment and background at location 202. As an example, the semantic information for a living room could indicate that the surrounding environment is well-lit, the curtains are open, and the television is broadcasting a news channel, etc.
[0059] In some embodiments, the semantic module 318 can generate semantic information and determine relevance via an artificial intelligence (AI) model. The semantic module 318 can use the AI model to determine parameters associated with one or more entities 204 and one or more IoT devices 208. Specifically, the semantic module 318 can determine a first set of parameters associated with one or more entities 204. The first set of parameters may include the location of one or more entities 204, actions being performed by one or more entities 204, etc. The semantic module 318 can also determine a second set of parameters associated with the operational state of one or more IoT devices 208. The second set of parameters may indicate the background scene and surrounding environment at location 202. Based on the first and second sets of parameters, the semantic module 318 can determine semantic information.
[0060] Based on semantic information, the triggering module 316 can determine whether one or more triggering conditions are true or false triggering conditions. To determine whether one or more triggering conditions are true or false, the triggering module 316 can consider readings from one or more non-imaging sensors 206 and the source of the detected event. For example, considering an event of an elderly person falling, an audio sensor can detect multiple sounds, and the source of the detected sounds could be sound from a television, the elderly person's voice, sound from equipment within location 202, etc. Therefore, the triggering module 316 can separate one or more sources of the detected event based on readings from one or more non-imaging sensors 206.
[0061] Furthermore, the trigger module 316 can determine a confidence score associated with each of the separated one or more sources. The confidence score can be determined based on one or more AI models. Considering the example of an elderly person falling, the trigger module 316 can determine that the confidence score for the sound from the television is 5%, the confidence score for the sound from household appliances is 2%, and the confidence score for the sound of a person falling is 90%.
[0062] Once the confidence score is determined, the triggering module 316 can compare the determined confidence score with a predetermined threshold. The predetermined threshold can be stored in the memory 304. Based on this comparison, the triggering module 316 can determine one or more trigger conditions as true trigger conditions or false trigger conditions. As mentioned above, classifying one or more trigger conditions as false or true trigger conditions can indicate whether the detected event is a false event or a true event. Therefore, the triggering module 316 can also classify the detected event as a false event or a true event based on the comparison of the confidence score with the predetermined threshold.
[0063] When the triggering module 316 determines one or more trigger conditions as false trigger conditions, the triggering module 316 can be configured to modify the event criteria. This modification can be considered as feedback for dynamically updating the event criteria, enabling increasingly accurate event detection over time. Therefore, the triggering module 316 prevents false alarms.
[0064] As an example, a cat (the physical entity) might be playing with toys in the living room (the location). A motion sensor can detect the event (the cat's movement). The state and semantic information of the IoT device 208 can be determined. Furthermore, the trigger module 316 can determine that the confidence score of the sound from the television is 80%, and the confidence score from the positioning sensor is 95%. Based on the confidence scores, the trigger module 316 can determine that the event is a false event and continue to update the event criteria. Based on the modification of the event criteria, the accuracy and efficiency of the system 210 are improved. For example, the same event may not be considered a trigger, thus avoiding further calculations.
[0065] When the triggering module 316 determines one or more triggering conditions as true triggering conditions (i.e., the event is determined to be a true event), further processing is performed. The vision module 320 can be configured to generate a visual representation of location 202 based on semantic information generated by the semantic module 318. The visual representation can indicate the occurrence of an event at location 202. In embodiments, the visual representation may include a representation of one or more entities 204, one or more IoT devices 208 in their operational state, and the surrounding environment of location 202.
[0066] In this embodiment, the vision module 320 can determine one or more ground images based on semantic information. The ground images may correspond to one or more entities 204 and the current scene at location 202. Specifically, the ground images may be sample images that can be stored in memory 304. For example, an image of location 202 may be available based on a physical map of location 202. Furthermore, sample images of one or more entities 204 may be available. Therefore, ground images can be considered to generate a visual representation.
[0067] In this embodiment, the visual representation is a static image or a set of static images. In this embodiment, the visual representation can be a sequence of frames generated based on semantic information. The frame sequence can be based on the current scene at location 202. For example, it can depict the operational state of one or more devices 208 and the actions of one or more entities 204. Therefore, a continuous visual representation of location 202 can be generated.
[0068] refer to Figure 5 This shows an example of visual representation. Figure 5 In the example, system 210 can detect the movement of a cat (an entity) in the bedroom of IoT environment 200. System 210 can detect the cat and further detect the operational status of IoT devices and the scene at that location. System 210 can detect the cat meowing near the bed. Furthermore, system 210 can detect that the television is on, the lighting is blue, and the blinds are closed. Based on semantic information, a visual representation 500 can be generated. As seen in visual representation 500, an avatar 502 of the cat can be shown near the bed 504. Furthermore, the television 506 can be shown as on, and the blinds can be shown as closed. Additionally, the lighting at this location can be displayed in blue. Therefore, the visual representation indicates that the current scene at the location of one or more entities (here, the cat) can be detected.
[0069] In one embodiment, the vision module 320 can remove sensitive information associated with one or more entities 204. In some embodiments, privacy rules can be predefined and stored in memory 304. In one embodiment, privacy rules can be defined by a remote user. Based on the predefined privacy rules, semantic information can be updated so as not to disclose the privacy of one or more entities 204.
[0070] In some embodiments, the semantic module 318 can be configured to detect semantic change conditions associated with semantic information. A semantic change can be the result of a predefined privacy rule that causes a change in the semantic information to remove sensitive information associated with one or more entities 204. That is, the semantic information can be updated by the semantic module 318 based on the detected semantic change conditions. Therefore, semantic change conditions can include units that remove and / or modify sensitive information from the semantic information. Furthermore, a visual representation can be generated by the visual module 320 based on the updated semantic information. For example, avatars of one or more entities 204 can be used in the visual representation. Therefore, the privacy of one or more entities is not disclosed. In embodiments, U-Net can be utilized to generate visual representations based on semantic information.
[0071] The vision module 320 can be configured to display a visual representation of the generated indication event on a user interface 222 associated with the user equipment 220 of the remote user. For example, the vision module 320 can send the generated visual representation to the user equipment 220, causing the visual representation to be displayed on the user interface 222. As previously described, the visual representation can be a sequence of frames representing the location. As a result, remote user monitoring of the location is achieved.
[0072] In some embodiments, system 210 can enable remote monitoring in a secure manner because all processing can occur on an edge unit or cloud unit with end-to-end encryption. In some embodiments, a remote user can grant access to a limited number of locations within the IoT environment to avoid privacy breaches. For example, a remote user can allow monitoring only of the living room and kitchen.
[0073] Reference Figure 6 An exemplary operation process 600 associated with remote monitoring of an IoT environment 200 is described. Figure 6 The operation can be performed by system 210, specifically by processor 302 and module 301 of system 210, but this disclosure is not limited thereto.
[0074] At operation 602, the entity is present at a location, for example, a cat is present in the living room and is meowing near the sofa. At operation 604, one or more non-imaging sensors detect the occurrence of an event associated with the entity (e.g., the cat). In the illustrated embodiment, an audio sensor detects the cat meowing at that location.
[0075] At operation 606, the surrounding environment at that location is monitored. That is, the operational status of the IoT devices is determined, and the current scene (surrounding environment) at that location is determined. At operation 608, the semantic information of the current scene at that location is determined. Semantic information can be generated based on the correlation between detected entities (e.g., a cat) and the operational status of one or more IoT devices at that location.
[0076] At operation 610, system 210 checks whether the event is a real or false event. This determination can be based on triggering conditions classified as real or false triggering conditions. If the event is a false event, the process does not continue further, and the event criteria used to detect the event can be modified at operation 612. If the event is a real event, a visual representation can be generated at operation 614. The visual representation can indicate the current scene at the location and can be annotated with enhanced information. In some embodiments, as described above, privacy rules can be used to generate the visual representation. The visual representation depicts semantic information at the location to provide accurate monitoring data to a remote user. At operation 616, the generated visual representation stream can be transmitted to the remote user via user equipment 220. In an embodiment, the visual representation can be encrypted before being sent to user equipment 220. In an embodiment, notifications and / or frame sequences can be displayed to the remote user. Therefore, the remote user can monitor entities at the location.
[0077] Figure 7 An exemplary process flow 700 for remote monitoring in an IoT environment according to an embodiment of the present disclosure is shown. In one embodiment, the steps of method 700 may be executed by system 210, for example, by processor 302 of system 210 in conjunction with multiple modules 301 and memory 304, but the present disclosure is not limited thereto.
[0078] At operation 702, one or more triggering conditions associated with the occurrence of an event at a location within the IoT environment are triggered, wherein the triggering conditions are associated with event criteria.
[0079] At operation 704, exemplary process flow 700 includes using one or more non-imaging sensors to detect one or more entities associated with the event present at that location.
[0080] At operation 706, exemplary process flow 700 includes determining the corresponding operational state of one or more IoT devices at that location when one or more entities are detected.
[0081] At operation 708, exemplary process flow 700 includes determining whether the one or more triggering conditions are false triggering conditions based on the relevance of the detected one or more entities to the corresponding operational states of the determined one or more IoT devices. In some embodiments, exemplary process flow 700 may include associating one or more entities with the corresponding operational states of one or more IoT devices by determining semantic information of the current scene indicating the location when the one or more entities are detected. In embodiments, a generative AI model may be used to generate the semantic information.
[0082] At operation 710, exemplary process flow 700 includes modifying event criteria based on determining that one or more triggering conditions are false triggering conditions.
[0083] Figure 8 An exemplary process flow 800 for remote monitoring in an IoT environment according to another embodiment of the present disclosure is shown. In one embodiment, the steps of method 800 may be executed by system 210, for example, by processor 302 of system 210 in conjunction with multiple modules 301 and memory 304, but the present disclosure is not limited thereto.
[0084] At operation 802, exemplary process flow 800 includes using one or more non-imaging sensors to determine the occurrence of an event at a location associated with the IoT environment.
[0085] At operation 804, exemplary process flow 800 includes detecting one or more entities associated with an event that are present at a location associated with the IoT environment.
[0086] At operation 806, exemplary process flow 800 includes determining the corresponding operational status of one or more Internet of Things (IoT) devices at that location.
[0087] At operation 808, exemplary process flow 800 includes determining semantic information based on the corresponding operational states of one or more detected entities and one or more IoT devices, the semantic information indicating the current scene at the location when one or more entities are detected.
[0088] At operation 810, exemplary process flow 800 includes generating a visual representation of a location associated with an IoT environment based on semantic information, the visual representation indicating the occurrence of an event at that location.
[0089] At operation 812, exemplary process flow 800 includes causing a visual representation of the generated indication event to be displayed on a user interface associated with a remote user, thereby enabling the remote user to monitor the location.
[0090] Although shown and described in a specific order Figure 7-8The steps described above are as described above, but according to various embodiments, these steps may occur in variations of these orders. Furthermore, with... Figure 7-8 Detailed descriptions of each step are already available in the relevant section. Figure 2-6 The relevant descriptions cover this, and for the sake of brevity, they are omitted here.
[0091] In an exemplary use case, an elderly person and a pet may be in the user's home. The user may have left the home to go to work. The home may not have cameras installed for monitoring. In the living room, a fall event can be detected via sensors integrated with speakers and a watch in the living room. The fall event can be detected as an elderly person possibly having fallen in the living room. In related technologies, a static notification can be sent to a remote user, and the user may be unsure whether the fall detection is for an elderly person or some other object that has fallen. However, in the method and system disclosed herein, the system detects that the event is not a false alarm and continues to generate a visual representation of the elderly person's fall, as well as the operational status of the IoT device. Therefore, the remote user can be fully aware of the event and make informed decisions.
[0092] In another exemplary use case, a child may be alone in the house, and the user may have left for work. The house may not have surveillance cameras installed. The child may be playing with toys in the living room. Motion can be detected via a presence sensor integrated with the television. In related technologies, static notifications can be sent to a remote user, who may be unsure what the child is doing. However, in the methods and systems disclosed herein, the system detects an event as a false alarm because the child is not in danger or distress. Therefore, the system does not trigger a notification to the remote user, thus preventing false alarms.
[0093] In yet another exemplary use case, a user may have left the house to do some work. The house may be empty, and there may be no one inside. Movement may be detected in the backyard. In related technologies, a static notification may be sent to a remote user, and the user may be unsure what is happening in the backyard. However, in the method and system disclosed herein, the system detects movement and generates a visual representation of the backyard with the detected entity and its surrounding environment. A remote user can view the visual representation to gain some context regarding the detected movement. As a result, user engagement is increased and potential intrusion is prevented.
[0094] In other exemplary use cases, the sound of flowing water can be detected by speakers in a house. The system can detect a leak in the house and decide to warn a remote user about the leak. A visual representation can then be generated to show the scene at that location (e.g., using ground images and house map details). Thus, a remote user can be alerted and an informed decision can be made.
[0095] This disclosure provides various technological advancements based on the aforementioned key features. The currently disclosed methods and systems enable intelligent surveillance of various locations. Locations without installed cameras can be monitored accurately and effectively. Furthermore, embodiments of this disclosure reduce the number of false alarms because remote users are not alerted for every event. Instead, the system and methods update event criteria to detect false events, thereby preventing false alarms. In other words, conditions leading to false alarms can be detected by associating the operational states of entities and IoT devices and subsequently reducing the number of false alarms. Therefore, user engagement is improved while the false alarm rate is significantly reduced.
[0096] Furthermore, remote users can have visualizations of events already detected at specific locations. The system displays an intuitive representation of the location to the user. For example, a user can receive a near real-time representation of AI-generated content at that location. Therefore, this system and method allow remote users to obtain continuous and reliable visualizations of events and related entities at monitored locations, particularly those not covered by cameras. Moreover, alerts and notifications sent to remote users include meaningful moving images and videos, rather than static notifications without any context. Additionally, predefined privacy rules can be used to maintain the privacy of the monitored entities.
[0097] While specific language has been used to describe the subject matter, it is not intended to create any limitation. It will be apparent to those skilled in the art that various working modifications can be made to the method to achieve the inventive concept taught herein. The accompanying drawings and the foregoing description provide examples of embodiments. Those skilled in the art will understand that one or more of the described elements can be well combined into a single functional element. Alternatively, certain elements may be divided into multiple functional elements. Elements from one embodiment may be added to another embodiment.
Claims
1. A method for remote monitoring in an Internet of Things (IoT) environment, the method comprising: Detect one or more triggering conditions associated with the occurrence of an event at a location within an IoT environment, wherein the one or more triggering conditions are associated with event criteria; Use one or more non-imaging sensors to detect one or more entities present at the location that are associated with the event; Determine the operational state of one or more IoT devices at the location when the one or more entities are detected; The correlation between the operational states of the one or more entities and the one or more IoT devices is used to determine whether the one or more triggering conditions are false triggering conditions; and The event criteria are modified based on determining that one or more triggering conditions are false triggering conditions.
2. The method according to claim 1, comprising: By using an artificial intelligence (AI) model to determine semantic information indicating the current scene at the location when the one or more entities are detected, the one or more entities are associated with the operational state of the one or more IoT devices.
3. The method according to claim 2, wherein, Determining whether the one or more triggering conditions are false triggering conditions includes: Use the one or more non-imaging sensors to separate one or more sources of the event; A confidence score is determined based on the AI model, wherein the confidence score is associated with each of the separated one or more sources; The confidence score is compared with a predetermined threshold; and Based on the comparison, the one or more triggering conditions are classified as false or true.
4. The method according to claim 2, comprising: The one or more triggering conditions are determined as true triggering conditions; When it is determined that one or more triggering conditions are true triggering conditions, a virtual visual representation of the location within the IoT environment is generated based on the semantic information, the virtual visual representation indicating the occurrence of the event at the location; as well as This enables the virtual visual representation to be displayed on a user interface associated with a remote user.
5. The method according to claim 2, wherein, Determining the semantic information includes: Determine a first set of parameters associated with the one or more entities, wherein the first set of parameters includes the location of the one or more entities and at least one of the actions performed by the one or more entities; Determine a second set of parameters associated with the operational state of the one or more IoT devices, wherein the second set of parameters indicates the background scene and surrounding environment at the location; and The semantic information is determined based on the first set of parameters and the second set of parameters.
6. The method according to claim 4, wherein, Generating the virtual visual representation includes: The semantic change conditions associated with the semantic information are detected based on multiple predefined privacy rules. The semantic information is updated based on the semantic change conditions, wherein the semantic change conditions are associated with removing a unit from the semantic information or modifying at least one of the units in the semantic information; and The virtual visual representation of the location is generated based on the updated semantic information.
7. The method according to claim 1, wherein, Detecting the one or more entities includes: using the one or more non-imaging sensors to detect the actions performed by the one or more entities, wherein the one or more non-imaging sensors include at least one of a motion sensor and an audio sensor.
8. The method according to claim 4, wherein, Generating the virtual visual representation of the location includes generating a frame sequence based on the semantic information.
9. A system for remote monitoring in an Internet of Things (IoT) environment, the system comprising: Memory; At least one processor communicatively coupled to the memory, the at least one processor being configured to: Detect one or more triggering conditions associated with the occurrence of an event at a location within the IoT environment, wherein the one or more triggering conditions are associated with event criteria; Use one or more non-imaging sensors to detect one or more entities present at the location that are associated with the event; Determine the operational state of one or more IoT devices at the location when the one or more entities are detected; The correlation between the operational states of the one or more entities and the one or more IoT devices is used to determine whether the one or more triggering conditions are false triggering conditions; and The event criteria are modified based on determining that one or more triggering conditions are false triggering conditions.
10. The system according to claim 9, wherein, The at least one processor is configured to: By using an artificial intelligence (AI) model to determine semantic information indicating the current scene at the location when the one or more entities are detected, the one or more entities are associated with the operational state of the one or more IoT devices.
11. The system according to claim 10, wherein, To determine whether the one or more trigger conditions are false trigger conditions, the at least one processor is configured to: Use the one or more non-imaging sensors to separate one or more sources of the event; A confidence score is determined based on the AI model, wherein the confidence score is associated with each of the separated one or more sources; The confidence score is compared with a predetermined threshold; and Based on the comparison, the one or more triggering conditions are classified as false or true.
12. The system according to claim 10, wherein, The at least one processor is configured to: The one or more triggering conditions are determined as true triggering conditions; When it is determined that one or more triggering conditions are true triggering conditions, a virtual visual representation of the location within the IoT environment is generated based on the semantic information, the virtual visual representation indicating the occurrence of the event at the location; as well as This enables the virtual visual representation to be displayed on a user interface associated with a remote user.
13. The system according to claim 10, wherein, In order to determine the semantic information, the at least one processor is configured to: Determine a first set of parameters associated with the one or more entities, wherein the first set of parameters includes the location of the one or more entities and at least one of the actions performed by the one or more entities; Determine a second set of parameters associated with the operational state of the one or more IoT devices, wherein the second set of parameters indicates the background scene and surrounding environment at the location; and The semantic information is determined based on the first set of parameters and the second set of parameters.
14. The system according to claim 12, wherein, In order to generate the virtual visual representation, the at least one processor is configured to: The semantic change conditions associated with the semantic information are detected based on multiple predefined privacy rules. The semantic information is updated based on the semantic change conditions, wherein the semantic change conditions are associated with removing a unit from the semantic information or modifying at least one of the units in the semantic information; and The virtual visual representation of the location is generated based on the updated semantic information.
15. The system according to claim 9, wherein, Detecting the one or more entities includes: using the one or more non-imaging sensors to detect the actions performed by the one or more entities, wherein the one or more non-imaging sensors include at least one of a motion sensor and an audio sensor.