Camera-based internet of things article search

By using machine learning models to analyze image and natural language input in the premises, the limitations and security issues of existing item tracking technology are solved, and efficient and secure item positioning and search services are achieved.

CN120234433APending Publication Date: 2025-07-01ROKU INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411912024.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-12-28
Filing Date
2024-12-24
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing item tracking technologies such as Bluetooth tracking devices and GPS/WiFi-based smartphone solutions have problems such as high cost, easy damage, high limitations and poor security, making it difficult to effectively help users locate small and easy-to-move items.

Method used

Using IoT cameras installed in the venue, we analyze image and natural language input through machine learning models, identify the location of items of interest, provide item search services, support online training and authentication mechanisms, and ensure user privacy and security.

Benefits of technology

It realizes efficient and secure positioning of items in the place without relying on the cloud, reduces dependence on the cloud, protects user privacy, and provides flexible item search and positioning functions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234433A_ABST
    Figure CN120234433A_ABST
Patent Text Reader

Abstract

Disclosed herein are system, method, and / or computer program product embodiments, and / or combinations and sub-combinations thereof, for providing item search services for a venue that includes a set of Internet of Things (IoT) cameras. An example embodiment operates by receiving a first user input with respect to an item of interest via a user interface of an item search service, where the first user input includes one or more of a voice input or a textual input; accessing a plurality of images of the venue captured by the set of IoT cameras; executing a machine learning model to identify one or more images of the plurality of images including the item of interest based at least on the first user input; generating item search results based on the identified one or more images; and providing item search results via a user interface of the item search service.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to item search services for a venue, and more particularly, to item search services for a venue implemented using one or more Internet of Things (IoT) cameras. Background Art

[0002] Keeping track of items within a venue can be challenging, especially if the items are relatively small and frequently moved. For example, in a home, essential or important items such as wallets, purses, bags, keys, smartphones, laptops, or remote controls can easily be misplaced, and the item owner needs to conduct a manual search throughout the home to retrieve the item. Such manual searches can be inconvenient, time-consuming, and stressful.

[0003] There are some technologies that help users retrieve lost or misplaced items. For example, there are battery-powered Bluetooth tracking devices that can be attached to items and then used to track the location of the items. However, some significant drawbacks of Bluetooth tracking devices include, but are not limited to: a separate Bluetooth tracking device must be purchased for each item that the user may wish to track, which can be expensive; some Bluetooth tracking devices do not allow battery replacement, which means that when the battery runs out, the user must purchase a brand-new tracking device; for Bluetooth tracking devices that can replace the battery, having to replace the battery regularly can be expensive and inconvenient; Bluetooth tracking devices can only be used to track items that have a form factor to which the tracking device can be attached; the connection means between the Bluetooth tracking device and the item can fail or be damaged; the item is only locatable within the range of another Bluetooth device; if the Bluetooth tracking device is small enough to be swallowed, or if the tracking device uses a button cell or coin cell and the device housing is not secure, the Bluetooth tracking device can pose a danger to children; and Bluetooth tracking devices can be misused to lock and track users and monitor their locations.

[0004] As another example of a technology that can help users retrieve lost or misplaced items, some smartphones may be configured to use built-in GPS-based and / or WiFi-based tracking technologies to help users locate the smartphones. However, this solution is very limited because it only applies to a very narrow category of items that include built-in location tracking technology. Summary of the Invention

[0005] The present disclosure provides system, apparatus, article, method, and / or computer program product embodiments for providing an item search service for a premises including a set of IoT cameras, and / or combinations and sub-combinations thereof. Example embodiments perform operations including: receiving, via a user interface of the item search service, a first user input regarding an item of interest, wherein the first user input includes one or more of a voice input or a text input; accessing a plurality of images of the premises captured by the set of IoT cameras; executing a machine learning model to identify, based at least on the first user input, one or more of the plurality of images that include the item of interest, generating an item search result based on the identified one or more images, and providing the item search result via the user interface of the item search service.

[0006] In some aspects, the first user input includes a natural language input, and the machine learning model includes a multimodal machine learning model trained on a set of images and natural language text respectively associated with each image in the set of images.

[0007] In some aspects, the operations further include receiving, via the user interface of the item search service, a second user input that specifies an image of the item of interest and a label assigned by the user to the item of interest, and training the machine learning model with the image of the item of interest and the label assigned to the item of interest.

[0008] In some aspects, the receiving, accessing, executing, generating, and providing operations are performed by one or more devices located within the premises.

[0009] In some aspects, the operations further include selecting a machine learning model from a plurality of different machine learning models, wherein each machine learning model in the plurality of different machine learning models is trained or fine-tuned for one of a specific premises type or a specific demographic.

[0010] In some aspects, the operations further include authenticating a user of the item search service and determining, based on the authentication, that the user is an authorized user of the item search service, and performing one or more of the receiving, accessing, executing, generating, and providing operations in response to determining that the user is an authorized user of the item search service.

[0011] In some aspects, the operations further include receiving, via the user interface of the item search service, a second user input that specifies an item that should not be searchable, and in response to receiving the second user input, applying a content filter that prevents the item search service from searching for items that should not be searchable or prevents the item search service from returning item search results for items that should not be searchable.

[0012] In some aspects, the operation further includes determining the identity of a user of the item search service, and performing the machine learning model includes performing the machine learning model to identify one or more images among a plurality of images that include an item of interest based at least on a first user input and the identity of the user.

[0013] In some aspects, generating an item search result based on the identified one or more images includes one or more of the following operations: generating a verbal or textual description of the location of the item of interest based on the identified one or more images, or generating an image showing the location of the item of interest based on the identified one or more images. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings are incorporated herein and constitute a part of this specification.

[0015] Figure 1 A block diagram of a multimedia environment according to some embodiments is shown.

[0016] Figure 2 A block diagram of a streaming media device according to some embodiments is shown.

[0017] Figure 3 A block diagram of a system for providing an item search service for a premises including a set of IoT cameras according to some embodiments is shown.

[0018] Figure 4 A flowchart of a method for providing an item search service for a premises including a set of IoT cameras according to some embodiments is shown.

[0019] Figure 5 A flowchart of a method for training a machine learning model for implementing an item search service for a premises according to some embodiments is shown.

[0020] Figure 6 A flowchart of a method for providing a secure item search service for a premises according to some embodiments is shown.

[0021] Figure 7 A flowchart of a method for preventing an item search service for a premises from searching for a specified item according to some embodiments is shown.

[0022] Figure 8 A flowchart of a method for providing an item search service for a premises that supports user-specific query interpretation according to some embodiments is shown.

[0023] Figure 9 An example computer system that can be used to implement various embodiments is shown.

[0024] In the accompanying drawings, like reference numerals generally denote like or similar elements. Additionally, generally, the leftmost digit of a reference numeral identifies the drawing in which that reference numeral first appears. Detailed Description

[0025] Embodiments of systems, apparatuses, devices, methods, and / or computer program products are provided herein for providing an item search service for a premises that includes a set of IoT cameras, and / or combinations and sub - combinations thereof, which address one or more of the above - mentioned problems associated with conventional solutions for tracking items of interest to a user. Providing an item search service may include (i) receiving a first user input regarding an item of interest via a user interface of the item search service, where the first user input includes one or more of a voice input or a text input; (ii) accessing a plurality of images of the premises captured by the set of IoT cameras; (iii) executing a machine - learning model to identify one or more images among the plurality of images that include the item of interest, at least based on the first user input; (iv) generating an item search result based on the identified one or more images; and (v) providing the item search result via the user interface of the item search service.

[0026] As will be described herein, the item search service can advantageously utilize various IoT cameras installed in or around a premises to assist a user in locating an item of interest. In some scenarios, by way of example only, the item search service can utilize various IoT cameras that have been installed as part of a home security system or a home automation system in a home, and / or various IoT cameras integrated within various user devices located throughout the home.

[0027] In some aspects, the item search service can utilize a voice input or a text input provided by a user as an input to a multi - modal machine - learning model that is capable of using such input to identify one or more images collected by the IoT cameras that include the item of interest. In a context where the machine - learning model is trained on a large number of images and associated natural - language descriptions, the machine - learning model can provide useful results for a natural - language query submitted by a user (e.g., "Where did I put my keys?"), even if the item that the user is looking for may not have been observed by the machine - learning model during the training process.

[0028] In some aspects, the machine - learning model used to provide the item search service can be selected from a plurality of different machine - learning models, where each of the plurality of different machine - learning models is trained or fine - tuned for one of a specific type of premises or a specific user demographic. This feature enables the selection of a machine - learning model that will provide optimal item - tracking performance for a specific premises and / or a specific user.

[0029] In some aspects, a user can provide input via a user interface of an item search service that specifies an image of an item of interest and a label assigned by the user to the item of interest. For example, the user can say "These are my keys" while holding up the keys in front of an IoT camera, or can label the keys in an image captured by the IoT camera and say or type "These are my keys". Then, the specified image of the item of interest and the label assigned to the item of interest can be used to train (e.g., train, retrain, or fine-tune) a machine learning model such that the model subsequently associates an image of the item with the label.

[0030] In some aspects, the item search service can infer a label for a particular item based on an association between a user and the particular item. For example, if the user who submits an image of a smartphone for training purposes is Bob, or if the item search service determines that the user Bob is most often observed holding the particular smartphone, the service can label the item as "Bob's smartphone". Then, these user-specified item labels can be used to train (e.g., train, retrain, or fine-tune) a machine learning model such that when a particular user (e.g., Bob) says "Where did I put my smartphone?", the machine learning model can identify the particular user's smartphone (e.g., Bob's smartphone) in an IoT camera image.

[0031] In some aspects, the machine learning model and other components of the item search service can be installed and executed on one or more devices (e.g., edge) within a premises rather than on a computing device (e.g., cloud) outside the premises. This implementation can be used to protect user privacy and data security, such that the processing requirements of the service can be distributed across multiple end-user devices and increases resilience as the service can still operate even when the internet connection is lost without relying on the cloud.

[0032] In some aspects, the item search service can include an authentication feature that ensures only authorized users can use the service and / or use the service to search for a particular item. This feature can advantageously prevent the system from being misused by bad actors (e.g., home intruders) and generally prevent certain users from using the service to locate certain items within a premises (e.g., prevent children in a home from accessing certain dangerous or prohibited items).

[0033] In some aspects, a user may provide, via a user interface of an item search service, an input specifying an item that should not be searched by the service. In response, the service may then apply a content filter that blocks the item search service from searching for the item or that blocks the item search service from providing item search results regarding the item. This functionality can advantageously enable users of the service to selectively "hide" certain items (e.g., wall safes, handguns) that the user does not want to be located.

[0034] In some aspects, an item search service may generate item search results in a manner that facilitates a user to easily track or retrieve items of interest. For example, based on one or more images identified by a machine learning model that include an item of interest, the item search service may generate a verbal or textual description of the location of the item of interest and / or generate an image showing the location of the item of interest, and such verbal / textual description and / or image showing the location of the item of interest may be presented to the user via a user interface of the item search service.

[0035] These and various other features and advantages of an item search service for an IoT camera-based premises will be described in detail herein with reference to various embodiments. The various embodiments of the present disclosure may be implemented using Figure 1 the multimedia environment 102 shown in Figure 1 and / or may be part of the multimedia environment 102 shown in

[0036] The multimedia environment

[0037] Figure 1 shows a block diagram of a multimedia environment 102 according to some embodiments. In a non-limiting example, the multimedia environment 102 may be for streaming media. However, the present disclosure is applicable to any type of media (instead of or in addition to streaming media), as well as any mechanism, means, protocol, method, and / or process for distributing media.

[0038] The multimedia environment 102 may include one or more media systems 104. The media system 104 may represent a home room, kitchen, backyard, home theater, school classroom, library, car, boat, bus, airplane, movie theater, stadium, auditorium, park, bar, restaurant, or any other location or space that needs to receive and play streaming content. One or more users 132 may operate the media system 104 to select and consume content.

[0039] Each media system 104 may include one or more media devices 106, and each of the one or more media devices 106 is coupled to one or more display devices 108. It should be noted that unless otherwise specified herein, terms such as "coupled", "connected to", "attached", "linked", "combined", etc. and similar terms may refer to physical, electrical, magnetic, logical, etc. connections.

[0040] The media device 106 may be, by way of example only, a streaming device, a DVD or Blu-ray device, an audio / video playback device, a cable TV box, and / or a digital video recording device. The display device 108 may be, by way of example only, a monitor, a television (TV), a computer, a smartphone, a tablet, a wearable device (e.g., a watch or glasses), a household appliance, an Internet of Things (IoT) device, and / or a projector. In some embodiments, the media device 106 can be part of its corresponding display device 108, integrated with its corresponding display device 108, operably coupled to its corresponding display device 108, and / or connected to its corresponding display device 108.

[0041] Each media device 106 may be configured to communicate with the network 118 via the communication device 114. The communication device 114 may include, for example, a cable modem or a satellite TV transceiver. The media device 106 may communicate with the communication device 114 via the link 116, where the link 116 may include a wireless connection (e.g., WiFi) and / or a wired connection.

[0042] In various embodiments, the network 118 may include, but is not limited to, a wired and / or wireless intranet, extranet, Internet, cellular, Bluetooth, infrared, and / or any other short-range, long-range, local, regional, global communication mechanism, means, method, protocol, and / or network, and any combination thereof.

[0043] The media system 104 may include a remote control 110. The remote control 110 can be any component, part, device, and / or method for controlling the media device 106 and / or the display device 108, such as a remote control, a tablet computer, a laptop computer, a smart phone, a wearable device, a screen control, an integrated control button, an audio control, or any combination thereof, by way of example only. In an embodiment, the remote control 110 wirelessly communicates with the media device 106 and / or the display device 108 using cellular, Bluetooth, infrared, etc., or any combination thereof. The remote control 110 may include a microphone 112, which will be further described below.

[0044] The multimedia environment 102 may include a plurality of content servers 120 (also referred to as content providers, channels, or sources 120). Although Figure 1 only one content server 120 is shown, in fact, the multimedia environment 102 may include any number of content servers 120. Each content server 120 may be configured to communicate with the network 118.

[0045] Each content server 120 may store content 122 and metadata 124. The content 122 may include music, video, movies, television programs, multimedia, images, still pictures, text, graphics, game applications, advertisements, program content, public service content, government content, local community content, software, and / or any combination of any other content or data objects in electronic form.

[0046] In some embodiments, the metadata 124 includes data about the content 122. For example, the metadata 124 may include associative information or ancillary information indicating or relating to the writer, director, producer, composer, artist, actor, summary, chapter, production, history, year, trailer, alternative version, related content, application, and / or any other information related to or associated with the content 122. The metadata 124 may also or alternatively include links to any such information related to or associated with the content 122. The metadata 124 may also or alternatively include one or more indexes of the content 122, such as, but not limited to, a trick mode index.

[0047] The multimedia environment 102 may include one or more system servers 126. The system server 126 is operable to support the media device 106 from the cloud. It should be noted that the structural and functional aspects of the system server 126 may exist in whole or in part in the same or different system servers 126.

[0048] The media device 106 can be present in thousands or millions of media systems 104. Thus, the media device 106 can be used for itself in a crowdsourcing embodiment, and thus, the system server 126 can include one or more crowdsourcing servers 128.

[0049] For example, using information received from media devices 106 in thousands or millions of media systems 104, one or more crowdsourcing servers 128 can identify similarities and overlaps between closed caption requests issued by different users 132 watching a particular movie. Based on such information, one or more crowdsourcing servers 128 can determine that turning on the closed caption can enhance the viewing experience of users at a particular part of the movie (e.g., when it is difficult to hear the movie's soundtrack), and turning off the closed caption can enhance the viewing experience of users at other parts of the movie (e.g., when the display of the closed caption obstructs key visual aspects of the movie). Thus, one or more crowdsourcing servers 128 are operable to cause the closed caption to be automatically turned on and / or off during future streaming of the movie.

[0050] The system server 126 can also include an audio command processing module 130. As described above, the remote control 110 can include a microphone 112. The microphone 112 can receive audio data from the user 132 (as well as other sources, such as the display device 108). In some embodiments, the media device 106 can be audio-responsive, and the audio data can represent verbal commands from the user 132 for controlling the media device 106 and other components in the media system 104 (e.g., the display device 108).

[0051] In some embodiments, the audio data received by the microphone 112 in the remote control 110 is sent to the media device 106, which then forwards it to the audio command processing module 130 in the system server 126. The audio command processing module 130 can operate to process and analyze the received audio data to identify the verbal commands of the user 132. The audio command processing module 130 can then forward the verbal commands back to the media device 106 for processing.

[0052] In some embodiments, alternatively or additionally, the audio data can be processed and analyzed by an audio command processing module 216 in the media device 106 (see Figure 2 ). Then, the media device 106 and the system server 126 can cooperate to select one of the verbal commands (the verbal commands identified by the audio command processing module 130 in the system server 126, or the verbal commands identified by the audio command processing module 216 in the media device 106) for processing.

[0053] Figure 2A block diagram of an example media device 106 according to some embodiments is shown. The media device 106 may include a streaming module 202, a processing module 204, a memory / buffer 208, and a user interface module 206. As described above, the user interface module 206 may include an audio command processing module 216.

[0054] The media device 106 may also include one or more audio decoders 212 and one or more video decoders 214.

[0055] Each audio decoder 212 may be configured to decode audio in one or more audio formats, such as but not limited to AAC, HE-AAC, AC3 (Dolby Digital), EAC3 (Dolby Digital Plus), WMA, WAV, PCM, MP3, OGG GSM, FLAC, AU, AIFF, and / or VOX, by way of example only.

[0056] Similarly, each video decoder 214 may be configured to decode video in one or more video formats, such as but not limited to MP4 (mP4, m4a, m4v, f4v, f4a, m4b, m4r, f4b, mov), 3GP (3gp, 3gp2, 3g2, 3gpp, 3gpp2), OGG (ogg, oga, ogv, ogx), WMV (wmv, wma, asf), WEBM, FLV, AVI, QuickTime, HDV, MXF (OP1a, OP-Atom), MPEG-TS, MPEG-2 PS, MPEG-2 TS, WAV, Broadcast WAV, LXF, GXF, and / or VOB, by way of example only. Each video decoder 214 may include one or more video codecs, such as but not limited to H.263, H.264, H.265, AVI, HEV, MPEG1, MPEG2, MPEG-TS, MPEG-4, Theora, 3GP, DV, DVCPRO, DVCPRO, DVCProHD, IMX, XDCAM HD, XDCAM HD422, and / or XDCAM EX, by way of example only.

[0057] Now refer to Figure 1 and Figure 2Both, in some embodiments, user 132 may interact with media device 106 via, for example, remote control 110. For example, user 132 may use remote control 110 to interact with user interface module 206 of media device 106 to select content such as movies, TV shows, music, books, applications, games, etc. Streaming module 202 of media device 106 may request the selected content from one or more content servers 120 via network 118. One or more content servers 120 may send the requested content to streaming module 202. Media device 106 may send the received content to display device 108 for playing the content to user 132.

[0058] In streaming embodiments, streaming module 202 may send content to display device 108 in real-time or near real-time while receiving content from one or more content servers 120. In non-streaming embodiments, media device 106 may store the content received from one or more content servers 120 in storage / buffer 208 for playing the content on display device 108 at a later time.

[0059] Item Search Based on IoT Cameras

[0060] Figure 3 A block diagram of a system 300 for providing an item search service for a venue including a set of IoT cameras is shown according to some embodiments. As Figure 3 shown, system 300 includes venue 302 and item search service 350. Each of these aspects of system 300 will now be described.

[0061] Location 302

[0062] As Figure 3 shown, system 300 includes venue 302 in which there are multiple IoT devices 306, 308, 310, and 312. Venue 302 may include, for example, but not limited to, a home, office, building, factory, warehouse, bar, restaurant, cinema, stadium, auditorium, car, bus, ship, or any other structure, location, or space where IoT devices may be present. Although for illustrative purposes only four IoT devices present in venue 302 are shown, it should be understood that venue 302 may include any number of IoT devices, including dozens, hundreds, or even thousands of IoT devices.

[0063] As used herein, the term "IoT device" is intended to broadly cover any device capable of digital communication with another device. For example, a device capable of digital communication with another device can include an IoT device as used herein, even if such communication is not via the Internet.

[0064] Each of the IoT devices 306, 308, 310, and 312 may include devices such as smartphones, laptops, notebooks, tablets, netbooks, desktop computers, video game consoles, set-top boxes, OTT streaming players. Additionally, each of the IoT devices 306, 308, 310, and 312 may include so-called "smart home" devices, e.g., smart bulbs, smart switches, smart refrigerators, smart washing machines, smart dryers, smart coffee makers, smart alarm clocks, smart smoke alarms, smart carbon monoxide detectors, smart security sensors, smart doorbell cameras, smart indoor or outdoor cameras, smart door locks, smart thermostats, smart plugs, smart TVs, smart speakers, smart remote controls, or voice controllers. Additionally, each of the IoT devices 306, 308, 310, and 312 may include wearable devices such as watches, fitness trackers, health monitors, smart pacemakers, or extended reality headsets. Additionally, each of the IoT devices 306, 308, 310, and 312 may include drones or other devices that are guided or otherwise navigated within or around the premises 302, or robots or other devices that are capable of moving on their own within the premises 302. However, these are only used as examples and are not intended to be limiting.

[0065] The IoT devices 306, 308, 310, and 312 may be communicatively connected to a local area network (LAN) 340 via suitable wired and / or wireless connections. The LAN 340 may be implemented using a hub-and-spoke or star topology. For example, according to such an implementation, each of the IoT devices 306, 308, 310, and 312 may be connected to a router via a respective Ethernet cable, wireless access point (AP), or IoT device hub. The router may include a modem that enables the router to act as an interface between the entity connected to the LAN 340 and an external wide area network (WAN) (e.g., the Internet). Alternatively, the LAN 340 is implemented using a full-mesh or partial-mesh network topology. According to the full-mesh network topology, each IoT device of a group of IoT devices in the premises 302 may be directly connected to each of the other IoT devices in the premises, such that each IoT device in a group of IoT devices can communicate with each of the other IoT devices without a router. According to the partial-mesh network technology, only some of the IoT devices in the premises 302 may be directly connected to other IoT devices, and indirect communication between unconnected IoT device pairs may be carried out through one or more intermediate devices. The mesh network implementation of the LAN 340 may also be connected to an external WAN (e.g., the Internet) via a router. However, these are only examples, and other technologies for implementing the LAN 340 may be used.

[0066] As Figure 3 Figure 3 As further shown, the IoT device 306 may include one or more processors 328, one or more sensors 330, a sensor data collector 332, one or more actuators 334, and one or more communication interfaces 336. The one or more processors 328 may include one or more central processing units (CPUs), microcontrollers, microprocessors, signal processors, application specific integrated circuits (ASICs), and / or other physical hardware processor circuits for performing tasks such as program execution, signal encoding, data processing, input / output processing, power control, and / or other functions.

[0067]

[0067] The one or more sensors 330 may include one or more devices or systems for detecting and responding to (e.g., measuring, recording) objects and events in the physical environment of the IoT device 306. By way of example and not limitation only, the one or more sensors 330 may include one or more of the following: cameras or other optical sensors, microphones or other audio sensors, radar systems, LiDAR systems, Wi-Fi sensing systems, global positioning system (GPS) sensors, temperature sensors, pressure sensors, proximity sensors, accelerometers, gyroscopes, magnetometers, infrared sensors, gas sensors, and / or smoke sensors. An IoT device including a sensor in the form of a camera may also be referred to herein as an IoT camera.

[0068]

[0068] The sensor data collector 332 may be configured to collect sensor data from the one or more sensors 330 of the IoT device 306 and provide such sensor data to the item search service 350 for performing a search for items within or around the premises 302 and for providing other features. For example, the sensor data collector 332 may continuously, periodically, or intermittently collect images captured by the camera of the IoT device 306 and provide such images to the item search service 350 such that the item search service 350 can perform a search for items. The sensor data collector 332 may provide sensor data to the item search service 350 by storing the sensor data in an IoT device sensor data storage device 364 accessible to both the IoT device 306 and the item search service 350. The data storage device 364 is intended to represent any physical storage device or system suitable for storing data.

[0069] One or more actuators 334 may include one or more devices or systems operable to effect a change in the physical environment of the IoT device 306. By way of example and not limitation only, one or more actuators 334 may include components that perform functions such as connecting a device to a power source, disconnecting a device from a power source, turning a light on or off, adjusting the brightness or color of a light, turning a sound alert on or off, adjusting the volume of a sound alert, initiating a call to a security service, turning a heating or cooling system on or off, adjusting a target temperature associated with a heating or cooling system, locking or unlocking a door, ringing a doorbell, starting to capture video or audio, changing the channel or configuration of a television, adjusting the volume of an audio output device, and the like.

[0070] One or more communication interfaces 336 may include components adapted to enable the IoT device 306 to communicate wirelessly with other devices via respective wireless protocols. The communication interface 336 may include, for example but not limited to, one or more of the following interfaces: a Wi-Fi interface that enables the IoT device 306 to communicate wirelessly with an access point or other Wi-Fi-enabled remote device according to one or more of the wireless network protocols based on the IEEE (Institute of Electrical and Electronics Engineers) 802.11 standard family; a cellular interface that enables the IoT device 302 to communicate wirelessly with remote devices via one or more cellular networks; a Bluetooth interface that enables the IoT device 304 to communicate in short-range wireless communication with other Bluetooth-enabled devices; or a Zigbee interface that enables the IoT device 306 to communicate wirelessly with other Zigbee-enabled devices.

[0071] One or more communication interfaces 336 may additionally or alternatively include components adapted to enable the IoT device 306 to communicate with other devices via a wired connection via respective wired protocols (e.g., universal serial bus (USB) connection and protocol or Ethernet connection and protocol).

[0072] As Figure 3 As further shown, the IoT device 306 may include an item search user interface (UI) 326. The item search UI 326 may include a UI that enables a user to interact with an item search service 350 to perform a search for items within the premises 302 and receive the results of that search. The item search UI 326 may also include a UI that enables a user to invoke other functions of the item search service 350, such as online training, content filtering, authentication, and item-based automation and item-based monitoring and alerts. These features will be described in more detail herein.

[0073] The item search UI 326 may include one or more input devices (e.g., one or more of a remote control, a set of buttons, a keypad, a keyboard, a mouse, a touchpad, a touch screen, a microphone, etc.), one or more output devices (e.g., one or more of a display screen, a speaker, etc.), and software components executed by one or more processors 328, the software components being configured to accept user input provided using the one or more input devices and present output to the user using the one or more output devices. Such software components of the item search UI 326 may also be configured to communicate with the item search service 350 using a suitable application programming interface (API) to invoke features of the item search service, pass user input thereto, receive output therefrom, etc. According to an embodiment, the item search UI 326 may include a graphical UI (GUI), a menu-driven UI, a touch UI, a voice UI, a form-based UI, a natural language UI, etc.

[0074] Each of the IoT devices 308, 310, and 312 may include components similar to those shown for the IoT device 306. Thus, for example, each of the IoT devices 308, 310, and 312 may include one or more processors, one or more sensors, a sensor data collector, one or more actuators, one or more communication interfaces, and an item search UI.

[0075] The user device 304 is intended to represent a personal computing device or media device associated with a user. For example, in an embodiment where a multimedia environment exists in the venue 302, the user device 304 may include the media device 106, and the user interface of the user device 304 may be presented to the user via the display device 108. The user device 304 may also include a smart phone, a laptop computer, a notebook computer, a tablet computer, a netbook, a desktop computer, a video game console, or a wearable device (e.g., a smart watch or an extended reality headset). The user device 304 may include one or more processors 316, one or more sensors 318, a sensor data collector 320, one or more actuators 322, one or more communication interfaces 324, and an item search UI 314.

[0076] One or more processors 316 may include one or more CPUs, a microcontroller, a microprocessor, a signal processor, an ASIC, and / or other physical hardware processor circuits for performing tasks such as program execution, signal encoding, data processing, input / output processing, power control, and / or other functions.

[0077] One or more sensors 318 may include one or more devices or systems for detecting and responding to (e.g., measuring, recording) objects and events in the physical environment of the user device 304. For example, one or more sensors 318 may include one or more of the sensor types described above with reference to the one or more sensors 330 of the IoT device 306.

[0078] The sensor data collector 318 may be configured to collect sensor data from one or more sensors 318 of the user device 304 and provide such sensor data to the item search service 350 for use in performing a search for items within or around the premises 302, as well as to provide other features described herein. For example, the sensor data collector 318 may continuously, periodically, or intermittently collect images captured by the camera of the user device 304 and provide such images to the item search service 350 so that the item search service 350 can perform a search for items. The sensor data collector 320 may provide sensor data to the item search service 350 by storing the sensor data in the IoT device sensor data memory 364 accessible to both the user device 304 and the item search service 350.

[0079] One or more actuators 322 may include one or more devices or systems operable to effect a change in the physical environment of the user device 304. For example, one or more actuators 322 may include one or more of the actuator types described above with reference to the one or more actuators 334 of the IoT device 306.

[0080] One or more communication interfaces 324 may include components adapted to enable the user device 304 to communicate with other devices via a wired or wireless communication medium using a corresponding wired or wireless communication protocol. For example, one or more communication interfaces 324 may include one or more of the communication interface types described above with reference to the one or more communication interfaces 336 of the IoT device 306.

[0081] The item search UI 314 may include a UI that enables a user to interact with the item search service 350 to perform a search for items within the premises 302 and receive the results of such a search. The item search UI 314 may also include a UI that enables a user to invoke other functions of the item search service 350, such as online training, content filtering, authentication, and item-based automation and item-based monitoring and alerting. These features will be described in more detail herein.

[0082] The item search UI 314 may include one or more input devices (e.g., one or more of a remote control, a set of buttons, a keypad, a keyboard, a mouse, a touchpad, a touch screen, a microphone, etc.), one or more output devices (e.g., one or more of a display screen, a speaker, etc.), and software components executed by one or more processors 316, which are configured to accept user input provided using the one or more input devices and present output to the user using the one or more output devices. Such software components of the item search UI 314 may also be configured to communicate with the item search service 350 using appropriate APIs to invoke features of the item search service 350, pass user input to it, receive output from it, etc. According to embodiments, the item search UI 314 may include a GUI, a menu-driven UI, a touch UI, a voice UI, a form-based UI, a natural language UI, etc.

[0083] Although Figure 3 only a single user device 304 is shown in [FIGURE REFERENCE], it should be understood that there may be multiple user devices in the premises 302, and each such user device may be configured in a manner similar to the user device 304.

[0084] Furthermore, although for simplicity Figure 3 only a single premises 302 is shown in [FIGURE REFERENCE], which has an associated user device 304 and a set of IoT devices 306, 308, 310, and 312, it should be understood that the system 300 may include any number of premises (including hundreds, thousands, tens of thousands, hundreds of thousands, millions, or hundreds of millions of premises), each having its own one or more user devices and IoT devices, and the item search service 350 may utilize the IoT devices deployed in the premises respectively to provide IoT camera-based item search services to each such premise.

[0085] Item Search Service 350

[0086] As Figure 3 further shown, the system 300 includes an item search service 350. As will be discussed below, the item search service 350 may utilize images captured by various IoT cameras present in the premises 302 to assist a user in locating items of interest in the premises 302 and provide other features.

[0087] The item search service 350 (and each of its various components) can be implemented as processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. The item search service 350 can be implemented by one or more devices (e.g., one or more servers) that are remote from the premises 302 but communicatively connected to the premises 302 via one or more networks (e.g., communicatively connected to the LAN 340). Alternatively, the item search service 350 can be implemented by a device within the premises 302, e.g., by the user device 304 or by one of the IoT devices 306, 308, 310, and 312. Additionally, the item search service 350 can be implemented in a distributed manner by two or more remote and / or local devices.

[0088] In some embodiments, the item search service 350 is installed and executed only on one or more devices (e.g., at the edge) within the premises 302 and does not rely on any computing devices (e.g., the cloud) external to the premises 302. This embodiment can protect the privacy and data security of users by avoiding the transmission of personal user data outside the premises 302, enabling the processing requirements of the item search service 350 to be distributed across multiple end-user devices, and increasing resiliency since the item search service 350 can still operate even when the Internet connection is lost without relying on the cloud.

[0089] As Figure 3 Further illustrated, the item search service 350 includes a plurality of components, which include an item search module 352, an online training module 354, a content filtering module 356, an authentication module 358, and an item-based automation / alarm module 360. Each of these components of the item search service 350 will now be described.

[0090] Item Search Module 352

[0091] The item search module 352 can perform a search for items of interest located within or around the premises 302. The item search module 352 can perform such an item search based on user input submitted to the item search service 350 via the UI (e.g., the item search UI 314 or the item search UI 326). For example, the user input can include voice input and / or text input regarding the item of interest. In some embodiments, the voice input or text input can include natural language input. For example, the natural language input might be "Where is my smartphone?" or "Help me find my keys."

[0092] To perform item search, the item search module 352 can access multiple images of the venue 302 captured by one or more IoT devices (e.g., one or more of IoT device 306, IoT device 308, IoT device 310, or IoT device 312) present in the venue 302. The multiple images can also include images captured by one or more user devices (e.g., user device 304). As described previously, these images can be captured by these devices and then stored in the IoT device sensor data storage 364, where they can be accessed by the item search service 350.

[0093] To perform item search, the item search module 352 can also execute a machine learning model 362, which is included within or otherwise accessible to the item search module 352. The machine learning model 362 can be used to identify one or more images among the multiple images of the venue 302 that include the item of interest, at least based on user input. For example, the machine learning model 362 can compare the encoded representation of the user input with the encoded representation of each image among the multiple images to identify which images (if any) might include the item that the user input pertains to, by way of example only. To utilize the machine learning model 362, the item search module 352 can first process the voice or text input provided by the user to put it in a form suitable for processing by the machine learning model 362. Similarly, the item search module 352 can also process each of the multiple images of the venue 302 to put each image in a form suitable for processing by the machine learning model 362.

[0094] For example, the machine learning model 362 can include a multimodal machine learning model trained on both images and text inputs. For example, the machine learning model 362 can include a multimodal machine learning model trained on a relatively large dataset of image-text pairs (e.g., hundreds of millions of image-text pairs), where the text associated with a given image includes natural language text. For example, the image-text pairs can include images from web pages and the natural language captions accompanying the images on the web pages. A non-limiting example of such a multimodal machine learning model is OpenAI ®Developed contrastive language-image pre-training (CLIP) neural networks. Models such as CLIP are designed to be used in a zero-shot manner, which means the model can associate observed and unobserved classes by means of some form of auxiliary information (encoding observable discriminative attributes of items) to identify item classes not observed during training. In this context, this means that even if the machine learning model 362 may not have observed a particular item that the user is looking for during training, the machine learning model 362 can advantageously provide useful results for a natural language query submitted by the user (e.g., "Where did I put my keys?").

[0095] Furthermore, leveraging the machine learning model 362 that has been trained on both visual and natural language modalities can greatly enhance the user experience, as the user can use natural language prompts to request item searches without having to follow a specific set of labels to perform the search. This can make invoking item searches a simple and straightforward experience for the user.

[0096] Alternatively, the machine learning model 362 can include a computer vision detection model trained to identify a predefined set of item classes. For example, such a computer vision detection model can be trained on a set of images, where each image is labeled (e.g., by a human) with one of a predefined set of item classes. According to this implementation, the predefined set of item classes can include common items that a user may want to search for in a venue.

[0097] In some implementations, the item search module 352 can select the machine learning model 362 from a plurality of different machine learning models, where each of the plurality of different machine learning models is trained or fine-tuned for one of a specific venue type or a specific user demographic. Such features can enable the selection of the machine learning model that will provide the best item tracking performance for a particular venue and / or a particular user.

[0098] For example, the item search module 352 may select the machine learning model 362 from a machine learning model trained or fine-tuned to identify common items in a home and a machine learning model trained or fine-tuned to identify common items in an office. As another example, the item search module 352 may select the machine learning model 362 from a machine learning model trained or fine-tuned to identify common items in a high-rise apartment and a machine learning model trained or fine-tuned to identify common items in a suburban residence. As yet another example, the item search module 352 may select the machine learning model 362 from a model trained or fine-tuned to identify common items used by a first age group and a model trained or fine-tuned to identify common items used by a second age group. As yet another example, the item search module 352 may select the machine learning model 362 from a model trained or fine-tuned to identify common items in a first geographic location and a model trained or fine-tuned to identify common items in a second geographic location.

[0099] The item search module 352 may select the machine learning model 362 from multiple different machine learning models based on information about the location 302 obtained by the item search service 350 (e.g., location type) and / or information about the user associated with the location 302 (e.g., demographic information about the user). For example, as part of registering or configuring the item search service 350, such information may be provided by the user associated with the location 302. In another embodiment, the item search service 350 may present a list of different models to the user associated with the location 302 and enable the user to select a model from the list. The item search module 352 may also use other methods to select the machine learning model 362 from multiple different machine learning models.

[0100] After the item search module 352 has utilized the machine learning model 362 to identify one or more images in the location 302 that include the item of interest, the item search module 352 may further operate to generate one or more item search results based on the one or more images and then provide the one or more item search results to the user via the UI (e.g., the item search UI 314 or the item search UI 326).

[0101] For example, the item search module 352 can analyze one or more images including an item of interest to generate a verbal or textual description of the item of interest based on the images. Further according to this example, the item search module 352 can analyze one or more images and, based on such analysis, generate a natural language text or voice response such as "Your keys are in the living room" or "Your keys are to the left of the sofa in the living room". As another example, the item search module 350 can utilize one or more images including an item of interest to generate an image showing the location of the item of interest. For example, the item search module 350 can take an image showing the user's keys to the left of the sofa in the living room and highlight the keys in the image before presenting the image to the user. As another example, the item search module 350 can extract the portion of the image showing the keys to the left of the sofa and present an enlarged representation of that portion of the image to the user.

[0102] Other methods of generating and presenting one or more item search results can also be utilized. For example, in some embodiments, the item search module 352 can be configured to activate various automated devices within the venue 302 to indicate to the user the location of the item of interest. For example, the item search module 352 can turn on one or more smart lights near the item of interest such that IoT devices near the item of interest emit an audible sound, etc. If the item of interest itself is a device controllable by the item search service 350, the item search module 352 can cause the device itself to emit a sound, vibrate, or produce some other stimulus to help the user more easily locate the item. If the item search module 352 determines that the item of interest is currently being owned or approached by a specific person within the venue 302, the item search module 352 can send a notification to the user device (e.g., smartphone) associated with this person to let them know that the user is currently searching for the item.

[0103] Since the multiple images stored in the IoT device sensor data memory 364 can include currently captured images (e.g., images captured immediately before and / or during the item search) as well as older images, the item search module 352 can identify the location where the item of interest is currently located and the location where the item of interest was last seen even if the item of interest cannot be found in any of the currently captured pictures. For example, although the item search module 352 may not be able to locate the item of interest in any of the currently captured images, the item search module 352 may be able to find the item of interest in older images. Thus, for example, the item search module 352 may be able to return text or voice item search results that indicate "your keys were last seen at 7:32 last night when you were putting them in your wallet" or "your smartphone was last seen in your hand when you entered the garage this morning." Similarly, the item search module 352 may be able to return an item search result that includes an image of the item of interest when it was last visible to the IoT cameras in the venue 302.

[0104] In some scenarios, based on the output of the machine learning model 362, the item search module 352 can identify the same item of interest in multiple different locations in the venue 302. For example, the machine learning model 362 can identify the item of interest in images generated by different IoT cameras based on the corresponding probability scores associated with each image. In this case, the item search module 352 can present the results to the user based only on the highest probability image. Alternatively, the item search module 352 can present all the results (e.g., in probability order or simultaneously) and ask the user to confirm which result is correct. Then, the feedback from the user can be used to further train and improve the machine learning model 362. In some embodiments, the item search module 352 can utilize statistical data about where items are typically located and / or older images previously collected by one or more IoT cameras (e.g., images showing the user carrying the item to one of the candidate locations an hour before the item search was performed) to select a set of candidate locations for the item of interest.

[0105] In some cases, the item search module 362 may not be able to identify an item of interest in any of the multiple images accessed in the IoT device sensor data store 364. In such cases, the item search module 362 can generate an item search result indicating that the item could not be found and provide such an item search result to the user via the UI (e.g., the item search UI 314 or the item search UI 326). In some embodiments, if the item search module 362 is unable to locate the item of interest, the item search module 362 can prompt the user to take one or more actions to assist the item search module 352 in locating the item of interest in subsequent searches, such as: providing a different or more detailed description of the item of interest, activating additional IoT cameras within the venue 302, changing the field of view of one or more IoT cameras within the venue 302, or invoking online training features to help the machine learning model 362 better identify the type of item being searched for.

[0106] Although the item search service 350 is described herein as being configured to search for items of interest, it is noted that the item search service 350 can also be configured to search for any entity or phenomenon that can be detected via image capture, such as an entity or phenomenon, such as an action (e.g., "Did I take my medicine this morning?"), a location (e.g., "Where is the men's restroom in this office?"), an item status (e.g., "Did I leave the TV on?"), etc.

[0107] Online Training Module 354

[0108] The online training module 354 can enable a user to train the machine learning model 362 to recognize a specific item. For example, the user can provide user input via the UI (e.g., the item search UI 314 or the item search UI 326) that specifies an image of a specific item and a label assigned to the specific item. Then, the online training module 354 can use the image of the specific item and the associated label to train (e.g., train, retrain, or fine-tune) the machine learning model 362 to recognize the specific item. This feature can advantageously enable the user to selectively expand the set of items that the item search service 350 can find within the venue 302. For example, the set of items can be expanded to include uncommon or even unique items within the venue. This feature can also advantageously enable the user to assign custom names to objects (e.g., a family pet can be labeled with its name, a child's blanket can be labeled with the nickname assigned to it by the child, etc.).

[0109] According to an embodiment, different methods can be used to specify an image and associated tags. For example, a user can upload an image of an item captured using a user device (e.g., user device 104), and can also submit a voice or text input describing the item. As another example, a user can present an item within the field of view of an IoT camera (such that an image of the item can be captured by the IoT camera and made accessible to the item search service 350), while also providing a voice or text description of the item. Thus, according to this example, a user can say "These are my keys" while picking up the keys in front of an IoT camera. As yet another example, a user can provide a voice / text description of an item, and the online training module 354 can present an image captured by an IoT camera within the premises 302 to the user, and can request the user to mark or otherwise indicate the item in the image. As another example, a user can provide a voice / text description of an item, and the online training module 354 can then present multiple images captured by one or more IoT cameras within the premises 302 to the user, and can request the user to identify which images (if any) include the item. However, these are only examples, and other methods can also be used by which a user can specify an image of an item and an associated tag.

[0110] Online training features can be based on item ownership or other user-item associations for differentiating between similarly named items. For example, the online training functionality can be used to train the machine learning model 362 to distinguish between "Dad's smartphone" and "Mom's smartphone". A user can explicitly make this distinction, for example, by providing an image of Dad's smartphone along with the tag "Dad's smartphone" to the online training module 354, and an image of Mom's smartphone along with the tag "Mom's smartphone". However, this distinction can also be inferred by the online training module 354. For example, the online training module 354 can infer that a submitted smartphone image with the tag "smartphone" should actually be labeled as "Dad's smartphone" because Dad is submitting the image as part of the online training process, or because the smartphone most frequently identified historically has been used by Dad (e.g., observed by one or more IoT cameras within the premises 302).

[0111] Thus, the item tags specified by the user can be used to train (e.g., train, retrain, or fine-tune) the machine learning model 362 such that when a particular user says "Where did I put my smartphone?", the machine learning model 362 can identify the particular user's smartphone (e.g., dad's smartphone) in the IoT camera image. That is, the item search module 352 can perform an item search based on both the voice / text input from the user ("Where did I put my smartphone?") and the identity of the user who submitted the input (e.g., "dad"), such that what is really being searched for is "dad's smartphone".

[0112] In some embodiments, the online training module 354 can be configured to generate a printable QR code or other fiducial mark that a user can attach to an item of interest such that the item is more easily recognizable / distinguishable when it appears in an image. For the purposes of online training, such a fiducial mark can be attached to the item before image capture. Further according to such an example, the online training module 354 can determine one or more aspects of the fiducial mark (e.g., the size and / or density of the QR code) based on the maximum resolution or other characteristics associated with one or more IoT cameras within the venue 302.

[0113] Performing online training of the machine learning model 362 may involve processing an image of the user-specified item to put the image in a form suitable for training the machine learning model 362, processing the label of the item provided by the user to put the label in a form suitable for training the machine learning model 362, and then using such transformed inputs to perform training of the machine learning model 362. Once the online training is complete, the online training module 354 can return information indicating the success of the online training or other indication to the UI (e.g., the item search UI 314 or the item search UI 326).

[0114] In some scenarios, the machine learning model 362 can include a model designed to support multi-modal one-shot learning such that only one image and label of the item need to be submitted to train the machine learning model 362 to subsequently identify the item in the IoT camera image, or a model designed to support multi-modal few-shot learning, in which case multiple (e.g., two to five) images of the item may need to be submitted along with the associated label.

[0115] In some embodiments, an initial version of the machine learning model 362 can be provided upon first activation or installation of the item search service 350, where the initial version is trained to identify a default set of items or a basic set of items. After such activation / installation, the user can invoke the above-described features of the online training module 354 to retrain the machine learning model 362 to identify additional items beyond the default or basic set. Further in accordance with such embodiments, the user is able to disable or opt out of the online training feature. For example, for reasons related to data privacy, the user can choose to disable or opt out of the online training feature.

[0116] In some embodiments, the online training module 354 can be configured to suggest to the user items that the user may wish to add to the default set of items or the basic set of items via online training. For example, such suggestions can be made based on collaborative filtering techniques that are keyed off of the type of venue 302, the geographical location of the venue 302, and / or demographic information associated with the user.

[0117] In certain embodiments, each user associated with the venue 302 can use the online training module 354 to independently train a different version of the machine learning model 362, thereby generating multiple user-specified machine learning models. In accordance with such scenarios, when performing an item search for a particular user, the item search module 352 can utilize the machine learning model associated with the particular user to perform the item search.

[0118] Content Filtering Module

[0119] The content filtering module 356 can enable a user of the item search service 350 to specify items that should not be searched by the item search service 350. Based on the item specification, the content filtering module 356 can apply a content filter to the item search service 350 that blocks the item search service 350 from searching for the item or that blocks the item search service 350 from providing item search results regarding the item. This feature can advantageously enable a user of the item search service to selectively "hide" certain items (e.g., wall safes, handguns) that the user does not wish to be located from the item search service 350.

[0120] For example, the user can specify items that should not be searched by the item search service by interacting with one of the item search UIs 314 or 326. Further in accordance with such examples, the user can interact with the UI to provide a verbal or textual description of the item that should not be searchable and / or to provide or specify an image of the item that should not be searchable.

[0121] In response to receiving such an input, the content filtering module 356 can activate a content filter that ensures that the item search module 352 cannot be used to locate items of interest. Such a content filter can operate on the input side, meaning that if the content filter determines, based on a user's item search query, that the target being queried is an item that should not be searchable, then the content filter will completely block the execution of the item search. However, such a content filter can also operate on the output side, meaning that if the content filter determines that an item search result includes an item that should not be searchable, then the content filter can prevent the item search result from being returned to the user.

[0122] Authentication Module

[0123] The authentication module 358 can implement authentication features that ensure that only authorized users can use the item search service 350 and / or search for specific items using the item search function 350. Such features can advantageously prevent malicious actors (e.g., home intruders) from abusing the item search service 350 and can generally prevent certain users from using the item search service 350 to locate certain items within the premises 302 (e.g., prevent children in the home from accessing certain dangerous or prohibited items).

[0124] When a user attempts to initiate an item search or when a user attempts to initiate an item search for a specific item (e.g., an item on a list of sensitive items or restricted items), the authentication module 358 can be activated. Any known system or technique for authenticating users can be used to perform the authentication. For example, authentication can be performed based on a user's biometric check, where some or all of the data collected for the biometric check can be obtained by one or more IoT devices or user devices within the premises 302. Further according to such an example, facial recognition can be performed based on an image captured by a camera in the IoT / user device, voice recognition can be performed based on audio captured by a microphone in the IoT / user device, fingerprint recognition can be performed by the user device, etc. Other forms of authentication can also be used, such as but not limited to password-based authentication, multi-factor authentication, etc.

[0125] If the authentication is successful, the authentication module 358 can determine that the user is an authorized user of the item search service 350, either generally or for the purpose of searching for a specific item. Then, in such a case, the authentication module 358 can enable the user to perform an item search using the item search module 352.

[0126] In some embodiments, the authentication module 358 may be configured to assign different roles to different users within the premises 302 (e.g., parents versus children, residents versus guests, etc.), and may implement a role-based access control (RBAC) scheme to determine which users can access which features of the item search service 350 and how these users can use these features. For example, parents may be allowed to set or delete content filters for other users and also perform item searches on all items, while children may not be allowed to set or delete content filters and may be prohibited from performing item searches on certain items. As another example, residents may be able to utilize the online training feature to teach the machine learning model 362 to recognize additional items, but guests cannot. However, these are only examples, and various other roles and associated access policies may be used.

[0127] Item-based Automation / Alert Module

[0128] The item-based automation / alarm module 360 may utilize the item search features of the item search service 350 to enable users to set various automation or monitoring / alarm workflows for the premises 302.

[0129] For example, with respect to automation, a workflow may be set by the user in which the presence or absence of a particular item in a particular location may cause one or more actions to be performed. For example, the user may formulate a workflow such as: If my keys are not on the kitchen counter at 7:00 am, conduct an item search for the keys; If I am seen taking these keys to the garage, automatically open the first garage door but not the second; If I am in bed or seen leaving the house, lock all the doors and activate the home alarm system; or If mom walks into the living room, turn on the living room lights. However, these are only examples and are not intended to be limiting.

[0130] With respect to monitoring and alarms, a workflow may be set by the user in which the presence or absence of a particular item in a particular location may cause a notification to be sent to the user. For example, the user may formulate a workflow such as: If my laptop is seen outside my home office, notify me; If the hidden wall safe becomes visible, notify me; If my handgun is seen outside my handgun safe, notify me; or If a baby is seen outside the nursery, notify mom or dad. However, these are only examples and are not intended to be limiting.

[0131] Figure 4FIG. 400 is a flowchart of a method for providing an item search service for a venue including a set of IoT cameras according to some embodiments. Method 400 can be executed by processing logic that can include hardware (e.g., circuitry, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps are required for the disclosure provided herein. Additionally, those of ordinary skill in the art will understand that some steps can be performed simultaneously or in a different order than Figure 4 shown.

[0132] Method 400 will be described with reference to Figure 3 However, method 400 is not limited to this example embodiment.

[0133] In 402, the item search module 352 of the item search service 350 receives a first user input regarding an item of interest via a user interface of the item search service 350 (e.g., item search UI 314 or item search UI 326), where the first user input includes one or more of a voice input or a text input.

[0134] In 404, the item search module 352 accesses a plurality of images of the venue 302 captured by a set of IoT cameras (e.g., one or more of IoT devices 306, IoT devices 308, IoT devices 310, or IoT devices 312) present in the venue 302. For example, the item search module 352 can access the plurality of images by accessing the IoT device sensor data memory 364 in which these images can be stored.

[0135] In 406, the item search module 352 executes a machine learning model 362 to identify one or more images among the plurality of images that include the item of interest, at least based on the first user input. In certain implementations, the first user input includes a natural language input, and the machine learning model 362 includes a multimodal machine learning model trained on a set of images and natural language text respectively associated with each image in the set of images. In additional implementations, the machine learning model 362 is selected from a plurality of different machine learning models, where each machine learning model among the plurality of different machine learning models is trained or fine-tuned for one of a specific venue type or a specific user demographic.

[0136] In 408, the item search module 352 generates an item search result based on the identified one or more images. In certain implementations, generating the item search result includes generating a voice or text description of the location of the item of interest based on the identified one or more images, or generating one or more of an image showing the location of the item of interest based on the identified one or more images.

[0137] In 410, the item search module 352 provides item search results to the user via the UI of the item search service 350 (e.g., via the item search UI 314 or the item search UI 326).

[0138] Figure 5 A flowchart of a method 500 for training a machine learning model for implementing an item search service for a venue is shown. The method 500 can be executed by processing logic, which can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps need to be performed according to the disclosure provided herein. In addition, those of ordinary skill in the art will understand that some steps can be performed simultaneously or in a different order than Figure 5 shown.

[0139] Reference will be made to Figure 3 describe the method 500. However, the method 500 is not limited to this example embodiment.

[0140] In 502, the online training module 354 of the item search service 350 receives a second user input via the user interface of the item search service 350 (e.g., the item search UI 314 or the item search UI 326), the second user input specifying an image of an item of interest and a label assigned by the user to the item of interest.

[0141] In 504, the online training module 354 uses the image of the item of interest and the label assigned to the item of interest to train (e.g., train, retrain, or fine-tune) the machine learning model 362.

[0142] Figure 6 A flowchart of a method 600 for providing a secure item search service for a venue is shown. The method 600 can be executed by processing logic, which can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps need to be performed according to the disclosure provided herein. In addition, those of ordinary skill in the art will understand that some steps can be performed simultaneously or in a different order than Figure 6 shown.

[0143] Reference will be made to Figure 3 describe the method 600. However, the method 600 is not limited to this example embodiment.

[0144] In 602, the authentication module 358 of the item search service 350 authenticates the user of the item search service 602.

[0145] In 604, the authentication module 358 determines that the user is an authorized user based on the authentication.

[0146] In 606, in response to determining in 604 that the user is an authorized user, the item search module 352 of the item search service 350 performs one or more of receiving 402, accessing 404, executing 406, generating 408, or providing 410 of method 400.

[0147] Figure 7 A flowchart of a method 700 for preventing an item search service at a venue from searching for a specified item is shown. Method 700 can be executed by processing logic that can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps need to be performed for the disclosure provided herein. Additionally, those of ordinary skill in the art will understand that some steps can be performed simultaneously or in a different order than Figure 7 shown.

[0148] will be referred to Figure 3 to describe method 700. However, method 700 is not limited to this example embodiment.

[0149] In 702, the content filtering module 356 of the item search service 350 receives a second user input via a user interface of the item search service 350 (e.g., item search UI 314 or item search UI 326) specifying an item that should not be searchable.

[0150] In 704, the content filtering module 356 applies a content filter that prevents the item search service 350 from searching for an item that should not be searchable or prevents the item search service 350 from returning item search results for an item that should not be searchable.

[0151] Figure 8 A flowchart of a method 800 for providing an item search service for a venue that supports user-specific query interpretation is shown. Method 800 can be executed by processing logic that can include hardware (e.g., circuits, dedicated logic, programmable logic, microcode, etc.), software (e.g., instructions executed on a processing device), or a combination thereof. It should be understood that not all steps need to be performed for the disclosure provided herein. Additionally, those of ordinary skill in the art will understand that some steps can be performed simultaneously or in a different order than Figure 8 shown.

[0152] will be referred to Figure 3 to describe method 800. However, method 800 is not limited to this example embodiment.

[0153] In 802, the item search module 352 of the item search service 350 determines the identity of the user of the item search service 350.

[0154] In 804, the item search module 352 executes a machine learning model 362 to identify one or more images among a plurality of images that include an item of interest, based at least on a first user input and the user identity determined in 802. For example, before providing the first user input to the machine learning model 362, the item search module 352 may modify the first user input (e.g., "Where is my smartphone") to incorporate the user's identifier (e.g., "Where is Dad's smartphone").

[0155] Example Computer System

[0156] For example, one or more well-known computer systems (e.g., Figure 9 the computer system 900 shown in ) can be used to implement various embodiments. For example, a combination or sub-combination of the computer system 900 can be used to implement one or more of the following: user device 304, IoT device 306, IoT device 308, IoT device 310, IoT device 312, item search service 350, item search module 352, online training module 354, content filtering module 356, authentication module 358, item-based automation / alarm module 360, machine learning model 362, or IoT device sensor data memory 364. Additionally or alternatively, for example, one or more computer systems 900 can be used to implement any of the embodiments discussed herein, as well as their combinations and sub-combinations.

[0157] The computer system 900 can include one or more processors (also referred to as central processing units or CPUs), such as processor 904. The processor 904 can be connected to a communication infrastructure 906 or a bus.

[0158] The computer system 900 can also include one or more user input / output devices 903, such as monitors, keyboards, pointing devices, etc., which can communicate with the communication infrastructure 906 through one or more user input / output interfaces 902.

[0159] One or more processors 904 can be a graphics processing unit (GPU). In an embodiment, the GPU can be a processor of a dedicated electronic circuit designed to process math-intensive applications. The GPU can have a parallel structure that is effective for parallel processing of large blocks of data (e.g., math-intensive data common for computer graphics applications, images, videos, etc.).

[0160] The computer system 900 may also include a main memory or primary storage 908, e.g., random access memory (RAM). The main memory 908 may include one or more levels of cache. The main memory 908 may store control logic (i.e., computer software) and / or data therein.

[0161] The computer system 900 may also include one or more secondary storage devices or memories 910. For example, the secondary memory 910 may include a hard disk drive 912 and / or a removable storage device or drive 914. The removable storage drive 914 may be a floppy disk drive, a tape drive, an optical disk drive, an optical storage device, a tape backup device, and / or any other storage device / drive.

[0162] The removable storage drive 914 may interact with a removable storage unit 918. The removable storage unit 918 may include a computer-usable or readable storage device on which computer software (control logic) and / or data is stored. The removable storage unit 918 may be a floppy disk, a tape, an optical disk, a DVD, a compact disc, and / or any other computer data storage device. The removable storage drive 914 may read from and / or write to the removable storage unit 918.

[0163] The secondary memory 910 may include other means, devices, components, tools, or other methods for allowing computer programs and / or other instructions and / or data to be accessed by the computer system 900. Such means, devices, components, tools, or other methods may include, for example, a removable storage unit 922 and an interface 920. Examples of the removable storage unit 922 and the interface 920 may include a program cartridge and a cartridge interface (e.g., those present in video game devices), a removable storage chip (e.g., EPROM or PROM) and an associated socket, a memory stick and a USB or other port, a memory card and an associated memory card socket, and / or any other removable storage unit and associated interface.

[0164] The computer system 900 may also include a communication or network interface 924. The communication interface 924 may enable the computer system 900 to communicate and interact with any combination of external devices, external networks, external entities, etc. (collectively and individually denoted by the reference numeral 928). For example, the communication interface 924 may allow the computer system 900 to communicate with an external or remote device 928 via a communication path 926, which may be wired and / or wireless (or a combination thereof) and may include any combination of LAN, WAN, the Internet, etc. Control logic and / or data may be sent to and from the computer system 900 via the communication path 926.

[0165] The computer system 900 can also be any one of a personal digital assistant (PDA), a desktop workstation, a laptop or notebook computer, a netbook, a tablet computer, a smart phone, a smart watch or other wearable device, a household appliance, a part of the IoT, and / or an embedded system (to name just a few non-limiting examples), or any combination thereof.

[0166] The computer system 900 can be a client or a server, accessing or hosting any application and / or data through any delivery paradigm, including but not limited to remote or distributed cloud computing solutions; local or in-house software (“in-house” cloud-based solutions); “as-a-service” models (e.g., Content as a Service (CaaS), Digital Content as a Service (DCaaS), Software as a Service (SaaS), Managed Software as a Service (MSaaS), Platform as a Service (PaaS), Desktop as a Service (DaaS), Framework as a Service (FaaS), Backend as a Service (BaaS), Mobile Backend as a Service (MBaaS), and Infrastructure as a Service (IaaS), etc.); and / or a hybrid model including any combination of the examples described above or other services or delivery paradigms.

[0167] Any applicable data structures, file formats, and schemas in the computer system 900 can be derived from standards, including but not limited to JavaScript Object Notation (JSON), Extensible Markup Language (XML), another markup language (YAML), Extensible HyperText Markup Language (XHTML), Wireless Markup Language (WML), message packets, XML User Interface Language (XUL), or any other functionally similar representation, either alone or in combination. Alternatively, proprietary data structures, formats, or schemas can be used alone or in combination with known or open standards.

[0168] In some embodiments, a tangible, non-transitory apparatus or article of manufacture including a tangible, non-transitory computer-usable or readable medium having control logic (software) stored thereon may also be referred to herein as a computer program product or a program storage device. This includes but is not limited to the computer system 900, the main memory 908, the secondary memory 910, and the removable storage units 918 and 922, as well as tangible articles embodying any combination of the foregoing. When executed by one or more data processing devices (e.g., the computer system 900 or one or more processors 904), such control logic can cause these data processing devices to operate as described herein.

[0169] Based on the teachings contained in this disclosure, for those skilled in the relevant art, how to use in addition to Figure 9It will be apparent that other data processing devices, computer systems, and / or computer architectures can be used to make and use embodiments of the present disclosure. In particular, embodiments can be implemented and operated using software, hardware, and / or operating systems other than those described herein.

[0170] Conclusion

[0171] It should be understood that the detailed description section, and not any other section, is intended to interpret the claims. Other sections can set forth one or more but not all of the exemplary embodiments contemplated by the inventors, and thus are not intended to limit the present disclosure or the appended claims in any way.

[0172] Although the present disclosure describes exemplary embodiments in exemplary fields and applications, it should be understood that the present disclosure is not limited thereto. Other embodiments and modifications thereof are possible and within the scope and spirit of the present disclosure. For example, without limiting the generality of this paragraph, embodiments are not limited to the software, hardware, firmware, and / or entities shown in the figures and / or described herein. Additionally, embodiments (whether explicitly described herein or not) have significant utility in fields and applications other than those exemplified herein.

[0173] Embodiments are described herein by way of functional building blocks that illustrate implementations of specified functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries can be defined as long as the specified functions and relationships (or their equivalents) are appropriately performed. Additionally, alternative embodiments can perform functional blocks, steps, operations, methods, etc. in a different order than that described herein.

[0174] References herein to "one embodiment", "an embodiment", "exemplary embodiment", or similar phrases indicate that the described embodiment may include a particular feature, structure, or characteristic, but not every embodiment must include that particular feature, structure, or characteristic. Moreover, these phrases do not necessarily refer to the same embodiment. Additionally, when a particular feature, structure, or characteristic is described in connection with an embodiment, incorporating such feature, structure, or characteristic into other embodiments (whether explicitly mentioned or described herein or not) is within the knowledge of those skilled in the relevant art. Further, some embodiments can be described using the terms "coupled" and "connected" and their derivatives. These terms are not necessarily synonyms of each other. For example, some embodiments can be described using the terms "coupled" and / or "connected" to indicate that two or more elements are in direct physical or electrical contact with each other. However, the term "coupled" can also mean that two or more elements are not in direct contact with each other but still cooperate or interact with each other.

[0175] The breadth and scope of the present disclosure should not be limited by any of the above exemplary embodiments, but should be defined only in accordance with the appended claims and their equivalents.

Claims

1. A computer-implemented method for providing an item search service for a venue including a set of Internet of Things (IoT) cameras, comprising: receiving, by at least one computer processor and via a user interface of the item search service, a first user input regarding an item of interest, wherein the first user input comprises one or more of a voice input or a text input; accessing a plurality of images of the location captured by the set of IoT cameras; executing a machine learning model to identify one or more images of the plurality of images that include the item of interest based at least on the first user input; generating item search results based on the one or more images of the identification; and The item search results are provided via a user interface of the item search service.

2. The computer-implemented method of claim 1 , wherein: The first user input comprises a natural language input, and wherein the machine learning model comprises a multimodal machine learning model trained on a set of images and natural language text respectively associated with each image in the set of images.

3. The computer-implemented method of claim 1 , further comprising: receiving, via a user interface of the item search service, a second user input specifying an image of the item of interest and a tag assigned by a user to the item of interest; as well as The machine learning model is trained using the images of the items of interest and the labels assigned to the items of interest.

4. The computer-implemented method of claim 1 , further comprising: The receiving, accessing, executing, generating, and providing steps are performed on one or more devices located within the premises.

5. The computer-implemented method of claim 1 , further comprising: The machine learning model is selected from a plurality of different machine learning models, wherein each of the plurality of different machine learning models is trained or fine-tuned for one of a specific venue type or a specific user demographic.

6. The computer-implemented method of claim 1 , further comprising: authenticating a user of the item search service; as well as determining, based on the authentication, that the user is an authorized user of the item search service; Wherein, one or more of the receiving, accessing, executing, generating and providing are performed in response to determining that the user is an authorized user of the item search service.

7. The computer-implemented method of claim 1 , further comprising: receiving, via a user interface of the item search service, a second user input specifying an item that should not be searchable; as well as In response to receiving the second user input, a content filter is applied that prevents the item search service from searching for items that should not be searchable or prevents the item search service from returning item search results for items that should not be searchable.

8. The computer-implemented method of claim 1 , further comprising: determining the identity of a user of the item search service; Wherein executing the machine learning model includes executing the machine learning model to identify one or more images of the plurality of images that include the item of interest based at least on the first user input and the identity of the user.

9. The computer-implemented method of claim 1 , wherein: Generating the item search results based on the one or more images of the identification includes one or more of the following operations: generating a spoken or textual description of the location of the item of interest based on the identified one or more images; or An image showing the location of the item of interest is generated based on the one or more images of the identification.

10. A system for providing an item search service for a location including a set of Internet of Things (IoT) cameras, comprising: one or more memories; as well as at least one processor, each processor being coupled to at least one of the memories and configured to perform operations including: receiving, via a user interface of the item search service, a first user input regarding an item of interest, wherein the first user input comprises one or more of a voice input or a text input; accessing a plurality of images of the location captured by the set of IoT cameras; executing a machine learning model to identify one or more images of the plurality of images that include the item of interest based at least on the first user input; generating item search results based on the one or more images of the identification; and The item search results are provided via a user interface of the item search service.

11. The system according to claim 10, wherein: The first user input comprises a natural language input, and wherein the machine learning model comprises a multimodal machine learning model trained on a set of images and natural language text respectively associated with each image in the set of images.

12. The system according to claim 10, wherein: The operations also include: receiving, via a user interface of the item search service, a second user input specifying an image of the item of interest and a tag assigned by a user to the item of interest; and The machine learning model is trained using the images of the items of interest and the labels assigned to the items of interest.

13. The system according to claim 10, wherein: The operations also include: The machine learning model is selected from a plurality of different machine learning models, wherein each of the plurality of different machine learning models is trained or fine-tuned for one of a specific venue type or a specific user demographic.

14. The system according to claim 10, wherein: The operations also include: authenticating a user of the item search service; and determining, based on the authentication, that the user is an authorized user of the item search service; Wherein, one or more of the receiving, accessing, executing, generating and providing are performed in response to determining that the user is an authorized user of the item search service.

15. The system according to claim 10, wherein: The operations also include: receiving, via a user interface of the item search service, a second user input specifying an item that should not be searchable; and In response to receiving the second user input, a content filter is applied that prevents the item search service from searching for items that should not be searchable or prevents the item search service from returning item search results for items that should not be searchable.

16. The system of claim 10, wherein: The operations also include: determining the identity of a user of the item search service; Wherein executing the machine learning model includes executing the machine learning model to identify one or more images of the plurality of images that include the item of interest based at least on the first user input and the identity of the user.

17. The system according to claim 10, wherein: Generating the item search results based on the one or more images of the identification includes one or more of the following operations: generating a spoken or textual description of the location of the item of interest based on the identified one or more images; or An image showing the location of the item of interest is generated based on the one or more images of the identification.

18. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by at least one computing device, cause the at least one computing device to perform operations for providing an item search service for a venue including a set of Internet of Things (IoT) cameras, the operations comprising: receiving, via a user interface of the item search service, a first user input regarding an item of interest, wherein the first user input comprises one or more of a voice input or a text input; accessing a plurality of images of the location captured by the set of IoT cameras; executing a machine learning model to identify one or more images of the plurality of images that include the item of interest based at least on the first user input; generating item search results based on the one or more images of the identification; and The item search results are provided via a user interface of the item search service.

19. The non-transitory computer readable medium of claim 18, wherein: The first user input comprises a natural language input, and wherein the machine learning model comprises a multimodal machine learning model trained on a set of images and natural language text respectively associated with each image in the set of images.

20. The non-transitory computer readable medium of claim 18, wherein: The operations also include: receiving, via a user interface of the item search service, a second user input specifying an image of the item of interest and a tag assigned by a user to the item of interest; and The machine learning model is trained using the images of the items of interest and the labels assigned to the items of interest.