Control of smart environments based on semantic control models
Semantic control models using generative AI and machine learning enable efficient and transparent control of smart devices by adapting settings based on user inputs and environmental conditions, addressing the inefficiencies in existing smart home and communication systems.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- KONINKLIJKE PHILIPS NV
- Filing Date
- 2024-07-10
- Publication Date
- 2026-07-29
AI Technical Summary
Existing smart home and communication systems face inefficiencies in controlling smart devices due to the need for specifying detailed settings and transferring large amounts of content, and coordinating multiple devices to achieve similar effects is complex, especially when environments and device characteristics vary.
Implementing semantic control models using generative AI and machine learning models, such as diffusion models and generative pre-trained transformers, to generate and transfer control prompts, allowing smart devices to adapt settings based on user inputs and environmental conditions.
Enables efficient and transparent control of smart devices across varying environments by deriving and reproducing desired settings through simple user inputs, reducing the complexity of setting up and maintaining consistent smart device experiences.
Smart Images

Figure 2026525269000001_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a system for controlling smart devices in a smart environment, such as a smart home system or a smart communication system. This invention can also be used to control smart devices in other environments, such as wireless communication devices.
Background Art
[0002] Smart home systems have smart devices (such as sensors and actuators for audio, video, lighting, etc.) and aim to provide a comfortable home experience. Standards such as "Matter" aim to achieve interoperability between devices by providing a unified Internet Protocol (IP) communication layer and a unified data model for controlling smart devices. Matter is an open-source connection standard for smart home and IoT devices, promising that smart devices can cooperate with each other regardless of the manufacturer, that smart home devices can continue to be used on the home network without an Internet connection, and that communication with smart home devices is secure.
[0003] However, currently, smart devices (such as smart lights, smart speakers, televisions (TVs), smart blinds, etc.) are controlled by specifying specific content to be played or specific settings to be applied. Furthermore, users need to set the desired output of smart devices (such as "natural light", "diamond", "rest", "brightness", etc. scenes for a smart light, or specific pixels within a scene that determine the color of the light) and provide it to the smart device. Additionally, the same lighting color and intensity may not produce the same effect in different environments. For example, when the wall color is different, the user's preferences are slightly different, or the devices (number, type, distribution) are different.
[0004] Furthermore, smart communication systems enable people to communicate with each other. For example, Facebook Portal TV allows people to communicate with each other comfortably from the sofa in their living room. Facebook Portal TV can be linked with a WhatsApp account, and the first Facebook Portal TV can record the environment of the first home, transfer it to a second Facebook Portal TV, and display it on a second TV in the second home.
[0005] However, such a system cannot transfer specific smart settings (for example, the first set of smart settings associated with the first set of smart devices in the first home) so that the second set of smart devices in the second smart home can use the same or similar smart settings.
[0006] Therefore, controlling smart devices (e.g., steering, creating settings for smart devices, users, and environments, encoding data affected by smart devices) can be inefficient or complex due to the need to specify and transfer large amounts of content. Furthermore, coordinating multiple smart devices to generate similar content is also a challenge. Ideally, the lighting system should provide the desired setting simply by the user expressing the desired type, such as the desired lighting.
[0007] Generally, a comprehensive smart home system should be able to control not only lighting fixtures but also other smart devices in a comprehensive manner, and this should be done in a transparent manner for the user, regardless of the characteristics of the environment or the devices present in it. For example, if a user says they want a "comfortable environment," the smart home system should create and execute settings in which all smart devices work together to create the desired comfortable environment. [Overview of the project] [Problems that the invention aims to solve]
[0008] The objective of this invention is to provide more efficient control of smart devices in a smart environment. [Means for solving the problem]
[0009] This objective is achieved by the apparatus of claim 1, the hub or smart device of claim 10, the smart device of claim 11, the system of claim 13, the methods of claims 14 and 15, and the computer program product of claim 16.
[0010] According to the first aspect (applicable to controllers in hubs or smart devices), a device is provided for controlling at least one smart device, in particular sensors and / or actuators, in a smart environment, the device being configured to generate, adapt, acquire or receive at least one semantic control model associated with at least one smart device and at least one semantic prompt associated with a scene, and to use at least one semantic control model and semantic prompt to control at least one smart device.
[0011] According to the second aspect, a hub or smart device having the apparatus according to the first aspect is provided.
[0012] According to a third aspect, a smart device is provided, which is configured to receive or obtain device settings, including a semantic control model and control prompts for controlling the smart device.
[0013] According to the fourth aspect, a system is provided having a hub according to the second aspect and one or more smart devices according to the second or third aspect.
[0014] According to a fifth aspect (applying to hubs or smart devices), a method is provided for controlling smart devices, in particular sensors and / or actuators, in a smart environment, and said method A step of generating, adapting, acquiring, or receiving a semantic control model in a hub or smart device of a smart environment, wherein the semantic control model is associated with or applied to the smart device. The process includes the step of using a semantic control model to control a smart device.
[0015] According to the sixth aspect (applying to smart devices), a method is provided for controlling smart devices, particularly sensors and / or actuators, in a smart environment, and said method A smart device receives device settings including a semantic control model and / or control prompts, The process includes the step of controlling a smart device based on device settings.
[0016] According to the seventh aspect, a computer product is provided, which includes coding means for generating the steps of the fifth or sixth aspect when executed on a computer device.
[0017] Therefore, (generative) artificial intelligence (AI) or machine learning (ML) models, semantic control models, such as diffusion models and generative pre-trained transformer (GPT) models (e.g., DALL-E), can be used to efficiently control smart devices ("prompt to scene"), and smart devices can be configured to store semantic control models. This allows smart devices to be individually configured to receive control (semantic) prompts and apply the received prompts to the received semantic control model to derive and reproduce smart device settings. Thus, the output or function of a smart device can be controlled by selected prompts defined by the device settings.
[0018] Semantic extraction models (such as CLIP and im2prompt) can also be used to efficiently extract semantic meaning from images, for example ("scene to prompt"). Such models can be used to extract semantic meaning from smart home scenes, such as scenes generated by smart devices (e.g., smart lighting, smart TVs, music). These can be used to extract semantic meaning from the environment (sunny / rainy day, hot / cold day, daytime, etc.). Extraction can be performed by taking one or more photos using a smart device (including functions such as a camera and microphone) and extracting the semantic meaning using semantic model extraction. This allows for the observation of a scene (for example, a smart home environment in a first smart home), the generation of one or more prompts describing that scene (using techniques such as those described in Yoad Tewel et al.: “ZeroCap: Zero-Shot Image-to-Text Generation for Visual-Semantic Arithmetic” (arXiv:2111.14447v2 [cs.CV] 31 Mar 2022) or Alec Radford et al.: “Learning Transferable Visual Models From Natural Language Supervision” (arXiv:2103.00020)), the transfer of those prompts to a second location, and the application of them to a semantic control model to generate similar scenes.
[0019] According to a first option which can be combined with any of the first to seventh aspects, the device can be configured to distribute at least one of a semantic control model and a semantic prompt to at least one smart device. This allows the smart device to be controlled through a prompt that is translatable in the semantic control model.
[0020] According to the first option or a second option which can be combined with any of the first to seventh aspects, the device can be configured to receive a command, derive or obtain a control prompt from that command, and apply that control prompt to a semantic control model to obtain device settings. Thus, a smart device can be configured based on a semantic control model via simple control prompts.
[0021] According to the first or second option, or a third option which can be combined with any of the first through seventh aspects, this device configuration can be used to control at least one smart device. Thus, the smart device can be controlled based on a semantic control model through simple control prompts.
[0022] According to the fourth option, which can be combined with any of the first to third options or any of the first to seventh aspects, the semantic control model and / or control prompts can be improved based on user feedback. Thus, the effects and results of the semantic control model can be evaluated and adjusted (improved) by the user.
[0023] According to one of the first to fourth options, or a fifth option which can be combined with any of the first to seventh aspects, semantic control models and control prompt indicators can be received from remote devices. This allows for the optimization of proposed model-based control and configuration using external information from other smart environments, external networks, model repositories, etc.
[0024] According to the 6th option that can be combined with any of the 1st to 5th options or any of the 1st to 7th aspects, it is possible to contact cloud-based services or blockchain and obtain or verify a semantic control model for smart devices. Thereby, by using cloud-based resources, the software and hardware requirements in the hub can be minimized.
[0025] According to the 7th option that can be combined with any of the 1st to 6th options or any of the 1st to 7th aspects, the semantic control model can be customized according to the smart environment of the smart device by considering one or more of the functions of available smart devices, user preferences, and the positions of smart devices. Therefore, the proposed model-based control and settings can be adapted to the individual smart environments of the smart devices to be controlled.
[0026] According to the 8th option that can be combined with any of the 1st to 7th options or any of the 1st to 7th aspects, smart devices in the second environment can be controlled based on device settings including the semantic control model and control prompts of the first environment. Thereby, it becomes possible to transfer model-based control and settings to other smart environments.
[0027] According to the 9th option that can be combined with any of the 1st to 8th options or any of the 1st to 7th aspects, smart devices can be configured to infer device settings by identifying other smart devices in the smart environment and specifying their settings. Thereby, model-based device settings can be directly transferred between smart devices without going through the hub.
[0028] It should be understood that the apparatus of claim 1, the hub or smart device of claim 10, the smart device of claim 11, the system of claim 13, the methods of claims 14 and 15, and the products of the computer program of claim 16 may have similar or identical embodiments, in particular, as defined in the dependent claims.
[0029] It should be understood that preferred embodiments of the present invention may be dependent claims or any combination of the above embodiments and their respective independent claims.
[0030] These and other aspects of the present invention will become apparent from and be explained with reference to the embodiments described below. [Brief explanation of the drawing]
[0031] [Figure 1] A diagram illustrating the system architecture of a feasible smart environment. [Figure 2] A schematic diagram illustrating the signaling and processes of an embodiment for managing and acquiring semantic control models for smart device registration and smart device control. [Figure 3] A diagram schematically illustrating the signaling and processes of an embodiment for operating a smart device in a first environment based on prompts and a semantic control model for smart device control. [Figure 4] A diagram illustrating the signaling and process of an embodiment for operating the smart device shown in Figure 3 in a second environment based on the smart device settings in a first environment. [Figure 5] A schematic diagram illustrating a smart device according to an embodiment of the present invention. [Modes for carrying out the invention]
[0032] Here, embodiments of the present invention are described as a smart home system. However, the present invention can also be used in conjunction with other smart environments and network systems such as cellular networks and Wi-Fi networks.
[0033] Through the following disclosures, “smart home” or “network” is understood as a network containing sensors and actuators that facilitate specific tasks (e.g., lighting or medical-related). This may include a home network hub (such as a data distribution entity) that manages the home network and allows multiple devices or nodes (such as sensors and actuators) to connect to the network. The home network hub may also be an entity that directs secure data distribution, such as data originating from the network. The home network hub may include or provide access to a router device for linking the home network to an external network (e.g., the internet), and / or enable the addition of devices to or removal of devices from the network. The network hub may be a smart device that manages the smart home, such as a smartphone or a smart home assistant device like Alexa or Google Home. These entities may be centralized or decentralized.
[0034] Furthermore, the term "metaverse" is understood to refer to a shared collection of potentially persistent, interactive spaces in which users can interact with each other using mutually perceived virtual features (i.e., augmented reality (AR)), or these spaces may be composed entirely of virtual features (i.e., virtual reality (VR)). VR and AR are sometimes broadly referred to as "mixed reality" (MR).
[0035] Furthermore, the term “data” is understood to refer to a representation of information in a known or agreed format that is stored, transmitted, or otherwise processed. The information may, in particular, include one or more channels of audio, video, images, haptics, motion, or other forms of multimedia information that can be synchronized. Such multimedia information may be derived from sensors (e.g., microphones, cameras, motion detectors, etc.) or partially or completely synthesized (e.g., a live actor in front of a synthesized background).
[0036] The term “data object” refers to one or more sets of data based on the definition above, optionally accompanied by one or more data descriptors that provide additional semantic information about the data that influences how it is processed at the sender and receiver. Data descriptors can be used, for example, to describe how the sender classifies the data and how the receiver should render it. For example, data representing an image or video sequence can be broken down into a set of data objects that describe the entire image or video together, and can be processed (e.g., compressed) individually, substantially independently of other data objects, in a manner best suited to the object and its semantic context. As a further example, a content program (as described in some embodiments below) can also be understood as a data object (e.g., a compressed semantic object).
[0037] Furthermore, the term “data object classification” is understood to refer to the process of dividing or segmenting data into multiple data objects. For example, an image or scene may be divided into multiple parts, such as a background forest and a foreground person. Criteria for data object classification are used to classify data objects. In this disclosure, such criteria may include at least one of the following: an indicator of the semantic content of the data object, the context of the data object, or a class of compression techniques that are best suited to preserving sufficient semantic content in a particular context.
[0038] Furthermore, a “semantic control model” is understood to refer to a repository of tools and data objects that can be used to assist in the control of smart devices and / or smart environments. For example, a model may include algorithms used for parsing and compressing data objects, or data objects that can be used as the basis for generative compression techniques. Advantageously, a model can be shared or owned by the sender and receiver, and / or updated or optimized according to the semantic content of the data being transferred. Also, a model can be adapted according to the device using it. A model can be personalized according to user preferences and the environment in which the user is located.
[0039] Furthermore, "prompt" is understood to refer to any type of gesture, physical action (such as guidance or touch), visual aid (such as diagrams or photographs), auditory aid (such as voices, music, songs, or spoken language), or other detectable stimulus.
[0040] Generative artificial intelligence (AI) can be understood as a set of algorithms (e.g., large-scale language models and diffusion models) that can generate seemingly new and realistic content such as text, images, and audio from training data. Powerful generative AI algorithms are self-supervisingly trained on vast amounts of unlabeled data and built upon foundational models that identify patterns underlying diverse tasks. New generative AI models can not only engage in sophisticated conversations with users, but they can also generate seemingly original content. Representative examples include "ChatGPT," which includes a generative text-based AI model; "DALL-E" or "MidJourney," which includes a generative "text-to-image" AI model; and "Gato," which includes a generative "text-to-video" AI model.
[0041] Generative AI enables the generation of multiple types of data based on short prompts. Stable Diffusion is one example. Stable Diffusion is an open-source machine learning model that was first made publicly available on August 22, 2022, and can generate images from text, modify images based on text, and add detail to low-resolution or low-detail images. It has been trained on billions of images and can produce results comparable to DALL-E 2 and MidJourney.
[0042] Generative AI is a branch of artificial intelligence that aims to generate new data from existing data such as text, images, and audio. Generative AI can be used in a variety of applications, including content creation, data augmentation, style transfer, and image completion. Generative AI models are typically based on two types of algorithms: large-scale language models and spreading models.
[0043] Large-scale language models (LLMs) are neural networks that can generate natural language from given inputs such as words, phrases, and images. LLMs are trained on vast amounts of text data, including books, articles, blogs, and social media posts, to learn statistical patterns and relationships between words and sentences. Using these patterns, LLMs can generate consistent, fluent text that matches the input and desired output. Examples of LLMs include GPT-3, BERT, and T5.
[0044] GPT-3 is one of the largest and most powerful LLMs, boasting 175 billion parameters and a vocabulary of 50,000 words. Given a few words or sentences as prompts, GPT-3 can generate text on virtually any topic. GPT-3 utilizes a transformer architecture, a type of neural network that processes continuous data such as text and speech using an attention mechanism. This attention mechanism allows the network to focus on the most important parts of the input and output, capturing long-range dependencies and context. GPT-3 is trained on a large and diverse text corpus called "Common Crawl," which covers various domains and languages.
[0045] BERT is another LLM that uses a transformer architecture, but its purpose and training method are different. BERT is designed to learn the bidirectional context of words in a sentence, meaning it can understand the meaning of a word from the words that precede and follow it. BERT is trained on two tasks: masked language modeling and next sentence prediction. Masked language modeling involves randomly masking some words in a sentence and asking the network to predict those words from the rest of the sentence. Next sentence prediction involves asking the network to determine whether two sentences are consecutive in a text. BERT is trained on a large text corpus called "BookCorpus," which consists of 11,038 books from various genres.
[0046] T5 is another LLM that uses a transformer architecture, but has a simpler, more general purpose and training method. T5 is designed to learn mappings between arbitrary input and output texts and can perform a variety of natural language tasks such as translation, summarization, question answering, and text classification. T5 is trained on a single task called text-to-text generation. Text-to-text generation asks the network to generate output text from input text with a prefix that specifies the task. For example, the prefix "translate English to French:" tells the network that the input text should be translated from English to French. T5 is trained on a large and diverse text corpus called the "Colossal Clean Crawled Corpus," which covers various domains and languages.
[0047] Diffusion models are neural networks that can generate realistic images from given inputs such as text descriptions, sketches, and low-resolution images. Diffusion models are trained on large amounts of images, such as faces, animals, landscapes, and artworks, learning the distribution and structure of pixels in different regions. These distributions can then be used to generate high-quality, diverse images that match the input and desired output. Examples of diffusion models include StyleGAN, CLIP, and VQGAN.
[0048] StyleGAN is one of the most advanced and realistic diffusion models, capable of generating high-resolution, diverse images of faces, animals, cars, and landscapes. StyleGAN employs a generative adversarial adversarial network (GAN) architecture, which consists of two competing neural networks: a generator and a discriminator. The generator attempts to produce fake images that look like real images, while the discriminator attempts to distinguish between real and fake images. The generator and discriminator are iteratively trained, with the generator learning to deceive the discriminator and the discriminator learning to catch the generator. Through this process, both networks improve their performance and generate realistic and novel images.
[0049] StyleGAN has several distinctive features that differentiate it from other GANs. One of these is its style-based generator, which allows the network to control the style and variations of the images it generates at different levels of detail. The style-based generator uses a mapping network that maps input noise vectors to a latent space representing the image style. This latent space is then input to a synthesis network, and an image is generated from the style. The style-based generator can also mix different styles from different latent vectors, improving the diversity and quality of the generated images. Another feature is progressive growing, which allows the network to start with low-resolution images and gradually increase the resolution as training progresses. This stabilizes the network during training and avoids mode collapse, which can be a problem when generating similar or identical images.
[0050] CLIP is another diffusion model that can generate realistic images from given text descriptions. CLIP uses a contrastive learning method, learning from correct and incorrect data pairs. CLIP is trained on a large and diverse image and caption dataset called the "Conceptual Captions" dataset, which contains 3.3 million image and caption pairs collected from the web. CLIP learns to associate semantically related images and captions and distinguish them from unrelated images and captions. CLIP uses a transformer architecture for both the image and text encoders, allowing the network to capture global and local features of images and text. CLIP can generate images that match text descriptions by using a diffusion model as a decoder.
[0051] VQGAN is another diffusion model that can generate realistic images from given text descriptions or sketches. VQGAN uses vector quantization, a technique that compresses high-dimensional image data into low-dimensional discrete representations. VQGAN has been trained on large and diverse image datasets, such as ImageNet, which contains 14 million images from 20,000 categories. VQGAN learns to encode images into a discrete latent space consisting of a fixed number of codebook vectors. Each codebook vector represents a visual feature or concept, such as color or shape object. VQGAN can use a transformer architecture as a decoder to generate images that match text descriptions or sketches.
[0052] Both LLM and diffusion models are examples of generative adversarial networks (GANs), consisting of two competing neural networks: a generator and a discriminator. The generator attempts to generate fake data that looks like real data, while the discriminator attempts to distinguish between real and fake data. The generator and discriminator are trained iteratively, with the generator learning to deceive the discriminator and the discriminator learning to catch the generator. Through this process, both networks improve their performance and generate real, novel data.
[0053] Please note that only the blocks, components, and / or devices related to the proposed data distribution functionality are shown in the accompanying drawings. Other blocks have been omitted for brevity. Furthermore, blocks designated with the same reference number are intended to have the same or at least similar functionality, and therefore their functionality will not be described again later.
[0054] Figure 1 schematically shows the system architecture of a smart environment in which various embodiments described below can be implemented.
[0055] The exemplary system architecture in Figure 1 includes a first smart environment 101 and a second smart environment 102, both of which can be smart homes. The first user 105 is located in the first smart environment 101, and the second user 106 is located in the second smart environment 102.
[0056] Furthermore, the system architecture in Figure 1 includes two smart hubs 111 and 112, each responsible for one or both of the respective smart environments 101 and 102. The smart hubs 111 and 112 can be operated centrally or distributed, locally or in combination with cloud services.
[0057] The smart environments 101,102 include smart home devices 121-126 such as smart lighting fixtures, smart speakers, smart blinds, smart TVs, smart sensors, and smart cameras. These entities can be centralized or decentralized. The smart hubs 111 and 112 can be configured to integrate multiple technologies and provide application control and automation for each of the smart devices 121-126.
[0058] The smart environment 101,102 may also refer to any type of 3GPP communication / wireless network system, such as cellular communication technology. The radio access network (RAN) includes access devices such as base stations (e.g., 5G gNB), mobile base stations, reflective intelligent surfaces (RIS), smart repeaters, and wireless sensing devices (e.g., those integrated into base stations). These provide communication and / or sensing functions / services to target devices (e.g., 5G user equipment (UE)) or entities (e.g., people using or being sensed by the UE). The RAN can be orchestrated by a core network (e.g., 5G CN).
[0059] Furthermore, a first cloud-based service 131 can be provided, configured to support at least one of the smart hubs 111,112. This support includes uploading data from smart devices 121-126 to the first cloud service 131 via their respective smart hubs, executing automation rules by the first cloud service 131, and downloading the resulting commands (e.g., light on / off commands) to one or more of the smart devices 121-126 via their respective smart hubs.
[0060] The advantage of this cloud-managed hub approach is that, because smart hubs 111 and 112 do not require much memory or processing power, it is possible to implement a mobile app-centric system that is less complex and costly, and does not require a desktop computer for setup or programming.
[0061] Furthermore, a second cloud-based service 132 configured to support AI and machine learning (ML) models may also be provided. For example, the second cloud-based service may be an independent party responsible for managing such AI / ML models on the smart device, or it may be a backend server for the device manufacturer.
[0062] Cloud-based services131,132 provide information technology (IT) as a service over the internet or a dedicated network, delivered on demand. Cloud-based services range from full applications and development platforms to servers, storage, and virtual desktops. These cloud-based services allow users to access, share, store, and securely manage information in the "cloud."
[0063] In the following, different embodiments will be described with reference to the respective process diagrams shown in Figures 2 to 4, where arrows indicate the direction of information flow between components shown at the top of the diagram, and time progresses from top to bottom. The steps in these flowcharts can be executed or initiated, at least partially, by the respective instructions of the software programs / routines that control the controllers provided to each device or component.
[0064] Figure 2 schematically shows a signaling and process diagram of an embodiment for managing and acquiring a semantic control model for smart device registration and smart device control. Not all steps or entities in the process are necessarily required. Steps can be performed once or multiple times and can be processed in different orders.
[0065] The entities involved in the process in Figure 2 (smart device 121, first smart hub 111, first and second cloud-based services 131, 132) correspond to, or are at least similar to, the entities with the same reference numbers as in Figure 1.
[0066] In step 201, the smart device 121 joins the smart environment (e.g., smart environment 101 in Figure 1) and the smart hub 111. This can be achieved through a pairing process and other authentication processes that exchange the information necessary for the devices to establish an encrypted connection. This includes authenticating the identities of the two devices to be paired, encrypting the link, and distributing keys to enable the restoration of security upon reconnection.
[0067] Next, in step 202, the smart hub 111 verifies the device and stores information related to the smart device 121, such as the type of AI model required.
[0068] In step 203, the smart hub 111 can contact the first cloud-based service 131 to obtain a first AI model and / or configuration suitable for the smart device 121 and / or the smart hub 111. Obtaining the AI model and appropriate configuration is performed in step 204.
[0069] In step 205, the smart hub 111 can contact a second cloud-based service 132 to obtain a second AI model and / or configuration suitable for the smart device 121 and / or the smart hub 111. Obtaining the AI model or appropriate configuration is performed in step 206.
[0070] In steps 203, 204, 205, and 206, the first cloud-based service 131 may be managed by the operator of the smart hub 111, and the second cloud-based service 132 may be managed by the smart device manufacturer. The first and second AI models / configurations may be complementary or overlapping. For example, they are complementary if the first AI model is more suitable for the smart hub 111 and the second model is more suitable for the smart device 121. They are overlapping if, for example, the first / second AI models can be applied to the smart hub 111.
[0071] In step 207, the AI model and / or settings are saved to the smart hub 111.
[0072] In step 208, the smart hub 111 distributes the AI model and / or settings to the smart device 121.
[0073] Finally, in step 209, the smart device 121 applies and / or saves the AI model and / or settings.
[0074] In related embodiments, the smart hub 111 or other entities can customize the generated AI model to suit a specific smart environment (e.g., smart environment 101) of a participating smart device (e.g., smart device 121), and it can be deployed taking into account one or more of the available smart devices (e.g., smart devices 121-123) (and their capabilities), user preferences, smart device locations, etc.
[0075] Figure 3 schematically illustrates the signaling and process diagram for operating a smart device (e.g., smart device 121) in a first environment (e.g., smart device 101), based on prompts and a semantic control model for smart device control. Not all steps or entities in the process are necessarily required. Steps can be executed once or multiple times and can be processed in different orders.
[0076] The entities involved in the process in Figure 3 (smart device 121, first user 105, first smart hub 111, first cloud service 131) correspond to, or are at least similar to, the entities with the same reference numbers as in Figure 1.
[0077] In step 301, the first user 105 gives a command (for example, a voice command) to the first smart hub 111.
[0078] In step 302, the first smart hub 111 processes the command. For example, it can perform local processing (checking for similar commands, associating commands with users, etc.) and derive / select prompts.
[0079] In step 303, the first smart hub 111 can contact the first cloud-based service 131 to obtain a prompt command appropriate to the received command, enhanced with the processing and data obtained in step 302 (for example, if the smart hub 111 cannot derive / select a prompt in step 302). Obtaining an appropriate prompt is performed in step 204.
[0080] In step 305, the first smart hub 111 can store received prompts and locally generated or acquired prompts.
[0081] In step 306, the first smart hub 111 sends a prompt to the smart device 121.
[0082] Finally, in step 307, the smart device 121 applies / saves the prompt along with its AI model and / or settings, and estimates the settings of the smart device.
[0083] In a related embodiment, the first smart hub 111 can store an AI model and apply prompts to it to infer smart device settings in step 307. The first smart hub 111 can send the smart device settings to the smart device in an additional step 308 (not shown in Figure 3), and the smart device 121 can apply them in a further step 309 (not shown in Figure 3).
[0084] In another related embodiment, the first smart hub 111 (or backend server) may adapt the model and prompts based on user feedback, and if the first user 105 does not like the output of the smart device 121 after step 307, the model and / or prompts may be adjusted to better suit the first user 105's preferences.
[0085] Figure 4 schematically shows the signaling and process diagram of an embodiment for controlling the smart device 121 of Figure 3 in a second environment (e.g., the second environment 102 of Figure 1) based on the smart device settings in a first environment (e.g., the first environment 101 of Figure 1). Not all steps or entities of the process described herein are necessarily required. Steps can be performed once or multiple times and may be processed in different orders.
[0086] The entities involved in the process in Figure 4 (smart devices 121, 122, 124, 125, first user 105, second user 106, first smart hub 111, second smart hub 112) correspond to, or are at least similar to, the entities with the same reference numbers as in Figure 1. For example, smart devices 121, 122, 124, and 125 could be smart devices such as smart TVs, or smart cameras and smart microphones for video conferencing based on Facebook's Portal TV.
[0087] In step 401, the first user 105 initiates interaction with the first smart device 122, triggering sharing of the first smart environment with the second user 106 of the second smart environment. Sharing can also be performed by saving the information to a database and reproducing it at a later point in time.
[0088] In step 402, the first smart device 122 sends a command to the first smart hub 111 to retrieve smart environment settings, which may include a list of AI / ML models and the currently used prompts.
[0089] As mentioned above, in the case of "scene to prompt," a device (e.g., smart device 122, first smart hub 111) can be used to observe scenes and environments (of a smart home), extract prompts that indicate the scenes, and operate the device in both or either a local or remote environment.
[0090] In step 403, the first smart hub 111 transfers the acquired smart environment settings to the first smart device 122.
[0091] In step 404, the first smart device 122 can create or utilize an existing (direct or indirect) connection to the second smart device 124 in the second smart environment (e.g., through the first smart hub 111 or a cloud service (e.g., the first or second cloud services 131, 132 in Figure 1)). After the connection is established, the first smart device 122 can provide the second smart device 124 with a list of AI / ML models and / or currently used prompts (i.e., smart environment settings) in the first smart environment. Additionally or alternatively, communication can also occur between the smart hubs 111 and 112 instead of the smart devices in the smart environment.
[0092] In step 405, the second smart device 124 transmits the acquired / extracted settings (smart environment settings) to the second smart hub 112.
[0093] Finally, in step 406, the second smart hub 112 provides the necessary smart environment settings (e.g., the necessary AI / ML models and / or the necessary prompts) to another smart device 125 in the second smart environment, so that the first smart environment settings can be replicated in the second smart environment.
[0094] In a related embodiment, this process can simply transfer standard smart device settings from a first smart environment to a second smart environment without the need to intervene with a semantic control model (e.g., an AI / ML model) and / or prompts.
[0095] In a related embodiment, a first smart device (e.g., 122) in a first environment can acquire a snapshot of the first environment (e.g., a 3D recording determining images, sounds, lighting conditions, mood conditions, acoustic conditions, etc.), identify smart devices that may be influencing the first smart environment, and apply a first model to obtain a semantic representation of the snapshot. The semantic representation / command, and optionally a list of smart devices in the first environment, can be shared with a second (smart) device in a second environment. The second (smart) device can capture the current second environment (e.g., lighting conditions, acoustics, etc., and / or smart devices in the second environment) and apply the semantic representation / command obtained from the first environment to a semantic control model of the second environment that takes the current second environment into account. For example, this can be applied to a use case where two users are on a video call, and the camera of a first device in the first environment captures the lighting and smart lights in that environment, extracts a semantic description of the lighting in the first environment (e.g., four smart bulbs that create a cozy, warm light to match the sunset visible from the window), and transmits this semantic description / prompt to the second environment. A second device in the second environment (e.g., a hub) can then use a semantic control model to determine which four smart bulbs (e.g., four out of ten) to use and with what settings to simulate the scene in the first environment. Furthermore, the second device (e.g., a hub) can also instruct a television positioned to match the window location in the first environment to send semantic commands (sunset, warm light) to recreate the sunset.
[0096] In another related embodiment, the first smart device 122 may infer the settings necessary to create a similar scene by not obtaining the settings of the smart environment, but by identifying other smart devices (e.g., smart device 121) in the first environment, determining their settings, and / or by observing the first environment itself.
[0097] In a further embodiment, a smart device, such as a smart TV, may be comprised of a generative model capable of generating content tailored to the environment and / or user. For example, a smart TV may have a semantic control model (e.g., a generative AI / ML model) capable of creating video or audio content based on prompts presented by the user or derived from the context. Prompts may specify genre, theme, mood, style, length, or other aspects of the desired content. The smart TV may also use sensors or cameras to detect ambient environmental conditions such as lighting, temperature, and noise, and adjust the content accordingly. Furthermore, the smart TV can personalize content using user feedback such as ratings, preferences, and viewing history to improve the quality of the generative model. The smart TV may also have different settings for the generative model depending on the device's capabilities, such as memory capacity, resolution, and sound quality. Moreover, the smart TV may have multiple generative models for different users and can select the appropriate model based on user ID and selections. Thus, the smart TV can generate and reproduce content locally generated by the generative model that meets the user's needs and expectations. For example, if a user sets a specific piece of music on a music system and dims the lights, the smart device / TV can use sensors and an AI model to recognize the type of music and lighting, and generate visual content that matches the recognized music type and lighting. This embodiment is shown in Figure 5, where 500 represents a smart device such as a smart TV. The smart device is equipped with sensors 501 (e.g., a microphone or camera) that allow the user to observe the environment (lighting conditions, sound, etc.) and extract the mood type of the environment using a "scene to semantics" model 503. The smart device may also be equipped with a radio 502 that can receive commands from other devices (e.g., a smart hub), such as instructing the smart TV to operate in AI generation mode.The smart device may have a local database or memory 504 that stores settings that can be adapted to different users and environments. The outputs of 502, 503, and 504 supply semantics to a control model operating on the processor 505, for example, a “semantics to video” model that generates content for at least the actuator 506 (e.g., the display of a smart TV). In other embodiments, a smart device (e.g., a device surrounding a control smart device such as a smart TV) can receive a prompt (e.g., the word “forest” received from the smart device) and generate a corresponding visual or audio output. For example, a smart audio device (e.g., a Sonos device) may have a semantic control model (e.g., a generative AI / ML model) that generates the sound of wind passing through leaves, or a smart lighting device may have a semantic control model (e.g., a generative AI (ML model)) that adapts the color of light to green. Additionally or alternatively, the control smart device may have a semantic model that enables it to create a control input for another smart device given a prompt.
[0098] Therefore, smart systems, such as smart lighting systems (e.g., Hue systems), can synchronize with data inputs such as movies and games by monitoring and analyzing data exchanged via high-definition multimedia interfaces (HDMI) or data received from other devices in the smart environment, either locally or remotely. This data representation can be standardized. For example, the data itself (e.g., a movie) (e.g., a Netflix series) can be encoded with prompts indicating how the smart device should behave. To achieve this, a smart device (e.g., a television) can be equipped with a smart home interface (e.g., a Matter interface) that can control the smart device based on its control data. This control data can be prompt-based.
[0099] In another related embodiment, an "image to semantics / text" model (Model 1) can be used to obtain the semantics of a video or image captured by a smart device (e.g., a camera) or transmitted to a device (e.g., a television). For example, the model can generate a text description or semantic representation of the video / image content, such as objects, actions, attributes, or scenes displayed on a television or captured by a camera. Then, a second model, such as a "semantics to control" model (Model 2), can be used to obtain control signals for another smart device (e.g., lighting or a speaker) based on the semantics extracted from the video / image. For example, the second model can generate prompts specifying how the smart device should behave according to the semantic input. The prompts can be encoded in a standardized format such as JSON or XML and can include parameters such as color, intensity, duration, and pattern. Alternatively, the prompts can be natural language sentences describing the desired effect. The smart device decodes the prompts and adjusts its settings accordingly. For example, if a video / image shows a sunset, the model might generate a prompt such as "a warm orange light is gradually fading," and the smart light can then adjust its color temperature and brightness accordingly. Model 1, the first model, works with devices such as televisions and smart cameras, while Model 2, the second model, works with smart controllers and end devices, so it is desirable to standardize the prompt format. The first and / or second models can be configured to be user-specific; for example, if the presence of a particular user is detected, a specific model version is selected, and / or the model is adjusted to include user information (e.g., user ID and user characteristics) to manipulate the output.The first and second models may operate on the same device or on different devices. For example, the first and second models may operate on a camera and a television, with the camera sending commands to smart lighting, or the first model operating on the camera and sending semantic commands to the lighting, which then executes the second model to obtain specific settings. For example, a security / surveillance camera installed in a smart home may execute the first model to determine whether the motion is that of a burglar, a female homeowner, a male homeowner, or an animal (e.g., a dog). Based on the identified subject, the smart camera can send semantic information to other smart devices (e.g., a smart lighting system or speaker). The smart devices, such as a smart lighting system or speaker, can then use the second model to determine the settings of the smart devices. For example, if an intruder is detected, the light intensity can be set to maximum; for known or authorized individuals, the light intensity and / or color can be set to values optimized for that person's mood, taking into account their daytime activities; and for animals, the light can be adjusted to suppress the presence of the identified animal; for example, a bright, flashing light may be preferable for cats, while the light intensity can be reduced for mosquitoes. As a further option, the device can be adapted to be configured with a control policy that can determine, for example, the following: Under what circumstances should Model 1 extract which semantic data? For example, depending on the context, Model 1 may need to extract variable amounts and / or different semantic information, such as extracting mood from the face of an acquaintance returning home, but not from someone who is not an acquaintance, or extracting the type of animal, for example; • Which semantic data can (or cannot) be shared with local and remote devices. For example, mood information is not shared with any devices, for example, a person's type (owner or thief) can be shared with devices on the local network, and for example, whether a person is human or animal can be shared with devices outside the local network (e.g., a cloud server).
[0100] In another relevant example, a first device (e.g., a camera) records the environment (e.g., mood, sound, movies / audio being played on a third device (e.g., a smart TV, a smart speaker)), and the first model can acquire the semantics of the environment. The first device can then send prompt commands to a third device (e.g., smart lighting), which can be determined by a second model, or the second model itself can be used to determine specific settings for the third device, and the obtained settings can be sent to the third device. For example, the first device could be a surveillance camera used in a home for security reasons. The first device could monitor a movie being watched and use this input to determine settings for a smart lighting system. In embodiments, a smart device may include a time synchronization module that enables the smart device to perform functions according to a global time standard. For example, the time synchronization module can receive time synchronization signals from a network source such as a server, router, or smart hub and adjust the device's internal clock accordingly. This allows the smart device to coordinate time-sensitive operation with other smart devices in the smart environment, ensuring that operations are performed at precise time intervals. For example, smart devices can control the lighting, heating, and security systems of a smart home and synchronize them according to the user's preferences, schedule, and location. For instance, a prompt command may include a time reference used to synchronize settings obtained from the prompt command.
[0101] In one embodiment, the first device needs to obtain the device profile of the second device before it can obtain the settings of the second device given a prompt command. The device profile can determine the type of settings required. The first device can use this device profile and the prompt command to obtain the settings of the second device given a semantic control model. Additionally or alternatively, the first device can examine the device profile to determine which functions should be controlled and / or within what range of values. For example, the first device may receive a device profile of the second device (e.g., a smart lighting device) indicating that several characteristics (e.g., hue, light temperature, light intensity) are adaptable, and the first device may determine that the second device wishes to adapt only some of its characteristics (e.g., light temperature and light intensity of the smart lighting device), and the prompt command may include the prompt itself (e.g., "warm light for children") or a subset of the characteristics to be adapted according to the device profile (e.g., light temperature and light intensity). Additionally or alternatively, a second device may obtain a configuration containing a subset of properties that need to be applied from a third device (e.g., a smart hub). In this case, the smart hub is used to determine which subset of properties needs to be controlled, the first device determines a semantic prompt, and the second device can obtain a specific configuration based on the semantic prompt and the determined subset of properties.
[0102] In the above embodiments (for example, in the case of a “scene to prompt” implementation), the smart device can be configured as a single device or as an array of multiple devices, and can typically be configured to capture a portion of the scene, but not necessarily in real time. Examples include cameras, microphones, and motion sensors. Some devices can capture stimuli outside the range of human perception (e.g., infrared cameras, ultrasonic microphones) and “downconvert” them into a form that is perceptible to humans. Some devices can have arrays of sensor elements that provide a more detailed or augmented impression of the environment (e.g., multiple cameras capturing a 360° view, or multiple microphones capturing a stereo or surround sound field). Sensors with different modalities (e.g., sound and video) can also be used in combination. In such cases, different data streams need to be synchronized. The transmitting device equipped with sensors can be VR / AR glasses or simply a UE.
[0103] Furthermore, (optional) rendering devices (e.g., audio, video) can be provided to render parts of the scene in real time. Examples include video displays or projectors, headphones, speakers, and haptic transducers. Some rendering devices may have arrays of multiple rendering elements that provide an augmented or more detailed impression of the captured scene (e.g., multiple video monitors or speaker arrays for rendering stereo or surround sound audio). It is also possible to use rendering devices with different modalities (e.g., sound and video) in combination. In this case, the rendering subsystem must ensure that all stimulus channels are rendered synchronously.
[0104] Furthermore, in embodiments, a "text-to-image" model (such as a latent diffusion model) can be used as a prompt.
[0105] Furthermore, in this embodiment, image classification can be performed at high speed on a low-capacity device.
[0106] Recent technologies are attempting to address the problem of semantic loss. Leading techniques in this field are represented by text-to-text conversion, a method for dynamically learning embeddings that represent previously unseen objects. Therefore, guide images are used to ensure that the learned embeddings adequately represent the observed reality (represented by the guide images). Such learned embeddings are called "learned prompts."
[0107] In concrete examples, smart devices (UEs, cameras, microphones, light sensors, etc.) can observe a scene using some kind of sensing (including audio, video, and image capture), segment that observation, and generate instances related to identified semantic objects. For well-reproducible parts, such as the atmosphere of a scene, the transmitter can directly generate or extract appropriate learning prompts. These prompts can then be used, for example, to reproduce the same scene, such as the atmosphere of a scene, in a remote location.
[0108] In some embodiments, the processes described above and in other embodiments can be performed predictively. For example, prompts can be generated predictively based on locally predicted motion in the scene or communication parameters with the receiving device (e.g., delay). These can be transmitted in advance, and then, when a true change is observed in the scene, only a short command (e.g., including a small correction factor) regarding the prompt to be used can be sent to correct the difference between the observed scene and the predicted scene (see Zhihong Pan et al.: “EXTREME GENERATIVE IMAGE COMPRESSION BY LEARNING TEXT EMBEDDING FROM DIFFUSION MODELS”). This makes it possible to reduce additional delay to near zero even with very high-bandwidth content.
[0109] In embodiments relating to multimedia data, prompts can be linked to metadata containing multiple parameters. The metadata can play a role in facilitating the reconstruction of compressed data, particularly in the case of data related to multimedia data or video.
[0110] In the first variation, prompts are linked to time ranges (e.g., to enable the reproduction of scenes, videos, and audio), and only a single prompt needs to be sent for a given period. For this purpose, prompts can be linked to metadata that includes parameters such as an estimated decay time (e.g., the number of frames in a video) and are expected to be, or known to be, valid during that period. Prompts can optionally be used beyond this decay period, but at the cost of increased semantic loss.
[0111] In the second variation, moving images or scenes can be linked to an initial prompt, an end prompt, a time range, and a moving pattern. The receiver can be configured to use a reconstruction model to reconstruct the motion (e.g., scene changes) from the prompts and metadata.
[0112] In a third variation, it may be necessary to synchronize prompts or compressed data related to different data types (e.g., audio and video). For example, a video prompt may contain metadata with a “link” to its audio prompt. This third variation is particularly important when both audio and video undergo generative compression and prompts need to be linked. Various ways of doing this include matching the “decay times” of audio and video prompts (as described above), and / or using a trained reconstruction model that reconstructs both audio and video from a shared latent space, designed prompts for that latent space, and / or using generative compression only for video and a different method for audio (in which case the audio may simply be temporally linked to video frames). For example, compressed data of different data types may be linked to metadata that determines the time frame at which the data is rendered. For example, if an image in a video is associated with the prompt "Alice Waking in the Street" with metadata "Time: [0.00'', 5.00'']" and audio in the video is associated with the prompt "Alice say: "Hi darling"" with metadata "Time: [2.00''-3.50''])", then synchronization should play the audio associated with the audio prompt, starting at 2 seconds and lasting for 1.50 seconds. This also affects the rendering of the video, as the audio prompt indicates that Alice is speaking, rendering it so that Alice says "Hi darling" between 2 seconds and 3 minutes 50 seconds.
[0113] In a fourth variation (e.g., audio and / or video prompts), the prompt metadata may include parameters to determine how the reconstructed data (e.g., audio) is mixed. For example, whose voice is louder when two people are speaking at the same time, or who is ahead of the other when two people are walking close together. Generally, this is a key feature of generative audio / image / video compression algorithms that use text conversion. The learned prompts may need to account for overlapping data objects, such as overlapping voices. In some cases, it may be more efficient to have multiple learned prompts (e.g., voice A, voice B, degree of overlap).
[0114] A fifth variation, for example in a metaverse scenario, allows a person's image to be linked to a prompt linked to an existing avatar. For example, a user's photo can be linked to prompt "S" which is linked to avatar Y. The avatar is a digital representation of the user (participant), and this digital representation can be exchanged with one or more users as mobile metaverse media (along with other media such as audio).
[0115] As an example of a fifth variation, an avatar call can be established, which is similar to a video call in that it is both visual and interactive, and provides participants with raw feedback about their emotions, attention, and other social information. Once an avatar call is established, the communicating parties can provide information in the uplink direction of the network. The terminal device (e.g., UE) can capture the facial information of the call participants and locally determine the encoding of the captured facial information (e.g., consisting of data points, coloring, and other metadata). This encoded information is transmitted in the form of a media uplink and provided to the other participants in the avatar call by the IP Multimedia Subsystem (IMS). Once the media is received by the participant's terminal device (e.g., UE), the media is rendered as a two-dimensional (or three-dimensional) digital representation.
[0116] In related variations, the audio prompt itself may indicate the meaning of the message or the length of the utterance, but the content may be generated locally and / or by other means, such as a family of generative pre-trained transformers (GPTs) for language models, like ChatGPT. The generated text can be converted to text-to-speech and fit within the required time interval.
[0117] As further concrete examples related to the structured vocabulary and generative models of prompts, the descriptive language can be a structured language / vocabulary, a human-readable language, or a pseudo-language (see, for example, Rinon Gal et al.: “An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion”).
[0118] Furthermore, in embodiments that can be used in combination with other embodiments or independently, a user may have an application that runs on a smartphone or computer and allows the user to create and edit content such as smart home scenes and audiovisual material such as videos, using at least a generative model and an interpretive programming language, or means supported by them. The user, who can be the content owner, can use prompts to create audio, video, images, etc., from a model containing target content (one example being audiovisual material, but which may also include other types of data and content). The user can then save the content and / or send the content locally or remotely to at least the (receiving) user / device.
[0119] In further related embodiments, the generated content may be considered or referred to as synthetic data specified by a "content program" which is taken as input by an interpreter of an "interpreted programming language" that relies on a generative model for interpretation.
[0120] In a further related embodiment, the semantic control model provides several options when the user attempts a prompt, allowing the user to select one of the model's outputs. Since multiple generative models may be involved, the user may also include a model identifier and / or version to enable the receiver to regenerate the same content. Because the generated audio / images / videos, etc., may not fully satisfy the user, the user can adjust the output and make changes to the data (during the content creation process). During the content creation process, the model is enhanced based on user-defined inputs by the user acquiring audio / video samples and assigning them to prompts. The user can create content (e.g., videos or smart home scenes) using a programming language, and the “content program” (or program) is written as in the following example: [New smart home scene, Duration 12 seconds, Resolution xyx] [Generative model ID xyz, version vxyz] [Background smart home scene: part in spring, sunny weather] [Background smart home sound: happy piano music] [Prompt: “funny video, small white cat”; Prompt_output: #3; Start time: 1 second; Duration: 10 seconds; Action: “walks from left to right”] [Prompt: “funny video, big fat dog”; Prompt_output: #2; Start_time: 2 second; Duration: 9 seconds; action: “walks from right to left”; action: “smells cat and smiles”; action: “follows cat”]
[0121] In this example of a "content program," new commands are given, for example, on new lines between parentheses. This requires a standard to determine which are new commands.
[0122] In this example, one or more (generative) models are shown, which may be identified by name or URL, for example. These generative models are used by the interpreter to generate the content specified by the "content program".
[0123] In this example, there may be several keywords that can help determine the type of content to be generated. Examples of these keywords include "new", "smart home scene", "duration", "Resolution", "Generative model", "Background image", "Background sound", "Prompt", "Prompt_output", "Start_time", and "Duration".
[0124] In this example, there may be a standard way to indicate which action is associated with a keyword. For example, the sequence [Keyword: Action;] indicates that a new keyword is started by "Keyword," and the action associated with that keyword appears after ":" and ends with ";".
[0125] This information can be entered as text or through other types of user interfaces, such as a graphical user interface.
[0126] Users can play back the generated smart home scene and edit it further until they are satisfied. At this point, users can release or publish it.
[0127] To verify the user who published the content, the data (or its hash value) used to generate the (audiovisual) content can be signed by the user. A fingerprint can also be made available (for example, attached to the content or made available in a public repository or blockchain), and that data can be made public so that other users can create further content based on it.
[0128] The "content program" can be played back in slightly different ways in different smart environments, depending, for example, on the type of smart device available.
[0129] When playing a "content program," the "content program" or the user can specify which smart devices can be used to play the content.
[0130] "Content programs" can be extracted from smart environments and scenes, and can be played back in other smart environments.
[0131] Currently, the use of semantic control models to control devices such as smart devices is being explained. If training is required, it necessitates collecting and utilizing data such as environmental records (e.g., photos, videos, audio), (smart) device configuration parameters, and user feedback and commands (e.g., comfortable environment, soft lighting). This information can be collected, for example, through a smart hub. This information can be transmitted to a server (if necessary) with the user's consent. This information can be used, for example, to train semantic control models based on diffusion technology.
[0132] Each embodiment can be combined with others or used independently as needed to address requirements and / or missing capabilities.
[0133] In summary, devices and methods for controlling smart environments are described, enabling efficient control of smart devices using semantic control models (such as generative AI or ML models), and allowing smart devices to be configured to store semantic control models. Furthermore, smart devices can be configured to receive prompts and apply the received prompts to the stored semantic control models to derive and reproduce smart device settings.
[0134] Furthermore, the present invention can be applied to mobile phones, vital sign monitoring / telemetry devices, smartwatches, detectors, vehicles (vehicle-to-vehicle (V2V) communication or more generally vehicle-to-everything (V2X) communication), V2X devices, IoT hubs, low-power medical sensors for health monitoring, IoT devices including medical (emergency) diagnostic and treatment devices for hospitals or first responders, virtual reality (VR) headsets, and the like.
[0135] Other modifications of the disclosed embodiments can be understood and implemented by those skilled in the art in carrying out the claimed invention, from a consideration of the drawings, disclosures, and appended claims. In the claims, the word “has” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude plurality. A single processor or other unit can fulfill the functions of several items enumerated in the claims. The mere fact that certain means are described in mutually different dependent claims does not imply that combinations of these means cannot be used advantageously. The foregoing description details certain embodiments of the invention. However, however detailed the foregoing may be in the text, it will be understood that the invention can be carried out in many forms and is therefore not limited to the disclosed embodiments. It should be noted that the use of certain terms in describing certain features or embodiments of the invention does not mean that the terms are redefined herein to limit them to certain features of the features or embodiments of the invention to which they relate. Furthermore, the expression “at least one of A, B, and C” should be understood disjunctively, i.e., “A and / or B and / or C”.
[0136] A single unit or device may perform the functions of multiple items cited in a claim. The mere fact that certain means are described in different dependent claims does not imply that combinations of these means cannot be used advantageously.
[0137] The operations shown in the above embodiments (for example, Figures 2 to 4) can be implemented as program code means for a computer program, or as dedicated hardware for associated network devices or functions, respectively. The computer program may be stored and / or distributed on a suitable medium such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but it may also be delivered in other forms such as the Internet or other wired or wireless telecommunications systems.
Claims
1. A device for controlling at least one smart device, in particular a sensor and / or actuator, in a smart environment, wherein the device is configured to receive at least one semantic control model associated with the at least one smart device and at least one semantic prompt associated with a scene, and to use the semantic control model and at least one of the semantic prompts, wherein the semantic control model is executed on the device.
2. The apparatus according to claim 1, configured to deliver the semantic control model and at least one of the semantic prompts to the at least one smart device.
3. The apparatus according to claim 1 or 2, configured to receive a command, derive or obtain a control prompt from the command, and apply the control prompt to the semantic control model to obtain device settings.
4. The apparatus according to claim 3, configured to use the apparatus settings to control the at least one smart device.
5. The apparatus according to claim 3, configured to improve the semantic control model and / or the control prompt based on user feedback.
6. The apparatus according to claim 3, configured to receive indicators of the semantic control model and / or control prompt from a remote device.
7. The apparatus according to claim 1 or 2, configured to contact a cloud-based service and / or blockchain to obtain and / or verify the semantic control model for the smart device.
8. The apparatus according to claim 1 or 2, configured to adjust the semantic control model to suit the smart environment of the smart device by taking into consideration one or more of the functions of the available smart device, user preferences, and location of the smart device.
9. The apparatus according to claim 1 or 2, configured to control a smart device in a second environment based on the semantic control model and apparatus settings including control prompts for a first environment.
10. A hub or smart device having the apparatus described in claim 1 or 2.
11. A smart device configured to receive or acquire device settings, including a semantic control model and control prompts for controlling the smart device.
12. The smart device according to claim 11, configured to infer the device settings by identifying other smart devices in a smart environment and determining their settings.
13. A system comprising the hub according to claim 10 and one or more smart devices according to claim 11 or 12.
14. A method for controlling smart devices in a smart environment, A step of receiving at least one semantic control model and semantic prompt associated with a smart device in the smart environment, wherein at least one of the semantic control model and semantic prompt is associated with or applicable to the smart device. A method comprising the steps of using the semantic control model and at least one of the semantic prompts to control the smart device, wherein the semantic control model is executed on the smart device.
15. A computer program that is executed on a computer and causes the computer to perform the method described in claim 14.