A method for conditionally transmitting a prompt as input to large language model
By conditionally transmitting prompts to LLMs based on user input complexity and augmenting with exemplars, the method optimizes computational efficiency and accuracy in smart home automation, addressing the challenges of LLMs in smart home control.
Patent Information
- Application Number
- PCT/EP2025/050172
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-25
- Filing Date
- 2025-01-06
- Publication Date
- 2025-07-17
AI Technical Summary
Large language models (LLMs) used for smart home automation face challenges such as potential errors, high computational costs, and resource requirements, necessitating optimized control performance and efficiency while maintaining accuracy and low operational costs.
A method for conditionally transmitting prompts to LLMs based on user input complexity, using local controllers to handle simple tasks and augmenting prompts with exemplar pairs for complex tasks, thereby optimizing computational resources and enhancing accuracy.
This approach reduces computational overhead and improves accuracy by selectively engaging LLMs for complex tasks, ensuring efficient and precise control of smart home devices like lighting systems.
Smart Images

Figure EP2025050172_17072025_PF_FP_ABST
Abstract
Description
[0001] A METHOD FOR CONDITIONALLY TRANSMITTING A PROMPT AS INPUT TO
[0002] LARGE LANGUAGE MODEL
[0003] FIELD OF THE INVENTION
[0004] The invention relates to a method, controller, system and computer program product for conditionally transmitting a prompt as input to a large language model.
[0005] BACKGROUND OF THE INVENTION
[0006] Large language models can be used to control various smart home devices, including lighting systems, thermostats, security systems, and entertainment systems and have the potential to revolutionize the way users interact with such smart home devices. These models can interpret natural language commands spoken by users and translate them into machine-readable instructions that control the devices. For example, a user could say, "Turn on the living room lights" and the model would understand the command and send a signal to the appropriate device to turn on the lights. By enabling natural language interaction with smart home devices, large language models have the potential to simplify and streamline the home automation experience for users. There is thus a growing interest in the use of large language models for smart home automation and related fields.
[0007] SUMMARY OF THE INVENTION
[0008] Using Large Language Models (LLMs) for controlling an actuating system, such as a lighting system, HVAC system, etc. can offer several benefits, including increased automation, improved accuracy, higher personalization to the needs of the specific user at hand (e.g., elderly, technology enthusiast) and enhanced energy efficiency. LLMs can be trained on vast amounts of data to understand language patterns and respond to user requests in a more natural and intuitive way. This can lead to more accurate and personalized lighting experiences, as the LLM can analyze user preferences and adjust the lighting accordingly. Additionally, the ability of LLMs to interpret and respond to natural language input can make them more accessible to a wider range of users, including those with disabilities or language barriers. However, there are also some drawbacks to LLMs. One of the main concerns is the potential for errors or inaccuracies in the model, which could lead to suboptimal control conditions. Additionally, LLMs may be expensive to call and use, and may require significant computational resources to operate effectively. Thus, there is a need to optimize the control performance and efficiency of an actuator system while maintaining accuracy and low operational cost. According to a first aspect, the object is achieved by a method for conditionally transmitting a prompt as input to a large language model, and for using an output of said LLM to control an actuation system, for example a lighting system. An actuation system refers to a network of interconnected devices, appliances, and systems that can be controlled and automated to perform various tasks and functions, such as lighting control, temperature control, security, entertainment, and more. The LLM receives a natural language textual user prompt to carry out a task associated with the control of one or more actuators (devices) of the actuation system, processes said prompt and generates a corresponding output to control the one or more actuators (devices) of the actuation system. The LLM output is based on the patterns and contexts it has learned from its training data and / or the reasoning abilities it has acquired during the training whereas the reasoning abilities enable the LLM to make out-of-distribution inferences e.g., the user stating “I am happy” and the LLM figuring out corresponding lighting settings for this specific user. Additionally, and / or alternatively the LLM may receive an image, video, audio user prompt (multimodal LLMs). The control output may comprise a set of actions (instruction code) that may be processed by a controller of the actuation system to control the one or more actuators. For example, if a user says, "turn on the lights in the kitchen", the LLM can analyze the input and generate an output command to turn on the lights in the kitchen. Similarly, if a user says, "dim the lights to 50% in the living room", the LLM can generate an output command to dim the lights to the specified level in the living room.
[0009] The method comprises receiving a textual user input indicative of an intention to control one or more devices of the actuation system. This can include tasks such as turning on a light, adjusting the temperature, opening a door, arming / disarming an alarm system, creating ambience with light, etc. For example, a textual user input “generate a meditation scene with lights” or just “I am very stressed” may involve the intention to control one or more lighting devices according to a certain color and intensity that will create a calming and relaxing atmosphere. The input may also involve controlling other devices, such as a sound system or smart blinds, to create a fully immersive meditation experience. The method further comprises analyzing said textual user input to determine at least one characteristic of said textual user input and determining a level of complexity of said textual user input based on said at least one characteristic. Only if said level of complexity is above a first complexity threshold, the method comprises generating a prompt, said prompt comprising at least said textual user input. For example, the prompt may comprise said textual user input and optionally configuration data of the lighting system. The method further comprises transmitting said prompt to said LLM for example via an output interface. The method further comprises receiving the output of said LLM and controlling said actuator system according to said output. The prompt comprises at least said textual user input and optionally further data indicative of one or more control characteristics of the actuation system, such as device types, device control parameters, LLM output format, relevant sensor information, information about the user (e.g., elderly, teenager, special medical conditions, etc.), preferences and habits of this specific user etc. If said level of complexity is below said first complexity threshold, the method refers from transmitting said textual user input to the LLM. That is, the method may comprise directly controlling the actuation system. For example, the method may comprise analyzing said textual user input to determine one or more actuation control parameters and controlling the actuation system according to said one or more control parameters. Said analysis may comprise analyzing the user input using a local LLM (for example a smaller LLM that runs locally on the edge).
[0010] According to a second aspect, the object is achieved by a controller for conditionally transmitting a prompt as input to an LLM, and for using an output of said LLM to control an actuation system, for example a lighting system, said controller configured to receive a textual user input, via an input interface, indicative of an intention to control one or more devices of the actuator system. The controller is further configured to: analyze said textual user input to determine at least one characteristic of said textual user input; determine a level of complexity of said textual user input based on said at least one characteristic; and only if said level of complexity is above a first complexity threshold, generate a prompt, said prompt comprising at least said textual user input, and transmit, via an output interface, said prompt to said LLM. The controller is further configured to receive, via the input interface, the output of said LLM and control said actuation system according to said output. The controller does not transmit (i.e., the controller deters from transmitting) a prompt to said LLM if said level of complexity is below the first complexity threshold. For example, the controller may directly control said actuation system, e.g., analyze said textual user input to determine one or more actuation control parameters and control the actuation system according to said one or more control parameters. According to a third aspect, the object is achieved by a system for conditionally transmitting a prompt as input to a LLM, said system comprising: an input interface, an output interface and at least one controller according to the second aspect.
[0011] According to a fourth aspect, the object is achieved by a computer program product for a computing device, the computer program product comprising computer program code to perform the method steps according to the first aspect when the computer program product is run on a controller according to the second aspect.
[0012] By conditionally calling a LLM based on the complexity of the user input, i.e., the complexity of the task at hand, can help optimize performance and efficiency while maintaining accuracy. For example, for simple lighting or temperature control tasks, such as turning on or off a light or adjusting the temperature by a few degrees, a less complex rulebased system may be sufficient and even more accurate, more interpretable / explainable and may not generate LLM hallucinations, referring to the phenomenon where a language model generates text that is nonsensical, irrelevant, or even disturbing. However, for more complex tasks, such as creating personalized lighting scenes or adjusting the temperature based on occupancy patterns, an LLM can provide more accurate and natural language-based control. By conditionally calling the LLM, only when necessary, one can reduce the computational resources required and improve overall system performance, for instance by lower latency and consistent performance. For instance, a user input (query) such as "turn on the lights" is a simple and straightforward request to control the lighting system that an e.g., internal, for example, rule-based system or controller can handle without calling an LLM. The system can directly turn on the lights based on, for example, a predefined rule or trigger. However, the query "I cannot see very well" is more complex as it requires the system to interpret the user's request in a natural language and respond appropriately. In this case, an LLM can be beneficial as it can analyze the user's request and respond with a more appropriate lighting scene that can improve visibility, such as increasing the brightness or changing the color temperature, for instance if the user is performing a task requiring high acuity such as threading a needle may result in a different LLM output compared to a user getting up from bed at night. The LLM can also take into account other factors, such as the time of day, the current context of the home or the user's preferences (e.g., sustainability goals, arbitration between needs of multiple users such as disturbance level to neighbors), to provide a more personalized and effective lighting solution. Therefore, in cases where the user's request involves more complex language or requires a more personalized response, an LLM can provide more accurate and natural language-based control. The at least one characteristic may be a readability level of said textual user input. The method may comprise analyzing the textual user input provided by the user to determine the readability level of the textual user input indicative of the ease in which the textual user input can be understood (comprehended) by an LLM. Different formulas may be used to assess readability ranging from simple metrics such as word frequency count, number of characters in the textual user input, number of different words in the textual user input, percentage of unique words, number of prepositional phrases, etc. to more complex formulas such as the Flesh Formulas, the Dale-Chall Formula etc. A higher level of readability may be associated with lower complexity level.
[0013] The at least one characteristic may be a number of tasks that need to be completed to carry-out said control of the one or more devices of the actuation (lighting) system. There are various techniques to determine these tasks (by the controller), such as analyzing the language used in the textual user input description, identifying keywords or phrases that indicate specific (pre-determined) actions, and using previous knowledge of similar past textual user inputs to infer what the tasks are. To determine the number of tasks required, the processor may analyze the user's request and break it down into individual tasks that need to be performed. For example, a user input query such as “Find optimal lighting conditions for the people in the living room” may require a series of complex tasks to be performed, such as determining the number of people in an area, classify the people in the area (e.g. elderly person, teenager, toddler), determining activities of the people in the area, rank lighting needs based on activity, identify best lighting conditions for each activity, etc. In this case, an LLM can be beneficial as it can interpret the user's request and perform the necessary tasks in an efficient and accurate manner.
[0014] The controller may be configured to apply a first machine learning model on the textual user input to determine an effectiveness score for said user input and the at least one characteristic is said effectiveness score. The first machine learning model may have been pre-trained based on labeled past textual user inputs and corresponding past LLM outputs, said past textual user inputs being associated with: a negative score, if an active response of the user has been received in response to the corresponding LLM past output; wherein an active response comprises receiving an input from the user, said input for controlling the lighting system; a positive score, if said active response is not received. This may help improve the user experience with the lighting system by predicting potential issues with LLM control outputs based on experience with past interactions with the LLM.
[0015] The at least one characteristic is a number of keywords, from a set of predetermined keywords, present in said textual user input. For example, if the textual user input is "turn on the lights", the controller may analyze said input to identify the keywords "turn on" and "lights" and map them to the intent of controlling the lighting system. On the other hand, for a user input such as, “I am stressed. Can you provide a relaxing atmosphere?”, the processor may identify one keyword, say “relaxing”. Similar, the user may indicate that he is distressed, depressed, feels overwhelmed, has migraine / pain or happy. In such cases, a higher level of complexity may be assigned to the textual user input that comprises a lower number of keywords. This would mean that for such inputs, the system may not easily map the keywords to specific control actions\outputs.
[0016] While LLMs have the ability to process natural language inputs and generate relevant control outputs, they rely on large datasets to learn the patterns and contexts of language. However, in the case of smart home automation, users may have unique or infrequently used commands for controlling their devices. In such cases, the LLM may not have enough data to accurately understand and respond to these commands. Augmenting the prompt to a LLM with exemplar input-output pairs enables the LLM to quickly adapt and learn the patterns and contexts of these commands, improving its accuracy and responsiveness to natural language commands related to smart home automation. On the other hand, there are drawbacks associated with adding exemplars to an LLM. Augmenting the prompt of a LLM (LLM) with exemplars can increase the computational cost of using the model as well as increasing the latency of the inference performed by the LLM. The amount of computational cost depends on the size and complexity of the LLM, and the number of exemplars added. In addition to the computational cost of in context learning, augmenting a prompt with exemplars may increase the runtime cost and latency of the LLM. This is because the model needs to perform additional computations to integrate the exemplars into its decision-making process. It is therefore an object to provide a method that balances the tradeoff between accuracy and computational efficiency. While increasing the amount of data and training examples can improve accuracy, it also requires more computational resources. On the other hand, simplifying the model or reducing the amount of data used for training can improve computational efficiency but may result in lower accuracy and predictability. The method may further comprise conditionally augmenting the prompt with exemplars as input to the LLM, and for using an output of said LLM to control an actuation system, such as a home or building automation system, a lighting system, etc. The method comprises receiving, via an input interface, a textual user input indicative of an intention to control one or more devices of the actuation system. The method further comprises analyzing said textual user input to determine at least one characteristic of said textual user input, such as the readability of said textual user input, the length (number of characters) of the textual user input, the number of sub-tasks (steps) needed to carry out the task of controlling said one or more devices, etc. If said level of complexity is below the first complexity threshold, the method refers from transmitting a prompt to the LLM. If said level of complexity is above the first complexity threshold, the method comprises generating a prompt, said prompt comprising said textual user input, and transmitting said prompt to said LLM via an output interface. If said level of complexity is further above a second complexity threshold, larger than said first complexity threshold, the method may further comprise augmenting the generated resulting in an augmented prompt, wherein said augmented prompt comprises at least said textual user input and one or more exemplar pairs (typically the textual user input appended after the one or more exemplar pairs), and transmitting said augmented prompt to said LLM via the output interface. That is, for simpler tasks that can easily be processed and understood by the LLM, it may not be required to augment the prompt with exemplars. On the contrary, augmenting the prompt may result in additional computational effort that may not be desired. However, more complex tasks may require task-specific exemplars to derive a satisfactory (correct) LLM output.
[0017] Each exemplar pair comprises an exemplary task-specific textual user input and a corresponding language model output. That is, an exemplar pair comprises exemplar conversational sentences between the user and the LLM (i.e., a textual user input and corresponding LLM output) that are specific to (represent) an actuation control task. They are representative sets of examples mapping textual user inputs to corresponding desired control output from the LLM. Each pair includes a textual input by a user associated with the intention of controlling (carry out a task associated with the control of) one or more actuators (devices) of the actuation system, and a control output by the LLM (i.e., a set of actions for controlling the one or more devices to carry out the task). For example, an example pair indicates: “Q: Turn on the lights in the bedroom; LLM output: location: bedroom, target: light, value: off”. By conditionally augmenting the prompt based on the complexity of the task at hand, a tradeoff between accuracy and computational efficiency can be achieved. For example, the LLM can easily process and generate a correct control action for a simple user command such as “turn on the lights in the living room”. In that case, it may be redundant and computationally expensive to augment the prompt with additional exemplar input / output pairs. However, the LLM may not be very successful in generating a correct control action for more complex user commands such as “create a romantic atmosphere in my living room” that involve more complex reasoning steps and a plurality of actuators. The optimal lighting for a romantic atmosphere is very different if the user is watching a romantic movie compared to having a date with another person. In the latter case, the system advantageously augments the prompt with suitable exemplar input / output pairs such that a correct set of instructions are generated by the LLM.
[0018] The method may further comprise identifying the one or more exemplar pairs to augment said prompt from a set of exemplar pairs stored in a memory by determining a sematic similarity metric between the textual user input and the exemplar task-specific textual user input from each exemplar pair from the set of exemplar pairs and identifying the exemplar pairs for which the sematic similarity metric is above a similarity threshold. As mentioned above, each exemplar pair comprises an exemplary task-specific textual user input and a corresponding language model output. The method may comprise determining a semantic similarity metric between the textual user input and each of the exemplary taskspecific textual user inputs from the exemplar pairs available in memory. The augmented prompt comprises said textual user input and the exemplar pairs for which the similarity metric is above the similarity threshold. The identified exemplar pairs may be inserted in descending order based on their corresponding semantic similarity metric, such that an exemplar with a higher semantic similarity metric value is positioned before an exemplar with a lower semantic similarity metric value.
[0019] The similarity threshold may be based on the computational resources and / or accuracy of the LLM and / or a safety level of the control output according to the LLM (the lighting task at hand). For example, LLMs with high accuracy and reasoning capabilities (which typically require more computational resources to generate an output) may be able to provide a satisfactory inference even if the user' s new prompt differs substantially from the exemplar provided along the user' s input. On the other hand, an LLM with low accuracy and less developed reasoning capabilities may be only able to provide a satisfactory inference only if the user input is very similar with the exemplar. Similarly, the similarity threshold may be based on the criticality (safety level) of the lighting task at hand. For instance, for a user input leading to a change in the functional lighting, a very high similarity threshold may be required as a change in functional lighting may result in unsafe situations (e.g., switching off the light while an elderly person is trying to go to the bathroom and thereby causing a fall). On the other hand, for a user input leading to a change in the ambient lighting, a low similarity threshold may be used as ambient lighting has a purely aesthetic role with no consequences on safety.
[0020] The method may further comprise inferring a performance gain for each exemplar pair from a set of exemplar pairs stored in a memory by applying a second machine-learning model on the textual user input augmented with each exemplar pair. The second machine-learning model may have been pre-trained based on labeled past textual user inputs and corresponding LLM outputs, said past textual user inputs being associated with: a negative score, if an active response of the user has been received in response to the corresponding LLM past output; wherein an active response comprises receiving an input from the user, said input for controlling the lighting system; a positive score, if said active response is not received. A negative score if the actuation resulted in the user stopping or changing his activity (e.g., user stopped knitting or moved from the couch to the dining table to have enough light for knitting). That is, the past inputs are labeled with a negative score if an active response from the user is detected (user intervened to change the output of the control system), and a positive score if no active response is detected. The method may further comprise identifying the one or more exemplar pairs for which the performance gain is above a performance threshold to augment the prompt. In other words, the performance gain for each exemplar pair is inferred by applying the second machine-learning model to the textual user input augmented with each exemplar pair, i.e., the input to the second machine learning model is textual user input concatenated with an exemplar pair. The exemplar pairs that exceed a performance threshold are then identified and used to augment the prompt (augmented prompt comprises the textual user input appended after the one or more identified exemplar pairs). The one or more identified exemplar pairs may be ranked based on the performance gain. In that case, the augmented prompt may comprise the one or more identified exemplar pairs inserted in descending order based on their corresponding performance gain followed by the textual user input. This is beneficial because the order in which the exemplars are presented can impact the effectiveness of the prompt. Thus, more relevant (effective) exemplars are presented first.
[0021] The method may further comprise identifying the one or more exemplar pairs to augment said prompt from a set of exemplar pairs stored in a memory by determining a level of complexity for each exemplar pair from said set of exemplar pairs, and identifying the one or more exemplar pairs for which the level of complexity is above a performance threshold. The level of complexity may be determined by analyzing said exemplar pair to determine a characteristic of said exemplar pair, for example to determine a number of characters of said exemplar pair (number of characters in past textual user input and corresponding past LLM output).
[0022] The exemplar pairs may comprise a number of intermediate (logical) reasoning steps towards outputting the language model output to the exemplary task-specific textual user input. That is, each exemplar pair comprises the exemplary task-specific textual user input followed by the set of logical intermediate steps or reasoning pathways that are used by the LLM to generate the corresponding language model output, followed by the corresponding language model output. For example, the exemplar pair may comprise [Q: “Is it best for my 85-year old dad to switch all lights off during the night?”, Intermediate steps: “It is common to sleep during the night”, “A person may wake-up during the night”, “Your father is 85-years old, thus, elderly”, “A low light level will prevent an elderly person from falling”, A: “A low light level is advisable during the night, for instance using the antistumble lighting fixture”]. The method may comprise determining a number of intermediate reasoning steps for each exemplar pair and determining the level of complexity for each exemplar pair based on the number of intermediate reasoning steps.
[0023] It should be understood that the system, method and computer program product may have similar and / or identical embodiments and advantages as the above- mentioned lighting devices.
[0024] BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The above, as well as additional objects, features and advantages of the disclosed systems, devices and methods will be better understood through the following illustrative and non-limiting detailed description of embodiments of devices and methods, with reference to the appended drawings, in which:
[0026] Fig. 1 shows schematically an example of a system for conditionally transmitting a prompt as input to a large language model;
[0027] Fig. 2 shows schematically a method for conditionally transmitting a prompt as input to a large language model;
[0028] Fig. 3 shows schematically a method for conditionally augmenting a prompt as input to a large language model. All the figures are schematic, not necessarily to scale, and generally only show parts which are necessary in order to elucidate the invention, wherein other parts may be omitted or merely suggested.
[0029] DETAILED DESCRIPTION
[0030] Fig. 1 shows an example of a system 100 for conditionally transmitting a prompt 450 as input to a large language model 550 (not part of the system 100). The system comprises an input interface 102, an output interface 104. The system 100 further comprises at least one data processor or controller 106. The system may further comprise at least one data repository or storage or memory 108 for storing computer program code instructions. The controller 106 may be communicatively coupled to the cloud. However, it should be noted that the controller 106 may be comprised in a central device (e.g., a smartphone, personal computer, a hub, a voice control system), a web-based portal, part of the lighting system, etc. The controller 106 may be communicatively coupled to the memory module 108. The input interface 102 is configured to receive an input 350 from the user. Said input may be a textual user input. Alternatively, the input interface 102 may be configured to receive a voice / audio input. For example, the system 100 may comprise a microphone for receiving sounds made by the user. Said sound input may be processed by the controller 106 to generate a textual user input 350 associated with said audio input for example by using Automatic Speech Recognition (ASR) or Speech-to-Text (STT) techniques. The controller 106 may be configured to receive a textual user input 350, via the input interface 102, indicative of an intention to control one or more devices 140, 142 of a lighting system 150. For example, the controller 106 receives a textual user input "turn on the lights" indicative of the intention of the user to control the device 140. The controller 106 is further configured to analyze said textual user input 350 to determine at least one characteristic of said textual user input and determine a level of complexity of said textual user input based on said at least one characteristic.
[0031] In an example, the controller 106 may analyze the textual user input 350 to determine a readability level of said textual user input and determine the level of complexity based on said readability level. Different formulas may be used to assess readability ranging from simple metrics such as word frequency count, number of characters in the textual user input, number of different words in the textual user input, percentage of unique words, number of prepositional phrases, etc. to more complex formulas such as the Flesh Formulas, the Dale-Chall Formula etc. Several models and formulas are available in the state of art for determining the readability of a text and as such will not be further discussed in the context of this application.
[0032] In another example, the controller 106 may analyze said textual user input 350 to determine a number of tasks that need to be completed to carry-out said control of the one or more devices 140, 142 of the lighting system 150 and determine the level of complexity based on said number. For example, for a textual user input such as “As a family of four, can you adjust the lights in our living room”, the controller 106 may determine that the (sub)tasks of: 1) determining activities of users in the living room 2) generate a set of lighting parameters for each activity 3) prioritize activities based on lighting needs 4) determine location of users with respect to lighting devices 5) adjust lighting parameters based on ranking. The controller 106 may determine the number of tasks required by using a machine learning model, such as a natural language processing model. The controller 106 may determine the level of complexity based on said number of tasks such that a higher level of complexity is assigned to a textual user input with a higher number of tasks.
[0033] In yet another example, the controller 106 may apply a first machine learning model to said textual user input to determine an effectiveness score for said user input. The first machine learning model may have been pre-trained based on labeled past textual user inputs and corresponding past large language model outputs (past interaction between the user and the LLM). A past interaction comprises a past user input transmitted as an input prompt to the LLM, the LLM generated a control output for the lighting system 150 corresponding to a lighting setting A, the processor 106 received the control output and controlled the lighting system according to the lighting setting A (past LLM output). The past textual user inputs may have been associated (labeled) with a negative score if an active response of the user has been received in response to the corresponding large language model past output. Here, an active response may comprise receiving an input from the user, said input for controlling the lighting system. For example, if during a past interaction between the user and the LLM, the user responded to said lighting setting A by controlling the lighting system according to a different lighting setting B (the user intervened to change the control output generated by the LLM), the past textual user input is labeled as negative (associated with a negative score). An active response may comprise determining a change in physiological and / or psychological parameters of the user indicative of a negative emotion, for example facial expression indicative of negative emotion, brain-wave activity indicative of negative emotion, hand movement / gestures, etc. On the contrary, if an active response is not received, that is, the user did not intervene to change the output of the lighting system corresponding to the control output of the LLM, the past textual user input is labeled as positive (associated with a positive score). In another example, the past textual user inputs may have been associated (labeled) with a further negative score if the controller 106 has determined that the corresponding past control output of the LLM resulted in a negative and / or harmful experience for the user, e.g., delay in the circadian rhythm due to lighting. If the user's input is similar to past inputs that resulted in a positive response (i.e., the lights turned on without the user intervening to the control output), the model would assign a positive score (high effectiveness score) to the current input. If the user's input is similar to past inputs that resulted in a negative response (i.e., the lights didn't turn on or there was an issue with the lighting system), the model would assign a negative score (low effectiveness score). The processor 106 may determine the level of complexity based on the effectiveness score, that is, the processor 106 may assign a high level of complexity to textual user inputs associated with a low effectiveness score and a low level of complexity to textual user inputs associated with a high effectiveness score. That is, the more effective inputs (high effectiveness score) are less complex for the LLM to handle.
[0034] In a further example, the controller 106 may analyze said textual user input 350 to determine a number of keywords, from a set of predetermined keywords, present in said textual user input and determine the level of complexity based on said identified (determined by the controller 106) number of keywords. For example, for a lighting system, the set of predetermined keywords may comprise [“turn”, “switch”, “on”, “off’, “lights”, “scene”, “meditation”, “romantic”, “relaxing”, “atmosphere”], or combinations of keywords such as “switch”+ “on”, “switch”+ “off’, “scene”+ “romantic”, etc. In an example, the textual user input may be “turn on the lights” and the controller 106 may analyze said textual user input to identify 3 keywords. A higher level of complexity may be assigned to textual user inputs with a smaller number of identified keywords (or a smaller ratio of identified keywords relative to the length of the textual user input).
[0035] The controller 106 is further configured to, if said level of complexity is above a first complexity threshold, generate a prompt, said prompt comprising at least said textual user input, and transmit said textual user input 350 as a prompt 450 to said large language model (for example LLM may run on the cloud), receive the output 500 of said large language model, via the input interface 102, and control said lighting system 150 according to said output. For example, the textual user input may be, “I am feeling like meditating, create a forest experience”, the controller 106 may analyze said textual user input to determine at least one characteristic of said textual user input, say a number of words, here 9, and determine a level of complexity based on said number. If said level of complexity is above a first complexity threshold, say threshold is 4, the controller 106 generates a prompt 450 that contains at least said textual user input 350 and optionally further information such as lighting system characteristics (e.g., lighting device type, location, lighting interface type, etc.), desired format output, e.g. JSON format, relevant sensor output, e.g., temperature data in the room, etc., and transmits said prompt to the LLM. For example, the generated prompt 450 may be:
[0036] {“textual user input”: “I am feeling like meditating, create a forest experience”,
[0037] “Lighting fixtures”: [
[0038] {“type”: “celling light”,
[0039] “location”: “bedroom”,
[0040] “interface”: “DALI”,
[0041] “params”: “level and color”},
[0042] {“type”: “down light”,
[0043] “location”: “bedroom”,
[0044] “interface”: “DALI”, “params”: “level”}],
[0045] “Sensor information”: “temperature”: “23 degrees”}
[0046] The LLM processes said prompt and generates an output 500 for controlling the actuation system. For example, the LLM output 500 may be in the form of: {"action": "command",
[0047] "location": "bedroom",
[0048] "target": "down light",
[0049] "color temperature value": 122,
[0050] "brightness value": 60,
[0051] "target": "thermostat",
[0052] "temperature value": 21 } The controller 106 may receive the output 500 of said large language model, via the input interface 102, and control said lighting system 150 according to said output.
[0053] If said level of complexity is below the first complexity threshold, the controller 106 does not generate a prompt for the large language model. The controller 106 may process said (simple) textual user input 350 and directly control said lighting system 150 accordingly. The controller may process the textual user input via a local LLM, for example a low parameter less accurate LLM. The system 100 may further comprise a speaker for providing an auditory response to the user.
[0054] Fig. 2 shows a method 200 of conditionally transmitting a prompt as input to a (remote) large language model. The method comprises the steps of:
[0055] - receiving (202) a textual user input indicative of an intention to control one or more devices of the lighting system;
[0056] - analyzing (204) said textual user input to determine at least one characteristic of said textual user input;
[0057] - determining (206) a level of complexity of said textual user input based on said at least one characteristic;
[0058] - if said level of complexity is above a first complexity threshold, generating (207) a prompt, said prompt comprising at least said textual user input, and transmitting (208) said textual user input as a prompt to said large language model, receiving (210) the output of said large language model and controlling (214) said lighting system according to said output;
[0059] - if said level of complexity is below the first complexity threshold, the method may (optionally) comprise further analyzing (212) said textual user input to determine one or more control parameters and controlling (214) said lighting system according to said control parameters. Said further analysis may comprise analyzing the textual user input via an local (locally run on the processor 106) LLM.
[0060] The method 200 may be executed by computer program code of a computer program product when the computer program product is run on a processing unit of a computing device, such as the controller 106.
[0061] A second embodiment of conditionally transmitting a prompt as input to a large language model is shown in Fig. 3. This embodiment is an extension of first embodiment of Fig. 2. In the embodiment of Fig. 3, steps 308, or 310 and 314 may be optionally performed in between steps 206 and 208 of Fig. 2. That is, the method of Fig. 3 comprises: - receiving 202, via an input interface 102, a textual user input 350 indicative of an intention to control one or more devices 140, 142 of a lighting system 150;
[0062] - analyzing 204 by the controller 106 said textual user input 350 to determine at least one characteristic of said textual user input. For example, the at least one characteristic may be a number of characters of said textual user input, an effectiveness score determined by applying the first machine learning model to the textual user input, a number of (sub) tasks that need to be completed to carry-out said control of the one or more devices of the lighting system, etc.
[0063] - determining 206 a level of complexity of said textual user input based on said at least one characteristic (or a combination of characteristics). For example, the controller 106 may determine a number of characters and an effectiveness score of said textual user input. The level of complexity may be based on the combination of these characteristics, for example an algorithm such as a decision tree or random forest may be used to determine the level of complexity given the values of the one or more characteristics, i.e., the length of the text and effectiveness in this example. If said level of complexity is below a first complexity threshold, the method comprises generating a prompt 207, said prompt comprising said textual user input and optionally further information related to the control of the actuation system, for example control parameters of the actuation system, output format of the LLM output, etc., and transmitting 208 said prompt to said large language model via an output interface 104.
[0064] In step 314, the method 300, comprises, if said level of complexity is above a second complexity threshold (the second threshold is larger than the first threshold), generating 314 an augmented prompt, wherein said augmented prompt comprises said textual user input 350 and one or more exemplar pairs 800 (for example the textual user input is appended after the exemplar pairs). The system 100 may comprise a memory 108 comprising a set of exemplar pairs 850. Each exemplar pair comprises an exemplary task-specific textual user input and a corresponding language model output. That is, each exemplar pairs comprises a textual user input and a corresponding language model output (exemplary conversation between the system 100 and the large language model 550) that is relevant for the control of one or more devices of the actuation system. In that context, a task refers to an action or set of actions related to the control of the actuation (lighting) system, such as (dynamically) turning the lighting devices on or off, adjusting brightness or color temperature or the flicker index or color point of the devices, setting a light scene, etc. In an example, the set of exemplar pairs may be previously processed past textual user inputs and corresponding LLM outputs. Similarly, exemplar pairs may be automatically created from a user manual of a lighting system (e.g. a detailed description of a predefined lighting scene to be selected for a specific user activity e.g. reading, relax, energize) or from movies / promotion videos on YouTube showing a change in the activity of a user in a room followed by a subsequent change (e.g. a church service showing a transition from praise and worship activity to a sermon).
[0065] The processor 106 may be configured to identify 310 the one or more exemplar pairs 800 to augment said prompt from a set of exemplar pairs 850 stored in memory 108 by determining a sematic similarity metric between the textual user input and the exemplar task-specific textual user input from each exemplar pair from the set of exemplar pairs. The processor 106 may identify (select) the exemplar pairs for which the similarity metric value is above a similarity threshold. In an example, the controller 106 may select a number of most similar example pairs (ranked by their corresponding similarity metric values). The exemplar number may be pre-determined, say 3 most similar pairs. Additionally, and / or alternatively, a maximum number of tokens allowed as input to the LLM 550 may limit the exemplar number. The augmented prompt may comprise the one or more exemplar pairs in descending order according to their corresponding similarity metric followed by the textual user input. In an example, the controller 106 may request the (pretrained) LLM 550 to determine the semantic similarity between the exemplar pairs and the textual user input. For example, by generating a prompt containing the textual user input and the set of exemplar pairs followed by a question to determine semantic similarity. In another example, two or more identified exemplar pairs may be fused into a single exemplar pair. A first exemplar pair may describe the getting out of bed routine of an elderly and the lighting choice for each of the sub activities. A second exemplar may describe lighting changes during a meal eating routine. The controller 106 may request the (pre-trained) LLM 550 to generate a single exemplar pair for a dynamic lighting recipe of the elderly getting out of bed and to the table to have dinner.
[0066] The similarity threshold may be determined based on the accuracy of the LLM. That is, an LLM with a many parameters and high accuracy would perform well with simple exemplars while a less accurate LLM needs high relevant exemplars.
[0067] In an example, the processor 106 may be configured to apply a second machine-learning model on the textual user input augmented with each exemplar pair to determine (infer) a performance gain for each exemplar pair. The second machine-learning model may have been pre-trained based on labeled past textual user inputs and corresponding large language model outputs, said past textual user inputs being associated with: a negative score, if an active response of the user has been received in response to the corresponding large language model past output; wherein an active response comprises receiving an input from the user, said input for controlling the lighting system; or a positive score, if said active response is not received. That is, if the user intervened to change the control output of the lighting system (corresponding to the LLM output), the controller 106 assigns a negative score (label) to the corresponding textual user input whereas if the user did not respond to the lighting control output, the controller 106 assigns a positive score (label) to the corresponding textual user input. The controller 106 may be configured to identify 310 the one or more exemplar pairs 800 to augment said prompt from the set of exemplar pairs 850 stored in memory 106 by selecting the one or more exemplar pairs 800 for which the performance gain is below a threshold. The one or more exemplar pairs may be ranked in terms of their corresponding performance gain and the controller 106 may be configured to select a (predetermined) number of exemplar pairs. Additionally, and / or alternatively, a maximum number of tokens allowed as input to the LLM 500 may limit the predetermined exemplar number. The augmented prompt comprises the one or more identified exemplar pairs inserted in descending order based on their corresponding performance gain followed by the textual user input.
[0068] The processor 106 may be configured to analyze each of said exemplar pairs from the set of exemplar pairs stored in the memory 108 to determine a level of complexity for each exemplar pair. For example, said level of complexity may be determined based on a number of characters in the exemplar pair.
[0069] Additionally, and / or alternatively, the exemplar pairs may comprise a number of intermediate reasoning steps towards outputting the language model output to the exemplary task-specific textual user input. For example, an exemplar pair may be in the form of: [Q: “Maria is reading, and Spiros is relaxing on the sofa. What is the best light setting?”, “Maria is performing a higher acuity activity”, “Maria has priority over Spiros”, “Maria is sitting in the dining table”, “The best lighting setting for Maria is to increase brightness in the dining table”, “Spiros is sitting in the living room”, “Given the lighting setting for Maria, adjust the lighting of Spiros in the dining table”, LLM output: “{"action": "command", "location": "dining room", "brightness value" : 90, "location": "living room", "brightness value": 40}”]. The controller 106 may be configured to analyze each exemplar pair to determine the number of intermediate reasoning steps towards outputting the language model output to the exemplary task-specific textual user input, that is, the number of steps besides the user query and the final LLM output. In the example above, 4. The controller 106 may determine the level of complexity of the exemplar based on the number of intermediate reasoning steps. The controller 106 may identify (select) the one or more exemplar pairs to augment the prompt for which the level of complexity is above an exemplar complexity threshold.
[0070] The method 300 further comprises transmitting 208 said augmented prompt to said large language model 550 via the output interface 104.
[0071] Step 308 is similar to step 207. The method 300 comprises, if said level of complexity is below the second complexity threshold (and above the first complexity threshold), generating 308 a prompt, said prompt comprising at least said textual user input and optional further information relating to the actuation system, and transmitting 208 said prompt to said large language model via the output interface 104.
[0072] The method 300 may further comprise receiving 210 the LLM output, via the input interface 102, and controlling 212 the lighting system 150 according to said LLM output. The LLM output may comprise control parameters for controlling the one or more devices 140, 142 of the lighting system 150.
[0073] It should be noted that the above-mentioned embodiments illustrate rather than limit the invention, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims.
[0074] In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. Use of the verb "comprise" and its conjugations does not exclude the presence of elements or steps other than those stated in a claim. The article "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer or processing unit. In the device claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0075] Aspects of the invention may be implemented in a computer program product, which may be a collection of computer program instructions stored on a computer readable storage device which may be executed by a computer. The instructions of the present invention may be in any interpretable or executable code mechanism, including but not limited to scripts, interpretable programs, dynamic link libraries (DLLs) or Java classes. The instructions can be provided as complete executable programs, partial executable programs, as modifications to existing programs (e.g. updates) or extensions for existing programs (e.g. plugins). Moreover, parts of the processing of the present invention may be distributed over multiple computers or processors or even the ‘cloud’. Storage media suitable for storing computer program instructions include all forms of nonvolatile memory, including but not limited to EPROM, EEPROM and flash memory devices, magnetic disks such as the internal and external hard disk drives, removable disks and CD-ROM disks. The computer program product may be distributed on such a storage medium, or may be offered for download through HTTP, FTP, email or through a server connected to a network such as the Internet.
Claims
CLAIMS:
1. A method for conditionally transmitting a prompt as input to a large language model, and for using an output of said large language model to control a lighting system, said method comprising:- receiving a textual user input indicative of an intention to control one or more devices of the lighting system;- analyzing said textual user input to determine at least one characteristic of said textual user input;- determining a level of complexity of said textual user input based on said at least one characteristic, and- only if said level of complexity is above a first complexity threshold, generating a prompt, said prompt comprising at least said textual user input,- transmitting said prompt to said large language model,- receiving the output of said large language model and controlling said lighting system according to said output.
2. The method according to claim 1, wherein said at least one characteristic is a readability level of said textual user input.
3. The method according to any preceding claim, wherein said at least one characteristic is a number of tasks that need to be completed to carry-out said control of the one or more devices of the lighting system.
4. The method according to any preceding claim, wherein said step of analyzing said textual user input comprises: applying a first machine learning model on the textual user input to determine an effectiveness score for said user input and wherein said at least one characteristic is said effectiveness score,wherein said first machine learning model has been pre-trained based on labeled past textual user inputs and corresponding past large language model outputs, said past textual user inputs being associated with: a negative score, if an active response of the user has been received in response to the corresponding large language model past output; a positive score, if said active response is not received.
5. The method according to any preceding claim wherein said at least one characteristic is a number of keywords, from a set of predetermined keywords, present in said textual user input.
6. The method according to any preceding claim, said method comprising: if said level of complexity is above a second complexity threshold, said second complexity threshold being larger than the first complexity threshold, generating an augmented prompt, wherein said augmented prompt comprises at least said textual user input and one or more exemplar pairs, and transmitting said augmented prompt to said large language model, wherein each exemplar pair comprises an exemplary task-specific textual user input and a corresponding language model output.
7. The method according to claim 6, wherein the method further comprises identifying the one or more exemplar pairs to augment said prompt from a set of exemplar pairs stored in a memory by: determining a sematic similarity metric between the textual user input and the exemplary task-specific textual user input from each exemplar pair from the set of exemplar pairs; identifying the one or more exemplar pairs for which the sematic similarity metric is above a similarity threshold.
8. The method according to claim 7, wherein said similarity threshold is based on an accuracy of the large language model and / or a safety level of the control output of said large language model.
9. The method according to claim 6, wherein the method further comprises identifying the one or more exemplar pairs to augment said prompt from a set of exemplar pairs stored in a memory by: inferring a performance gain for each exemplar pair from said set of exemplar pairs by applying a second machine-learning model on the textual user input augmented with each exemplar pair, wherein said second machine-learning model has been pre-trained based on labeled past textual user inputs and corresponding large language model outputs, said past textual user inputs being associated with: a negative score, if an active response of the user has been received in response to the corresponding large language model past output; a positive score, if said active response is not received. identifying the one or more exemplar pairs for which the performance gain is above a performance threshold.
10. The method according to claim 7 or 9, wherein the method further comprises ranking the one or more identified exemplar pairs based on the performance gain or the sematic similarity metric and wherein said augmented prompt comprises the one or more identified exemplar pairs inserted in descending order based on their corresponding performance gain or sematic similarity metric followed by the textual user input.
11. The method according to claim 6, wherein said method further comprises identifying the one or more exemplar pairs to augment said prompt from a set of exemplar pairs stored in a memory by: determining a level of complexity for each exemplar pair from said set of exemplar pairs, and identifying the one or more exemplar pairs for which the level of complexity is above an exemplar complexity threshold.
12. The method according to claim 11, wherein said exemplar pairs comprise a number of intermediate reasoning steps towards outputting the language model output to the exemplary task-specific textual user input and wherein said step of determining the level of complexity for each exemplar pair comprises determining the level of complexity based on the number of intermediate reasoning steps.
13. A controller for conditionally transmitting a prompt as input to a large language model, and for using an output of said large language model to control a lighting system, said processor configured to:- receive a textual user input, via an input interface, indicative of an intention to control one or more devices of the lighting system;- analyze said textual user input to determine at least one characteristic of said textual user input;- determine a level of complexity of said textual user input based on said at least one characteristic;- only if said level of complexity is above a complexity threshold, transmit, via an output interface, said textual user input as a prompt to said large language model, receive the output of said large language model and control said lighting system according to said output.
14. A system for conditionally transmitting a prompt as input to a large language model, said system comprising: an input interface, an output interface; at least one controller according to claim 13.
15. A computer program product for a computing device such as the controller of claim 13, the computer program product comprising computer program code to perform the method of claims 1-12 when the computer program product is run on a processing unit of the computing device.
Citation Information
Cited By
Natural language input for controlling a lighting system
GB2701700A
Device, method and computer program product for generating light effect
US12538407B2
Device, method and computer program product for generating light effect
US20250301553A1
Cybersecurity Command Line Assessment
US20250373642A1