Computer implemented method and controller for controlling implementation of a lighting configuration task

A computer-implemented method and controller use multimodal models to process natural language and image data to efficiently decompose lighting configuration tasks into sub-tasks, addressing the inefficiencies of existing methods by reducing labor and cost, and enabling adaptable lighting solutions.

WO2025252439A1PCT designated stage Publication Date: 2025-12-11SIGNIFY HOLDING BV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/063658
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-17
Filing Date
2025-05-19
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Existing methods for automating lighting configuration tasks in smart homes and commercial buildings are labor-intensive, require programming expertise, and are economically unfeasible for complex but limited-value tasks due to the need for well-curated datasets and specialized commissioning flows.

Method used

A computer-implemented method and controller that utilize a multimodal large language model to process natural language instructions and image data to determine sub-tasks for lighting configuration, generating output data that includes control instructions for lighting units, leveraging off-the-shelf models and specialized modules to decompose complex tasks into simpler steps.

Benefits of technology

Enables efficient and customizable lighting configuration by breaking down tasks into manageable sub-tasks, reducing labor and cost, and allowing for rapid adaptation to diverse lighting needs without extensive retraining or dataset curation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025063658_11122025_PF_FP_ABST
    Figure EP2025063658_11122025_PF_FP_ABST
Patent Text Reader

Abstract

The invention is directed to a computer implemented method for controlling implementation of a lighting configuration task (LT) for one or more lighting units (11, 12) of a lighting arrangement (10), that comprises receiving a first user input (104) e.g., a natural language instruction (105), indicative of the lighting configuration task, receiving a second input (108) comprising image data (110), indicative of an environment (112) wherein the lighting configuration task is to be implemented, determining, using the first user input and the image data received, one or more sub-tasks (116.1, 116.2,… 116.N) that are associated with the implementation of the lighting configuration task in the environment; and generating, based on the first input, output data (118), e.g., a textual description, indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment. The method enables a composition of specialized models for lighting configuration task.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Computer implemented method and controller for controlling implementation of a lighting configuration task

[0002] FIELD OF THE INVENTION

[0003] The invention is directed to a computer implemented method for controlling implementation of a lighting configuration task for one or more lighting units of a lighting arrangement. The invention is also directed to a controller for implementing a lighting configuration task and to a computer program.

[0004] BACKGROUND OF THE INVENTION

[0005] While prior art vision models typically solve a well-defined but rather narrow problem, the machine learning tasks associated to the real world of lighting / smart home are often broader and loosely defined, due to the diversity of the lighting world. Manually creating these customized machine learning programs for automating each one of the complex tasks, which are encountered regularly in a smart home / commercial building or during the commissioning phase, not only requires programming expertise but is also slow, labor intensive, and ultimately insufficient to cover the space of all desired lighting tasks with their own custom machine learning (ML) applications.

[0006] In current state of art, the predominant approach to building specialized arterial intelligence (Al) systems has been massive-scale unsupervised pre-training followed by supervised multitask training. Similar to what happened with GPT-4 / ChatGPT for language, what is now needed to build a multi-task application for vision is a Vision Transformer model (which is pre-trained on over 100 million image examples from the internet) and combine it with a small number of new example images specific to your industrial application, resulting in a custom ML application. However, this approach requires a well curated dataset for each new task that makes it challenging to scale to the large amount of complex-but-limited-economic-value tasks of ML systems for the lighting / smart-home domain. Similarly, the development of tooling for specialized commissioning flows required for a certain application currently is economically unfeasible, leaving those applications without connected systems. SUMMARY OF THE INVENTION

[0007] It is therefore an object of the present invention to enable a composition of specialized models for lighting configuration tasks that tackles the above mentioned issues.

[0008] A first aspect of the present invention is formed by a computer implemented method for controlling implementation of a lighting configuration task for one or more lighting units of a lighting arrangement. Lighting configuration task refers in general to tasks pertaining to the operation or control of the lighting units of the lighting arrangement, and that include, among others, configuration, commissioning and / or operation of the lighting units, either for the direct operation of the lighting units or for a simulation of said operation. The method comprises:

[0009] - receiving a first user input indicative of the lighting configuration task;

[0010] - receiving a second input comprising image data, indicative of an environment wherein the lighting configuration task is to be implemented;

[0011] - determining, using the first user input and the image data received, one or more sub-tasks that are associated with the implementation of the lighting configuration task in the environment; and

[0012] - generating, based on the first user input, output data indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment.

[0013] Lighting arrangements comprise a plurality of lighting units that are installed in a given environment such as a room, a building, a parking lot, a street or group of streets, etc.

[0014] Lighting configuration task refers in general to tasks pertaining to the operation or control of the lighting units of the lighting arrangement, and that include, among others, configuration, commissioning and / or operation of the lighting units, either for the direct operation of the lighting units or for a simulation of said operation.

[0015] The image data shows or represents at least part of the environment where the lighting configuration task should be implemented. The environment can be, as indicated above, a representation of a room, a building or part thereof, such as an aisle in a warehouse, an outdoor environment such as a street or a parking lot. The image data can be image data directly obtained from the environment or a schematic representation thereof, such as a floor plan of the environment. Using the image data and the first user input the method determines one or more sub-tasks that are associated with the implementation of the lighting configuration task in the environment, and subsequently, generates, based on the first user input, output data indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment.

[0016] The sub-tasks correspond to the steps that would be required to implement the lighting configuration task indicated in the first user input.

[0017] The output file indicative of the one or more subtasks to be executed or carried out forms a suitable guide that can be used for controlling the implementation of the lighting configuration task for the one or more lighting units of the lighting arrangement in the given environment.

[0018] In the following, embodiments of the method of the first aspect of the invention will be described.

[0019] Preferably, the first user input comprises a natural language instruction. The instruction received via the first user input is a natural language instruction. A natural language instruction is an instruction provided in natural language. Natural language, or ordinary language, is a language that occurs naturally in a human community by a process of use, repetition, and change without conscious planning or premeditation and are distinguished from constructed or formal languages, such as those used to program computers or study logic. Additionally, or alternatively, the first user input may also include other data file types, such as an image file or an audio file, that indicate the lighting configuration task. For instance, in the case of multi-modal large language models (LLM), such as for instance Google’s Gemini and Open Al’s ChatGPT, that accept text files as well as other types of files such as image or audio files, some LLM are advantageously configured to process visual representations of images or video, or sound representations in audio files, to obtain relevant information that can be indicative of a lighting configuration task. For instance, a vision system can be trained (e.g., supervised, semi-supervised, selfsupervised) to turn an input image into a high-level representation. This high level representation of the input image is basically a list of tokens that are similar to the kind of tokens that a LLM takes as an input. The tokens are then feed that to the LLM in addition to the text, and, during training, the LMM is able to use the high level representations for determining the one or more sub-tasks. The same concept can be applied for audio input as first user input.

[0020] Thus, in an embodiment, the first user input comprises text, such as a text file. Additionally, or alternatively, the fist user input may comprise an audio signal, such as an audio file. Also additionally, or alternatively, the first user input may comprise an image, such as an image file, from which a text indicative of the natural language instruction can be extracted. Also additionally, or alternatively, the first user input may comprise a video, such as a video file, from which a text indicative of the natural language instruction can be extracted, or from which an audio signal indicative of the natural language instruction can be extracted.

[0021] Preferably, in an embodiment, the determination of the one or more subtasks is performed by a first language model with the capability of analyzing image data, e.g., in the form of an image or a sequence of images, such as a video file. A non-limiting example of a suitable language model is VisProg.

[0022] In a preferred embodiment, the output data comprises human-readable or human-understandable text, for instance a textual description of the one or more sub-tasks. Alternatively, or additionally, the output data can include one or more tokens indicative of a LLM-readable representation of information that are similar to the kind of tokens that a multimodal LLM takes as an input (e.g. when ingesting an image). Tokens can be in the form of non human-understandable text strings, that are understandable by a given large language model, as it will be explained with more detail below.

[0023] In a particular embodiment, the method further includes:

[0024] - identifying, or otherwise determining, for the one or more sub-tasks, module data indicative of corresponding relevant software modules that are associated with the respective sub-tasks. In this embodiment, the output data (e.g., the textual description and / or the output tokens) provided by the method is further indicative of the identified module data. Relevant software modules are preferably identified or selected from a pool or library of available software modules. The relevant software modules can include, for example, mathematical modules or models pertaining to the implementation of lighting functions of the lighting units.

[0025] In another embodiment, the method additionally or alternatively comprises the step of identifying, or otherwise determining, for the one or more lighting units, lighting unit data indicative of operation parameters and / or capabilities of the respective lighting units. The respective lighting units are, for instance, the lighting units in the environment and / or the lighting units otherwise referenced in the natural language instruction. In this embodiment, the output data (e.g., the textual description and / or the output tokens) provided is further indicative of the identified lighting unit data. The lighting unit data is thus associated to the lighting units and indicative of the parameters of the lighting unit and / or parameters under which they can be operated. Lighting data unit may include data pertaining to a dimension or size of the corresponding lighting unit, possible mounting positions, type of illumination that can be provided by the respective lighting unit, for instance in terms of intensity, temperature, color, direction of illumination, directivity, including fixed values and / or possible ranges.

[0026] In another embodiment, which may include any of the technical features described above with reference to other embodiments, the image data comprises an image file or a video file, or a depth map or a floor plan drawing, or a three-dimensional representation of a built space. For instance, depth maps can be generated by LiDAR, be a stereoscopic camera system, or by applying Al to 2D-images. Also, a neural radiance field (NeRF) can be used, which is created by a deep learning model from sparse 2D-images. Such a NeRF model of the environment can be subsequently used to generate, for instance, novel view synthesis, scene geometry and reflectance properties of the scene or space. The image data can be directly provided from a memory device as a stored file, or provided as a link to a file, or provided directly from a camera.

[0027] In yet another embodiment, the method further includes the step of:

[0028] - generating a multimodal prompt including the image data, the first user input (e.g., the natural language instruction), and the output data (e.g., the textual description and / or the output tokens) indicative of the sub-tasks, and providing said multimodal prompt to a natural language model.

[0029] Preferably, when the method includes identifying the module data indicative of corresponding relevant software modules that are associated with the respective subtasks and / or lighting unit data indicative of operation parameters and / or capabilities of the respective lighting units, the multimodal prompt further includes the identified module data and / or the identified lighting unit data. The multimodal prompt is thus an augmented prompt that includes the relevant information generated by the method, for instance as pre-pended data, as postpended data or as data otherwise integrated in the augmented prompt. The multimodal prompt is then provided to a natural language model, which can be the same language model used to determine the one or more sub-tasks or a different language model.

[0030] In another embodiment, which includes identifying the module data indicative of corresponding relevant software modules that are associated with the respective sub-tasks, the method further comprises:

[0031] - selecting, from a predetermined set of available example data, target example data indicative of examples of use of the relevant modules; and

[0032] - generating the multimodal prompt further including the target example data.

[0033] The target example data includes one or more examples of how the respective relevant module can be used. These examples can be than pre-pended to the multimodal prompt provided to the natural language model.

[0034] Alternatively, or additionally, in another embodiment, the method further comprises:

[0035] - generating the multimodal prompt further including a high level textual description of one or more of the relevant modules.

[0036] In this particular embodiment, the high level description, or high level definition, of at least one of the relevant modules is also included, e.g. pre-pended, postpended or otherwise included, to the multimodal prompt that is provided to the natural language model.

[0037] In an exemplary embodiment, the high level textual description of a module comprises a description of input parameters required by the respective module and a description of module operation for obtaining an output parameter based on said input parameter. The high level textual description of the module may additionally or alternatively comprise examples of suitable and / or unsuitable use of the respective module. For instance, the description of operation of the module can include a first positive example of using the module such a positive result is achieved. It may also include a second, negative example to warn the natural language model (e.g., a LLM receiving the multimodal prompt) that using the module in a certain way will lead to poor or unpredictable results.

[0038] In yet another embodiment, the method further comprises:

[0039] - determining a complexity measure of the first user input, in particular of the natural language instruction, and

[0040] - including the example data and / or the high level textual in the multimodal prompt in dependence on the determined complexity measure.

[0041] Thus, whether or not the examples data and / or the high level description is included in the multimodal prompt depends on the level of complexity of the first user input, in particular in the form of a natural language instruction. The level of complexity is associated to the complexity measure, which can be, for instance, related to the number of sub-tasks or a readability level of the natural language instruction. Different formulas may be used to assess readability ranging from simple metrics such as word frequency count, number of characters in the textual user input, number of different words in the textual user input, percentage of unique words, number of prepositional phrases, etc. to more complex formulas such as the Flesh Formulas, the Dale-Chall Formula etc. Preferably, when the complexity measure is above a predetermined threshold value for high complexity, and thus the natural language instruction is deemed to be relatively complex, a high definition of the relevant modules is included in the multimodal prompt. In cases where the complexity measure is below a predetermined threshold value for low complexity, and thus the natural language instruction is deemed to be relatively simple, the example data is included in the multimodal prompt.

[0042] In another embodiment, the method further comprises:

[0043] - generating, using the provided multimodal prompt and the natural language model, an output file indicative of the lighting configuration task applied on the environment indicated by the image data. In this particular embodiment, the multimodal prompt is provided to the natural language model, which generates the output file that is associated to, and indicative of, the lighting configuration task applied on the environment indicated by the image data.

[0044] In a particular embodiment of the method of the first aspect of the invention, the output file comprises a text file or latent space representation including control instructions for controlling the one or more lighting units. The text file can for instance contain human readable text. Additionally, or alternatively, text file can contain latent space representations of control instructions e.g. a latent space representation of “warm white 800 lumen” or “cozy lighting scene”. The same readable text “warm white 800 lumen” in English will be at the same latent space representation as the German phrase “Warmes weisses Licht mit 800 lumen Lichtstrom” The text file text can also be indicative of light intensity values at points on a grid of the space. In the simplest case, the text file is an IES file that describes the intensity of a single light source in the room at points on a spherical grid.

[0045] Additionally, or alternatively, in another embodiment, the method further comprises generating and providing a control signal indicative of the control instructions to the lighting arrangement for controlling the one or more lighting units in accordance with the lighting configuration task. The control signal can be then provided to a controller unit of the lighting arrangement for controlling operation of the lighting units in accordance with the natural language instruction and the lighting configuration task.

[0046] Additionally, or alternatively, in another embodiment, the output file comprises an image file and / or a video file and / or a neural radiance field and / or a lightfield indicative of the lighting configuration task applied to the environment. For instance, the output file can be a modified image file or video file based on the image data received as second input data where the effects of applying the task on the environment are simulated. A lightfield is defined as a vector function that describes the amount of light flowing in every direction through every point in a space.

[0047] A second aspect of the present invention is formed by a controller for controlling implementation of a lighting configuration task for one or more lighting units of a lighting arrangement. The controller comprises:

[0048] - a first input unit configured to receive a first user input, in particular comprising a natural language instruction, indicative of the lighting configuration task;

[0049] - a second input unit configured to receive a second input comprising image data, indicative of an environment wherein the lighting configuration task is to be implemented;

[0050] - an input processing unit configured to determine, using the first user input (e.g., the natural language instruction) and the image data received, one or more sub-tasks that are associated with the implementation of the lighting configuration task and to generate, based on the first user input, output data (e.g., a textual description and / or output tokens) indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment.

[0051] The controller of the second aspect of the invention thus shares the advantages of the method of the first aspect of the invention or of any of its embodiments.

[0052] In the following embodiments of the controller of the second aspect will be described.

[0053] The first input unit and the second input unit can be implemented as data interfaces for receiving the first input, for instance in the form of a text file, and image or video file or an audio file, and the second user input, for instance in the form of a file, or a link to a file, or directly from a camera.

[0054] In an embodiment, the input processing unit is preferably implemented based on a language model with capability of analyzing image data, e.g. image files or video files, for example VisProg.

[0055] In an embodiment, the input processing unit is configured to identify, or otherwise determine, for the one or more sub-tasks, module data indicative of corresponding relevant software modules that are associated with the respective sub-tasks. The relevant software modules are selected from a predetermined pool or library of available software modules. Here, the textual description is further indicative of the identified module data.

[0056] In another embodiment, the input processing unit is additionally or alternatively configured to identify, or otherwise determine, lighting unit data indicative of operation parameters and / or capabilities of the respective lighting units. Here, the textual description is further indicative of the identified lighting unit data.

[0057] In another embodiment, the input processing unit is configured to generate a multimodal prompt including the image data, the natural language instruction, the textual description indicative of the sub-task, and optionally, the identified module data and / or the identified lighting unit data, and to provide said prompt to a natural language model.

[0058] In yet another embodiment, the input processing unit is configured to select, from a predetermined set of available example data, target example data indicative of examples of use of the relevant modules and to generate the multimodal prompt further including the target example data. Alternatively, or additionally, in another embodiment, the input processing unit is configured to generate the multimodal prompt further including a high level textual description of one or more of the relevant modules. The high level textual description preferably includes a description of input parameters required by the respective module and a description of module operation for obtaining an output parameter based on said input parameter.

[0059] In yet another embodiment, the input processing unit is also configured to determine a complexity measure of the natural language instruction, and to include the example data and / or the high level textual description in the multimodal prompt in dependence on the determined complexity measure.

[0060] Preferably, in another embodiment, the controller comprises a prompt execution unit that is configured to receive the multimodal prompt and to generate, based thereon, an output file indicative of the lighting configuration task applied on the environment indicated by the image data. The output file may comprise a text file including control instructions for controlling the one or more lighting unit in accordance with the natural language instruction. The prompt execution unit can be configured to provide a control signal indicative of the control instructions to the lighting arrangement for controlling the one or more lighting units in accordance with the lighting configuration task and / or an output file comprising an image or video file indicative of the lighting configuration task applied to the environment.

[0061] A third aspect of the present invention is formed by a computer program comprising instructions which, when executed by a controller in accordance with the second aspect cause the controller to perform the method of the first aspect.

[0062] It shall be understood that the computer implemented method of claim 1, the controller of claim 14, and the computer program of claim 15, have similar and / or identical preferred embodiments, in particular, as defined in the dependent claims.

[0063] It shall be understood that a preferred embodiment of the present invention can also be any combination of the dependent claims or above embodiments with the respective independent claim.

[0064] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In the following drawings:

[0066] Fig. 1 shows a schematic block diagram of an exemplary controller for controlling the implementation of a lighting configuration task in accordance with one embodiment of the invention;

[0067] Fig. 2 shows a schematic block diagram of another exemplary controller for controlling the implementation of a lighting configuration task in accordance with another embodiment of the invention;

[0068] Fig. 3 shows a schematic block diagram of another exemplary controller for controlling the implementation of a lighting configuration task in accordance with another embodiment of the invention;

[0069] Fig. 4 shows a schematic block diagram of another exemplary controller for controlling the implementation of a lighting configuration task in accordance with another embodiment of the invention;

[0070] Fig. 5 shows a schematic flow diagram of an exemplary method for controlling the implementation of a lighting configuration task in accordance with an embodiment of the invention.

[0071] DETAILED DESCRIPTION OF EMBODIMENTS

[0072] Fig. 1 shows a schematic block diagram of an exemplary controller 100 for controlling implementation of a lighting configuration task LT for one or more lighting units 11, 12 of a lighting arrangement 10. The controller comprises a first input unit 102 configured to receive a first user input 104, which, in this particular example, comprises a natural language instruction 105 indicative of the lighting configuration task LT. Other first user inputs may include other data files types, such as an image file or an audio file, that indicate the lighting configuration task. For instance, in the case of multi-modal large language models (LLM), such as for instance Google’s Gemini and Open Al’s ChatGPT, that accept text files as well as other types of files such as image or audio files, some LLM are advantageously configured to process visual representations of images or video, or sound representations in audio files, to obtain relevant information that can be indicative of a lighting configuration task. For instance, a vision system can be trained (e.g., supervised, semi-supervised, self-supervised) to turn image into a high-level representation. This high level representation is basically a list of tokens that are similar to the kind of tokens that a LLM takes as an input. The tokens are then feed that to the LLM in addition to the text, and, during training, the LMM is be able to use the high level representations for determining the one or more sub-tasks. The same concept can be for audio input as first user input.

[0073] Thus, The first user input 104 can be provided, for instance, in the form of a text file including text indicative of the lighting configuration task, or in the form of an image or video file, also including text indicative of the lighting configuration task, or in the form of a video or an audio file, including a sound signal indicative of the lighting configuration task. Lighting configuration task LT refers in general to tasks pertaining to the operation or control of the lighting units 11, 12, of the lighting arrangement 10, and that include, among others, configuration, commissioning and / or operation of the lighting units, either for the direct operation of the lighting units or for a simulation of said operation. Thus, the lighting units can be existing lighting devices whose operation can be controlled by the lighting configuration task, or virtual lighting devices, whose operation can be simulated according to the lighting configuration task LT.

[0074] A second input unit 106 is configured to receive a second input 108 that comprises image data 110, and that is indicative of an environment 112 wherein the lighting configuration task is to be implemented. In the present example, the image data is an image file depicting a dining area 112 as an environment illuminated by a hanging luminaire 11 and a floor lamp 12 that form part of a lighting arrangement 110. The image date 110 can be, for instance, provided as an image file or, alternatively, as a video file.

[0075] The controller 100 also comprises an input processing unit 114, for instance including a large language model LLM (e.g. GPT, Gemini, LLaMA, Claude, Mistral, etc.) configured to determine, using the first user input 104, e.g., as a natural language instruction 105, and the image data 110 received, one or more sub-tasks 116.1, 116.2,... 116.N that are associated with the implementation of the lighting configuration task, and to generate, based on the first input, output data 118 indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment. The output data can for instance comprises human-readable or human-understandable text, for example a textual description of the one or more sub-tasks. Alternatively, or additionally, the output data can include one or more tokens that are similar to the kind of tokens that a multimodal LLM takes as an input (e.g. when ingesting an image). Tokens can be in the form of non human-understandable text strings, that are understandable by a given large language model.

[0076] For instance, the following natural language instruction 105 is provided, as a text or audio file, via the first user input 102: “Turn of the lights above the dining table”. The natural language instruction includes or makes reference to a lighting configuration task, namely the appliance of a given lighting instruction (turn-off) on some lighting units, placed in a given environment. Additionally, the image data 108 depicting the environment 112, and in this particular example, the lighting units 11 and 12 is provided via the second user input 108.

[0077] To control the implementation of the lighting configuration task, the input processing unit 114 is configured to understand the intent of the instruction 105 and then identify a sequence of steps or sub-tasks necessary to implement the lighting configuration task. In the current example, the sub-tasks 116 include, for example, detect the dining room in the image data (e.g., 116.1), detect objects in the dining room (e.g., 116.2), detect the table among the objects and determine lighting units or lighting fixtures among the objects. These steps can be performed by state of art multimodal models like GPT4. Further subtasks include determine lighting settings for the determined lighting unit (e.g., 116.N) based on the instruction, namely turn off the lighting unit located above the table. In order to implement this step, and as it will be explained below, lighting specific models typically not openly available might be required.

[0078] The output data 118, e.g., the textual description, provided indicates the steps or sub-tasks 116 that are necessary to implement the lighting configuration task on the lighting units in the environment.

[0079] Instructions given by user of lighting arrangements, such as Signify Hue lighting arrangement, can be highly creative, in terms of what they ask to a Generative Al model. Due to the high diversity of lighting needs, it is not particularly advantageous to creating systems that can create one task or five tasks or even ten tasks. Systems that can create a large number of tasks that users want to do are desirable. Similarly, the lighting industry needs more specialized commissioning flows (instead of the current cookie cutter flow) to both lower the commissioning cost and enable new niche-specific features in lighting systems, e.g., InterAct systems. The controller 100 is advantageously configured to use aLLM, e.g., an off-the-shelf LLM, to build customized lighting specific computer vision solutions on the fly. The image data can be for instance provided by making use of the vision sensing modalities included in smartphones or AR / VR headset.

[0080] Many lighting-related use cases involve compositional visual reasoning subtasks such as understanding the furniture setup of a room to base thereupon a commissioning or controlling of the lighting units or the arrangement and / or generate a well-matching lighting scene. The controller 100 is advantageously configured to tackle the long tail of complex lighting-related commissioning & controls tasks by decomposing a given lighting configuration task (as described in natural language by the user) into simpler steps or subtask that are provided as a textual description 118. This textual description 118 may be handled outside of the LLM by a multitude of specialized end-to-end trained machine learning modules or other programs such as conditional logic.

[0081] Fig. 2 shows a schematic block diagram of another exemplary controller 100 for controlling the implementation of a lighting configuration task LT in accordance with another embodiment of the invention. For the sake of clarity and simplicity, those features of the controller 100 of Fig. 2 that have similar or identical functionalities will be referred to using the same reference numbers as those used in the description of the controller 100 of Fig. 1. For some lighting configuration tasks, specific models or modules which might not be openly available, might be required. Thus, advantageously, the controller 100 of Fig. 2 is configured to identify, for one or more of the one or more sub-tasks 116, module data MD1, MD2 indicative of corresponding relevant software modules MD.l, MD.2,... MD.N, that are associated with the respective sub-tasks. The resulting textual description 118 is further indicative of the identified module data. Relevant software modules MD.1, MD.2 are preferably identified from a pool or library 115 of available software modules. The relevant software modules can include, for example, mathematical modules or models pertaining to the implementation of lighting functions of the lighting units.

[0082] Further, the controller of Fig. 2 is also configured to identify, for the one or more lighting units 11, 12, lighting unit data LD.1, LD.2 indicative of operation parameters and / or capabilities of the respective lighting units. The textual description 118 is thus further indicative of the identified lighting unit data. The lighting unit data can be identified from a pool or library 117 of available lighting data. The lighting unit data is thus associated to the lighting units and indicative of the parameters of the lighting unit and / or parameters under which they can be operated. Lighting data unit may include data pertaining to a dimension or size of the corresponding lighting unit, possible mounting positions, type of illumination that can be provided by the respective lighting unit, for instance in terms of intensity, temperature, color, direction of illumination, directivity, including fixed values and / or possible ranges.

[0083] For example, if the natural language instruction 105 is “Set the lights above the dining table according to a morning walk light scene”, apart from the sub-tasks discussed above (detect the dining room, detect objects in the dining room, detect the table among the objects, determine lighting fixtures among the objects), the sub-task “Determine lighting setting for the determined fixture based on instruction” requires lighting specific modules, which are the relevant modules associated to this sub-task. Further after having identified the lighting units, lighting unit data indicative of the capabilities of said lighting unit is also provided.

[0084] The relevant software modules include, in this particular example, a “Text- to-light” module that takes a text description as input and provides one or more colors corresponding to said description as output, and a “dynamics” modules that takes one or more colors and a desired level of dynamics (low, medium, high) as input and provides a dynamic lighting setting based on the one or more colors and level of dynamics as output.

[0085] The first module converts a textual input into one or more RGB colors. In this example, the input processing unit 114 (e.g. a multimodal LLM) can derive that the “morning” word is the important parameter from the user input for the choice of colors and this should be the string input for the first module. The output of the first module is then used as an input for the second module, along with the desired level of dynamics. The input processing unit may further derive that a “morning walk” is typically associated with a lower level of dynamics compared to for example a morning run, which is typically more high-intensity and energetic. Thus, the input processing unit may derive that “low” is the first input for the dynamics model, while the output RGB colors of the first module shall be used as the second input for the second module. The second module generates a dynamic lighting setting that includes the one or more RGB colors corresponding to a “morning” scene and corresponding timestamps as an output. This is provided in the textual description 118, which, as it will be explained below, can then be used to control the lighting arrangement to create the desired morning walk light scene. Fig. 3 shows a schematic block diagram of another exemplary controller 100 for controlling the implementation of a lighting configuration task in accordance with another embodiment of the invention. For the sake of clarity and simplicity, those features of the controller 100 of Fig. 3 that have similar or identical functionalities as the controllers described with reference to Fig.1 and Fig. 2 will be referred to using the same reference numbers as those used in the description of the controller 100 of Fig. 1 and Fig. 2. The input processing unit 114 is here advantageously configured to generate a multimodal prompt 120 that includes the image data 110, the natural language instruction 105, the textual description indicative of the sub-tasks, including the module data and the lighting unit data. The multimodal prompt 120 is thus an augmented prompt that includes the original natural language instruction, the image data and the textual description, which can include the identified module data and / or the identified lighting unit data.

[0086] In Fig. 3, the exemplary controller comprises, or is connected to, a prompt execution unit 122, that is configured to receive the multimodal prompt 120 and to generate, based thereon, an output file 124 indicative of the lighting configuration task applied on the environment indicated by the image data 108. In the example shown in Fig. 3, the prompt execution unit comprises a natural language model that which receives, as a prompt, the multimodal prompt 122 provided by the input processing unit 114.

[0087] As an example, the natural language instruction 105 provided to the input processing unit is “substitute in the provided image the hanging lamp in the dining room with a bar pendent light model X of company Y and show the brightest illumination settings possible”. Further, an image of the dining room with the hanging lamp 11 and a floor lamp 12 is provided. The input processing unit is configured to determine the relevant sub-tasks for implementing the lighting configuration task of substituting, in the provided image, the hanging lamp with the desired substitute lamp. The input processing unit access the pool or library of lighting unit data for obtaining, for example, relevant information about the characteristics (physical and operational) of the light model X of company Y.

[0088] Relevant modules are also identified and provided as part of the textual description 118. For instance, “Text-to-light” module that takes a text description as input and provides one or more colors corresponding to said description as output, and a “dynamics” modules that takes one or more colors and a desired level of dynamics (low, medium, high) as input and provides a dynamic lighting seting based on the one or more colors and level of dynamics as output.

[0089] The resulting multimodal prompt is provided to the prompt execution unit 122, which generates and provides the output file 124, for instance in the form of a modified image data 108’ where the ceiling lamp (lighting unit 11) has been identified, removed, and replaced by a representation of a bar pendent light model X of company Y (lighting unit 13), showing how the brightest illumination setings would look like when applied to the environment 112 shown in the original image data 108. For example, the dimensions, the installation mode, and the maximum brightness has been extracted from the lighting unit data LD. In this case, the lighting unit 13 is a virtual lighting unit and the output file shows a simulation of how the virtual lighting unit would look like and operate within the environment 112.

[0090] Alternatively, the output file may comprise a text file including control instructions for controlling the one or more lighting unit in accordance with the natural language instruction. For instance, the natural language instruction can be indicative of a lighting configuration task directed to implementing a given light scene on the lighting arrangement. The prompt execution unit can be configured to provide a control signal indicative of the control instructions to the lighting arrangement for controlling the one or more lighting units in accordance with the lighting configuration task.

[0091] In a particular example, a controller runs a first LLM (e.g., GPT 4) in the input processing unit 114 for orchestrating and a second LLM (e.g., a fine-tuned Llama model) for commissioning. The second LLM is acting as a specialized software module. The first LLM does not necessarily have to output a human readable text which is then fed to the second LLM. Instead, the first LLM may output non human-readable / understandable text strings as input for the second LLM. The non human-understandable text strings are however tokens that are understandable by the second LMM, which are then fed to said second LLM. Normally, the second LLM is fed with a human-understandable text. However, as multi-modal LLMs like Gemini have shown it is possible to feed those non- human-understandable tokens (which however are LLM-understandable) to an off-the- shelf LLM. Fig. 4 shows a schematic block diagram of another exemplary controller 100 for controlling the implementation of a lighting configuration task in accordance with another embodiment of the invention. For the sake of clarity and simplicity, those features of the controller 100 of Fig. 4 that have similar or identical functionalities as the controllers described with reference to Fig.1, 2 and Fig. 3 will be referred to using the same reference numbers as those used in the description of the controller 100 of Fig. 1, 2 and Fig. 3.

[0092] In the exemplary controller 100 of Fig. 1, the input processing unit 114 is further configured to select, from a predetermined set of available example data, target example data EXI, EX2 indicative of examples of use of the relevant modules and to generate the multimodal prompt 120 further including the target example data EXI, EX2. Additionally, or alternatively, the input processing unit is also configured to generate the multimodal prompt 120 further including a high level textual description of one or more of the relevant modules MD1, MD2. The high level textual description of a module comprises a description of input parameters required by the respective module and a description of module operation for obtaining an output parameter based on said input parameter.

[0093] Preferably, the unit processing unit 114 is configured to determine a complexity measure of the natural language instruction, and to include the example data EXI, EX2 and / or the high level textual description in the multimodal prompt 120 in dependence on the determined complexity measure.

[0094] For instance, the controller 100 can advantageously configured to leverage an off-the-shelf language model, to create, as a possible output file, Python programs using in-context learning. Instead of building unified multi-modal models (e.g. UnifiedlO or Flamingo) to act as a single unified “lighting world model”, Python programs are generated naturally from natural language descriptions of the lighting-specific task provided by the user. While unified world models such as Google' s Flamingo are difficult to keep up to date in the rapidly moving machine learning field, the proposed approach builds upon always leveraging the latest version of existing specialized modules / models for the sub-tasks.

[0095] Optionally, in-context exemplars of similar tasks are used, which comprise i) natural language instructions; and / or ii) descriptions of python programs, both of which can be prepended to the user-provided natural language instruction, also referred to as user prompt. Hence, thanks to the prepended information, the prompt execution unit 122 is able to firstly understand the user's commissioning or lighting controls task, secondly is able to understand the capabilities of machine learning modules (e.g. software modules) available for its disposal and thirdly understand how these modules may be used in a python program. Unlike prompting a LLM (e.g., GPT-4) with solely the user-prompt alone, the LLM will be now able to produce a correct Python program invoking the right lighting-specific sub-modules (outside of the LLM) to solve the lighting task described in the user’s prompt.

[0096] While fine-tuning a single end-to-end model to a novel task needs thousands or millions of examples to leam a new lighting-specific task, the proposed visual programming method produces great results on new tasks with as little as 20 in-context examples and the method involves neither fine-tuning (i.e. no changing of weights) of GPT-4 nor fine-tuning of any of the other machine learning models used in the specialized modules.

[0097] In addition, the machine learning modules invoked by the system can be asked to summarize the step’s computation -for instance, visually- using html (e.g. after the ML system made the inference it shows to the user on an image how it first removed the spotlighting effect on an image and then replaced it with a downlighting effect, along the Python code invoking machine learning modules for doing so). By using HTML to visualize the computation in each individual step, the controller 100, acting as a visual prompting lighting ML system, produces highly interpretable execution trees which serve for the user as easy-to-understand visual rational for the prediction of the overall task.

[0098] The controller 100 can be configured to generate high-level Python programs (scripts) that invoke trained state-of-the-art neural models and other python (scripting language) functions at intermediate steps as opposed to generating an end-to-end neural network for the specialized task at hand.

[0099] Fig. 5 shows a schematic flow diagram of an exemplary method 500 for controlling implementation of a lighting configuration task for one or more lighting units of a lighting arrangement. The method can be advantageously carried out by a controller 100, as described with reference to Figs. 1 to 4 above. The method 500 comprises, in a step 502, receiving a first user input comprising a natural language instruction indicative of the lighting configuration task. The method 500 also includes, in a step 504, receiving a second input comprising image data indicative of an environment wherein the lighting configuration task is to be implemented. The method 500 further comprises, in a step 506, determining, using the natural language instruction and the image data received, one or more subtasks that are associated with the implementation of the lighting configuration task in the environment, and, in a step 508, generating, based on the first input, a textual description indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment.

[0100] Optionally, as indicated by the dashed boxes in Fig. 5, the method 500 can comprise one or more of the following steps:

[0101] - identifying, in a step 510, for the one or more sub-tasks, module data indicative of corresponding relevant software modules that are associated with the respective subtasks, and wherein the textual description is further indicative of the identified module data; and / or

[0102] - identifying, in a step 512, for the one or more lighting units, lighting unit data indicative of operation parameters and / or capabilities of the respective lighting units, e.g. the lighting units referred to in the lighting configuration task, and wherein the textual description is further indicative of the identified lighting unit data; and / or

[0103] - generating and providing to a natural language model, in a step 514, a multimodal prompt including the image data, the natural language instruction and the textual description.

[0104] The method 500 can also optionally include, in a step 516, determining a complexity measure of the natural language instruction, and, in a step 518, including example data and / or a high level textual description of the relevant modules in the multimodal prompt in dependence on the determined complexity measure.

[0105] The method 500 can also include, in a step 520, generating, using the provided multimodal prompt and the natural language model 123, an output file indicative of the lighting configuration task applied on the environment indicated by the image data, and / or, in a step 522, generating and providing a control signal indicative of control instructions to the lighting arrangement for controlling the one or more lighting units in accordance with the lighting configuration task. In the following, non-limiting examples of use of the method and the controller according to the invention will be discussed.

[0106] For instance, in the example of Fig. 4, the relevant feature refers to decision criteria for the unit processing unit 114 when to include in the multimodal prompt 120 to the prompt execution unit 122, comprising an LLM 123: i) a specialist exemplar as example data (e.g., with python code) calling a subset of the available sub modules (please note that the exemplar has to be relatively close to the current task at hand to be useful demonstration for the LLM) vs. ii) when to include a high level model-capability description (the high level textual description) of the relevant module rather than an exemplar (a model-capability - description is a more generalist approach and hence will also work for previously unseen, out-of-distribution requests by the user), vs. iii) when to provide both a specialist exemplar(s) EXI, EX2 and a high level model capability description(s) to the prompt execution unit 122.

[0107] For instance, the input processing unit 114 can be configured to provide to the LLM 123 a generalist description of the module’s capabilities (hence following approach i)). Alternatively, a generalist description of the relevant module can be provided in addition to an exemplar for the user’s new task at hand.

[0108] On the other hand, approach iii) can be implemented whenever a moderately well matching exemplar is available.

[0109] Approach ii) can be selected for instance when only ill-matching exemplars for the new task are available.

[0110] Similarly, approach iii) can be selected when it is recognize that the task at hand is complex and hence, in addition to a well matching exemplar EXI, EX2, also the generalist description of the module is required for the LLM to successfully automatically create the specialized machine learning application for the task.

[0111] This is preferably done by determining a complexity measure of the natural language instruction, and including the example data and / or the high level textual description in the multimodal prompt in dependence on the determined complexity measure. Complexity measure can be correlated to the number of sub-tasks, readability level (Different formulas may be used to assess readability ranging from simple metrics such as word frequency count, number of characters in the textual user input, number of different words in the textual user input, percentage of unique words, number of prepositional phrases, etc. to more complex formulas such as the Flesh Formulas, the Dale- Chall Formula etc.).

[0112] Thus the input processing unit 114 can be configured to task a generic, off- the-shelf LLM such as GPT-4 to generate the overarching python program across the various computational sub-modules. The LLM can be tasked, as the final step, to autoprompt itself for the final answer to the user’s question whereas the LLM is utilizing the previous answers from sub-modules for its final answer.

[0113] Thus, a modular and interpretable neuro-symbolic system is proposed for compositional visual reasoning in an image involving both a room scene (environment) and a lighting arrangement. To serve the long tail of complex lighting tasks in the wild, an Al solution is proposed that avoids the need for any task-specific model training (e.g., involving modification of the model weights). Instead, it is proposed to use the in-context learning ability of large language models to generate, for instance, python-like modular programs, which are then executed to obtain both the solution as well as a comprehensive and interpretable rationale explaining step by step what the LLM 123 decided to do. These easy-to-understand Python programs can subsequently be easily verified by the commissioning person and / or end-user for logical correctness. Alternatively or additionally, the python program could be further processed and represented to the user using a higher lever representation, e.g., a block diagram that might be more appropriate for non-expert users to review the logical reasoning steps of each module utilized for solving the user’s required lighting configuration task.

[0114] Each line of the generated Python program may invoke one of several off- the-shelf computer vision models, image processing subroutines, or python functions to produce intermediate outputs that may be consumed by subsequent parts of the program.

[0115] Thanks to the impressive in-context learning ability of large-scale LLMs, it is possible to create these customized Python programs solely by the user prompting the LLM (in the input processing unit 114) with a simple natural language instruction 105 (or a visual question or a statement the user wants to be verified) along with a few examples of similar instructions and their corresponding Python programs. These examples / exemplars EXI, EX2 may for instance be stored on a Hue / InterAct bridge or be available in a library. The proposed approach hence removes the need to train specialized program generators for each of the novel lighting tasks that a user base will come up with. Similarly, for InterAct, it is no longer needed to manually create specialized commissioning flow for niche segments.

[0116] The proposed approach is also well suited for neuro-symbolic problemsolving approaches, which use besides the LLM also specialist non-neural modules such as conditional logic (expert system). The neuro-symbolic approaches allow to solve complex and compositional visual lighting tasks given natural language instructions.

[0117] Given a few exemplars of natural language instructions and the desired high-level python programs invoking sub-modules to step-by-step solve the task, a visual programming may generate a new program for any new user-instruction provided during test time. For instance, the python programs can be generated by solely using in-context learning in GPT-4 and then execute the newly generated Python program on the input image(s)108 from smartphones and / or AR headsets to obtain with the help of the submodules a prediction on a previously unseen new image 108’. The Python program then invokes these submodules to sequentially solve each of the sub-task generated by the input processing unit 114.

[0118] If a user-provided prompt is deem to be challenging, for example in accordance with a determined complexity measure, a larger number of exemplars can be for instance pre-pended in multimodal prompt 120 to the LLM 123 (instead of a single exemplar EXI).

[0119] In addition, the visual programming tool can be also configured to summarize the intermediate outputs (intermediate text; image edited by a sub-module; bounding boxes; segmentation masks) into an interpretable visual rationale which can be easily reviewed by a non-AI-expert user or the commissioning person wanting to create, for example, a custom InterAct system for a customer.

[0120] The set of relevant modules may include a first module for light-effect understanding and a second module for image manipulation (e.g. altering a light effect in a synthetic image of the room), a third module for deciding which lights to actuate to generate a desired lightplan within the room, a fourth module for knowledge retrieval, and a fifth module arithmetic and logical operations (e.g. non-neural python subroutines).

[0121] Specialist relevant modules may be for instance be off-the-shelf computer vision models, language models, CLIP as an open-vocabulary image classifier, image processing subroutines in OpenCV (e.g. module for compositional visual question answering; zero-shot natural language visual reasoning (NLVR) on image pairs), factual knowledge object tagging from natural language instructions, language-guided image editing (remove objects; add light effects to the image), spatial mapping of the room from panoramic scans with the iPhone (e.g. Apple Room Plan), extracting information from a building’s floorplan / technical building drawing or arithmetic / logical operators for calculation of energy consumption.

[0122] The controller 100 can also call on lighting-design specific modules, which may include software packages such as AGi32, CalcuLuX, DIALux, mi mi, Radiance, Microlux, LightCalc, Visual 3D.

[0123] Additionally, the controller 100 can also call on Signify-in-house created lighting specific modules such as HueStudio (capable of color extraction from images, scene creation ruleset - e.g., what colors work best with each other, etc.) and the Hue Entertainment library (color extraction and color to location mapping), stochastic models for dynamic content generation (e.g., based on Perlin noise).

[0124] The relevant modules may consume inputs that are produced by executing previous lines of the LLM-generated Python code and output intermediate results that can be consumed by other modules downstream in the Python program generated by the LLM. Each module may consume multiple arguments including strings, numbers, arithmetic and logical expressions, or arbitrary python objects (such as list() or dict() instances containing bounding boxes or segmentation masks) produced by previous steps.

[0125] The proposed approach builds upon all the available work rather than attempting to make one single gigantic world model for the lighting domain (which is typically not feasible due to the lack of highly specialized lighting training data).

[0126] The proposed solution can be applied to diverse lighting tasks: Firstly, it can perform compositional visual question answering (e.g. furniture setup in the room; user activities in the room e.g. watching TV), which serves then as input for the conditional logic deciding the grouping of the luminaires (for commissioning purposes) and / or the to-be-administered lighting scene according to the task.

[0127] Secondly, it can perform language-guided editing of the light scene rendered in the room (the light scene may be observed with the help of a camera) Thirdly, it can perform zero-shot reasoning on image pairs. For instance, the image pair may consist of a first image depicting a living room with specific atmosphere the user likes and a second image may be of the user’s own living room. The user requests from our system to replicate the lighting atmosphere of the first room in his actual living room; the LLM hence decides to call upon a first specialized module capable of identifying the luminaires in the images and a second module capable of compare their light rending abilities based on the IES file documents found by the LLM. The LLM then uses another module to select which luminaires in the user's room to use for creating the targeted lighting atmosphere.

[0128] Fourthly, it can perform factual knowledge object tagging on lighting objects e.g. narrow light beam tag for a spot light fixture and omnidirectional light beam tag for a downlight fixture)

[0129] The advantage of the proposed approach is that the scope of Al systems can be easily and effectively expanded to serve the long tail of lighting tasks desired by the user base or bringing our InterAct platform to niche applications beyond generic office layouts and generic manufacturing.

[0130] In an example, the user’s prompt 105 instructs a vision system such as the controller 100 to “Tag the luminaire types in this image and identify what role the each of the lights play in the user’s activity depicted on the image.” To perform this lighting configuration task, the controller first needs to understand the intent of the instruction and then perform a sequence of steps - detect the light sources, retrieve list of luminaire architypes from a knowledge base (e.g. by using GPT-4 as a knowledge retrieval system), classify luminaires on the image using the list of luminaire architypes, and tag the image with recognized luminaire’s bounding box and names. In addition, to user' s task there is also the need to use an open-vocabulary image classifier and open-vocabulary localization module as well modules specialized for spatial prepositions (such as the first light is located above, left, etc. of the user and to the right of the second light).

[0131] While different vision and language systems exist to perform each of these steps, executing this task described in natural language is beyond the scope of end-to-end trained systems. In the example above, the visual Python program generated by the LLM multimodal prompt 120 will correctly invoke all these sub-modules to produce the desired output for the user without any problems. Based on a decision criteria, e.g., a complexity measure, the controller 110 can decide when to provide the LLM 123 merely with the application note document for each module versus when to provide the LLM 123 with a more detailed description of what the module can do for the specific room scene context or room type at hand (e.g. if it is known that there is no dog in the room scene at hand; the portion of the module’s application note describing what the module can do for dogs is not relevant), versus when to provide both the complete application note (for instance the hueStudio application note), versus when to provide the LLM 123 a crip custom-summary of the most relevant aspects of the module’s capabilities (e.g., customized for the user task at hand).

[0132] For instance, the multimodal prompt 120 we describe in words to the LLM 123 how a first hueReLight module, which can be based upon the Apple-RoomPlan-based re-lighting model, can be used given the re-lighting task requested by the user. The following high level text can be prepended to the user’s prompt 105 to explain the module’s capabilities to the LLM 123:

[0133] “This hueReLight module —given an image or video stream and given virtual light source(s) not present on the image and their desired locations-, can visualize for you said virtual light sources on the image or video stream using a given light effects. This module may be useful for instance if the user requests “Here is a picture / stream of my living room add a Hue Signe in the left comer of my room and recall Savanna sunset scene”. Alternatively, you can use the hueReLight module to re-light physical light sources on the image / video stream; in this second case the type and location of physical light sources (on the image) could be given as input. The output of the hueRelight module in this case is an image / video stream wherein new (suggested) light effects are rendered.”

[0134] A second module may be for generating light scene based on key words. The capabilities of the model can be described in the following way to the LLM 123:

[0135] “This Text2Light module -given a set of keywords describing a desired atmosphere in the room and / or user activity as input—, can generate for you five palettes each consisting of 5 colors, brightness, and dynamic levels. We give you now an exemplar how to use the hueReLight and hueKeyWord modules for a Python program for solving the following user request (prompt): “Visualize how my living room will look if I add one pixelated Hue Signe in the left comer of the room and have a relaxing sunset like scene”. The resulting program (using description format of VisProg) looks like this: [

[0136] OBJO=Seg(image=IMAGE, query = “Left comer”) / / localizing left comer OBJl=Light(query=”pixelated Hue Signe”) / / geting Hue Sign model from database (knowledge retrieval)

[0137] OBJ2=Palete(query=”relaxing sunset”) / / module for generating palete based on keywords

[0138] IMAGE l=Relight(image=IM AGE, locations = OB JO, object=OBJl, scene=OBJ2) / / relight module could also be split into a few more basic modules depending on the implementation

[0139] RESULT = IMAGE 1

[0140] ]

[0141] Another module may be able to assess the efficacy of a lighting intervention and modify the intervention by re-enforcement techniques.

[0142] Another module may be specialized on home security. For instance, a description can be provided to the LLM 123 that the dynamic-lighting-effect definition module is knowledgeable how to drive burglars away with the help of stroboscopic light effects and / or how to use lights for presence mimicking while the owner is away.

[0143] Providing a high level text description of the module’s capabilities, rather than providing the LLM 123 with exemplars EXI, EX2, is much more generalizable to diverse smart home use cases in the wild. Hence, text description of the modules’ capabilities is a preferred embodiment, in particular for highly complex user prompts 105.

[0144] In a further developed embodiment the LLM 123 is provided with an explanation, under which circumstances the relevant module MD1, MD2, works well and under which it does not work. In the text prepended to the user’s prompt 105 and part of the multimodal prompt 120, performance metrics of the module under different circumstances can be included (e.g. performance under low light condition vs regular lighting conditions). For instance, again using hueReLight and Text2Light modules as example, the multimodal prompt 120 to the LLM 123 can include the following;

[0145] "The palete provided by Text2Light should be within a color gamut of the physical hue lights in the user’s room (that can be given as a parameter). For the Relight module, please analyze the resulting image whether it has colors that are equal or close to the palete from Text2Light, and whether all colors are present in the part of the image where pixelated Signe is located, so a subset of image / video defined by Seg module could be further analyzed.”

[0146] In general, in the prepended text provided to the LLM 123, the LLM is told to always match each module output to the capabilities of the lighting arrangement, for instance as indicated by lighting unit data pertaining to the lighting units (real or virtual) of the lighting arrangement (e.g., a Hue system). For example, if a relevant software module is supposed to generate light settings for each light source, does the number of settings generated by the module match the number of physical hue lights in the room? Or, do they match the type (e.g., multiple colors for pixelated lights, CCT for white lights)? The explanation regarding the circumstances under which the modules work is then advantageously prepended in the multimodal prompt 120.

[0147] The controller 100 can be also used to map and tag image data-base of room scenes so that a used subsequently can, during interactions with architects / end-users, quickly retrieve images of lighting projects which are similar to an envisioned lighting design (an example of lighting configuration task). To create tags for an image depicting a reading scene, one user question may be: “Is the small spotlight luminaire located to the left or to the right of the person that is reading a book”. The input processing unit 114 can create a visual programming program which first localizes “person reading a book”, then crops the region to the left (or right) of this person, checks if there is a “small spotlight” on that side, and return “left” if so and “right” otherwise. The LLM can then use a specialized question-answering module based on a Vision-and-Language Transformer such as VILT, but instead of simply passing the user’s complex original question to ViLT, the visual prompting system invokes the ViLT for simpler tasks like identifying only the contents within an image patch.

[0148] As a result, the Python program generated by the LLM 123 is not only more interpretable than merely using VILT with the user’s original question 105 but also more accurate. Alternatively, one could completely eliminate the need for the QA module to use a question-answering (QA) model like ViLT; instead the LLM may create its own QA module on invoking other systems like CLIP and object detectors.

[0149] Further, an UI for Visual question answering for lighting design and commissioning which shows the reasoning steps is proposed. For instance, the UI shows that the reasoning firstly involves determining the number of lamps, then secondly determining the locations of each of the lamps and users, then thirdly determining the type of activity (context), and finally generating a context-based lighting scene consisting of a color palette and assignment of specific colors to specific lamps (e.g., based on a set of keywords that describe activity and context).

[0150] Also, Zero-Shot Reasoning on Image Pairs VQA models can be trained to answer questions about a single image. In practice, one might require a visual-questionanswering system to answer questions about a collection of images. For example, a user may ask a system to parse images from a camera (e.g., Raven) and answer the lighting configuration task related question, such as, for example: “Which lighting scene was selected when I woke up in the morning after dining in my home with Sarah?” Other examples could be “What scene was active before we switched on the video streaming?”, or even “Who forgot to switch off the light last night?” or “Who was in the room when this scene was activated?”

[0151] Instead of assembling an expensive dataset and training a multi-image model, the proposed solution allows to use an off the shelf single-image VQA system to solve a task involving multiple images without training on multi-image examples. This is done by decomposing a complex statement into simpler questions about individual images and a Python expression involving arithmetic and logical operators and answers to the image-level questions.

[0152] By the visualizing the reasoning step and output by each of the involved submodules, the UI reveals the reason for failures of the Al system to a user who can then suggest to the Al how to correct its errors. Additionally, the proposed UI will allow the user to modify his original instruction to improve performance.

[0153] Also, users often want to identify lighting units in images whose names are unknown to us. For instance, a user might want to identify the luminaire archetypes (downlight, spotlight, etc.) in his room and / or their luminaire make (Acuity, GE Current, Signify, etc.. This is for instance important in refurbishment business, where in an old building there are no records which luminaires are installed.

[0154] Solving this task requires not only localizing luminaires but also looking up factual knowledge in an external knowledge base (Internet; proprietary Signify customersupport database) to construct the set of categories for the specific classification task at hand, such as name and architype of each luminaire in the public Acuity catalogue available on the Internet.

[0155] Another task is Factual Knowledge Object Tagging or Knowledge Tagging for short. For solving a lighting-related Knowledge Tagging task, we use GPT-4 as an implicit knowledge base that can be queried by the end user with natural language prompts such as “List the downlights in the product catalogue of Signify Genlyte separated by commas.” This generated category list can then be fed by the Python code (which has been auto generated by the LLM) to a specialized a CLIP image classification module that classifies image regions which are produced by localization and object detection modules.

[0156] Our visual program Python generator automatically determines whether to either use a specialized luminaire object detector module or an open-vocabulary object localizer module depending on the context in the natural language instruction.

[0157] The controller 100 can also be configured to provide lighting-specific Image Editing with natural language. A user may provide as a natural language instruction 105, the following exemplary prompt: “Hide the Acuity luminaires in the image with a heart emoji”, for instance for de-identification or privacy preservation. Another example of deidentification prompt could be “Hide the luminaires in the image but preserve the light effect in the room”, for instance to share the images without disclosing the exact look of a new, not yet released luminaire or as the lighting designer still needs to design a 3D printed housing, in this case “hide” could be additionally explained to the controller as in: “Hide the prototype luminaires with off the shelf luminaires of the same type and preserve the original light effect”, or ’’Create a color pop of the luminaires and blur the background” (object highlighting), or “Replace the Pixelated LED strip with the same LED strip model with a diffuse optics” (object replacement). For instance, the LLM 123 understands thanks to the explanation provided regarding modules / exemplars (which are prepended as text in front of the actual user prompt 105) that the latter task first requires identifying the object of interest, generating a mask of the object to be replaced and then invoking an image in painting model (for instance Stable Diffusion) with the original image, mask specifying the pixels to replace, and a description of the new pixels to generate at that location. The LLM generates the output file for instance as a Python code for sequentially invoking these sub-modules accordingly. In summary, the invention is directed to a computer implemented method for controlling implementation of a lighting configuration task for one or more lighting units of a lighting arrangement, which comprises receiving a first user input, e.g., a natural language instruction, indicative of the lighting configuration task, receiving a second input comprising image data, indicative of an environment wherein the lighting configuration task is to be implemented, determining, using the natural language instruction and the image data received, one or more sub-tasks that are associated with the implementation of the lighting configuration task in the environment; and generating, based on the first input, output data, e.g., a textual description, indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment. The method enables a composition of specialized models for lighting configuration task.

[0158] Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims.

[0159] In the claims, the word "comprising" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality.

[0160] A single unit or device may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.

[0161] A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium, supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.

[0162] Any reference signs in the claims should not be construed as limiting the scope.

Claims

CLAIMS:

1. A computer implemented method (500) for controlling implementation of a lighting configuration task (LT) for one or more lighting units (11, 12) of a lighting arrangement (10), the method comprising:- receiving (502) a first user input (104) indicative of the lighting configuration task;- receiving (504) a second input (108) comprising image data (110), indicative of an environment (112) wherein the lighting configuration task is to be implemented;- determining (506), using the first user input and the image data received, one or more sub-tasks (116.1, 116.2,... 116.N) that are associated with the implementation of the lighting configuration task in the environment; and- generating (508), based on the first user input, output data (118) indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment- generating (514) a multimodal prompt (120) including the image data (110), the first user input (105), and the output data, and- providing (516) said multimodal prompt to a natural language model (123).

2. The method of claim 1 further comprising:- identifying (510), for the one or more sub-tasks, module data (MD.1, MD.2) indicative of corresponding relevant software modules (MD.l, MD.2) that are associated with the respective sub-tasks, and wherein the output data (118) is further indicative of the identified module data.

3. The method of claim 1 or 2, further comprising:- identifying (512), for the one or more lighting units, lighting unit data (LD.l, LD.2) indicative of operation parameters and / or capabilities of the respectivelighting units, and wherein the output data (118) is further indicative of the identified lighting unit data.

4. The method of any of the preceding claims, wherein the image data (110) comprises an image file or a video file or a depth map or a floor plan drawing, or a three- dimensional representation of a room scene.

5. The method of any of the claims 2 to 4, further comprising- selecting, from a predetermined set of available example data, target example data indicative of examples of use of the relevant modules; and;- generating the multimodal prompt further including the target example data.

6. The method of any of claims 2 to 5, further comprising;- generating the multimodal prompt further including a high level textual description of one or more of the relevant modules.

7. The method of claim 6, wherein the high level textual description of a module comprises a description of input parameters required by the respective module and a description of module operation for obtaining an output parameter based on said input parameter, and / or comprises examples of suitable and / or unsuitable use of the respective module.

8. The method of claims 5 and 6 or 7, further comprising:- determining (516) a complexity measure of the first user input, and- including (518) the example data and / or the high level textual description in the multimodal prompt in dependence on the determined complexity measure.

9. The method of any of the preceding claims 1 to 8, further comprising:- generating (520), using the provided multimodal prompt and the natural language model, an output file indicative of the lighting configuration task applied on the environment indicated by the image data.

10. The method of claim 9, wherein the output file comprises a text file or latent space representation including control instructions for controlling the one or more lighting units.

11. The method of claim 10, further comprising, generating and providing a control signal indicative of the control instructions to the lighting arrangement for controlling the one or more lighting units in accordance with the lighting configuration task.

12. The method of any of the preceding claims 9 to 11, wherein the output file comprises an image file and / or a video file and / or neural radiance field and / or a light field indicative of the lighting configuration task applied to the environment.

13. Controller (100) for controlling implementation of a lighting configuration task (LT) for one or more lighting units (11, 12) of a lighting arrangement (10), the controller comprising- a first input unit (102) configured to receive a first user input (104) indicative of the lighting configuration task (LT);- a second input unit (106) configured to receive a second input (108) comprising image data (110), indicative of an environment (112) wherein the lighting configuration task is to be implemented;- an input processing unit (114) configured to determine, using the first user input and the image data received, one or more sub-tasks (116.1, 116.2,... 116.N) that are associated with the implementation of the lighting configuration task and to generate, based on the first user input, output data (118) indicative of one or more sub-tasks to be executed for implementing the lighting configuration task on the environment, generate a multimodal prompt (120) including the image data (110), the first user input (105), and the output data, and provide said multimodal prompt to a natural language model (123).

14. Computer program comprising instructions which, when executed by a controller in accordance with claim 13, cause the controller to perform the method of any of the preceding claim 1 to 12.

Citation Information

Patent Citations

  • Methods and systems for a user interface for illumination power, management, and control

    US11842194B2

  • Systems, methods, and apparatuses for distributing computational resources over a network of luminaires

    US20210298157A1