A semantic understanding-based method and system for controlling display lighting effects
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-14
AI Technical Summary
[0003]但是,现有方案在自然语言意图复杂、环境照度变化明显或多灯设备能力差异较大时,容易出现语义理解结果与实际灯效参数脱节的问题;同时,灯珠数量、颜色通道、色域、功率上限、供电状态、通信延迟和空间拓扑等设备约束未被统一纳入生成过程,导致生成的颜色、亮度、闪烁频率和相位参数可能不能稳定执行;此外,低置信语义解析结果仍可能触发过亮、过快闪烁或过饱和的灯效输出,影响视觉舒适度和多灯同步效果
1、通过将用户自然语言输入解析为语义槽位并生成语义控制令牌,使场景类型、情绪氛围、活动阶段、亮度偏好、颜色风格、动态节奏和眩光抑制级别能够进入后续灯效生成流程,减少仅依赖关键词或固定模板造成的意图表达损失。
Smart Images

Figure CN122579412A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent lighting control, and in particular to a method and system for controlling display lighting effects based on semantic understanding. Background Technology
[0002] With the development of smart lighting fixtures, ambient light strips, and multi-light linkage devices, users are increasingly describing their desired atmosphere through voice or text, such as for watching movies, parties, sleep aids, or gaming. Existing lighting control solutions typically rely on preset templates, fixed color schemes, or simple keyword matching, directly mapping user input to a limited range of lighting effect modes.
[0003] However, existing solutions are prone to discrepancies between semantic understanding results and actual lighting effect parameters when the natural language intent is complex, the ambient illuminance varies significantly, or the capabilities of multiple lamp devices differ greatly. At the same time, device constraints such as the number of lamps, color channels, color gamut, power limit, power supply status, communication latency, and spatial topology are not uniformly incorporated into the generation process, which may result in unstable execution of the generated color, brightness, flicker frequency, and phase parameters. In addition, low-confidence semantic parsing results may still trigger excessively bright, excessively fast flickering, or oversaturated lighting effect output, affecting visual comfort and the synchronization effect of multiple lamps.
[0004] Therefore, there is a need for a display lighting effect control method and system that can overcome the shortcomings of the existing technology. Summary of the Invention
[0005] One objective of this invention is to propose a display lighting effect control method and system based on semantic understanding. Addressing the problem in existing technologies that struggle to translate user natural language lighting effect intentions into executable lighting effect control schemes by combining luminaire capabilities, ambient light intensity, and communication constraints, this invention proposes a technical solution that generates semantic control tokens through a semantic parsing model, characterizes luminaire execution conditions through a device capability graph, and outputs candidate lighting effect sequences through a conditional lighting effect sequence generation model, which are then corrected by projection onto the feasible region. This invention improves lighting effect intention matching, visual comfort, and the stability of multi-lamp synchronization.
[0006] This invention provides a display lighting effect control method based on semantic understanding, including: S1. Obtain user natural language input, ambient illuminance sequence, optional music rhythm features, and device status parameters of multiple lamps. The device status parameters include the number of lamp beads, color channels, color gamut, power limit, power supply status, power level, installation location, communication delay, communication bandwidth, packet loss rate, and spatial topology. S2. The semantic parsing model trained by the user's natural language input outputs semantic slots corresponding to scene type, emotional atmosphere, activity stage, brightness preference, color style, dynamic rhythm and glare suppression level, and generates semantic control tokens based on the semantic slots. S3. Construct a device capability diagram based on the device status parameters, and encode the device capability diagram into a device capability diagram embedding; S4. Input the semantic control token, the embedded device capability map, the ambient illuminance sequence, and the rhythm input feature formed by the optional music rhythm feature or empty rhythm flag into the conditional lighting effect sequence generation model trained to generate candidate lighting effect sequences. The candidate lighting effect sequences include color parameters, brightness curves, flashing frequency, gradation mode, cycle period, and lamp group phase parameters. S5. Based on brightness comfort, power consumption, color gamut, number of LEDs, communication bandwidth, and synchronization delay constraints determined by communication delay and spatial topology, perform feasible domain projection correction on the candidate lighting effect sequence, and output lighting effect control frames with timestamps, key frame compressed packages, and local interpolation rules.
[0007] Optionally, S1 includes: Speech recognition is performed on voice input to obtain multiple candidate texts, while the original text is retained for text input, and the multiple candidate texts or the original text is used as the user's natural language input. Ambient illumination sequences are collected according to a unified timeline. Music rhythm features are extracted when audio input is present, and empty rhythm markers are generated when no audio input is present. The system reads the number of LEDs, color channels, color gamut, power limit, power supply status, power level, and installation location from the lighting control interface. It also obtains the communication delay (represented by the median round-trip delay within a sliding time window), the communication bandwidth (represented by the available load per unit time), and the packet loss rate through link detection. The system then generates a spatial topology based on the installation location.
[0008] Optionally, S2 includes: The multiple candidate texts or the original text, the historical preferences obtained from the historical control record, and the environmental noise features collected from the input environment are input into the semantic parsing model to obtain the semantic slot probability and the slot position confidence interval. The semantic control token is generated based on the semantic slot probability, slot position confidence interval, and historical preference. When the lower limit of the slot position signal is less than the preset signal threshold, the brightness range, flicker frequency range and saturation range are respectively limited to the preset comfort boundary; Furthermore, the semantic control token includes scene encoding, emotion encoding, stage encoding, brightness range, color style encoding, rhythm encoding, glare suppression level, flicker frequency range, saturation range, and confidence range; The lower bound of the slot position information is determined by the quantile value of the semantic slot probability and the text consistency score. The text consistency score of multiple candidate texts is determined by the semantic consistency among multiple candidate texts, and the text consistency score of the original text is determined by the consistency between the original text and the semantic slot prediction result. The preset comfort boundary includes a brightness boundary, a frequency boundary, and a saturation boundary. When the lower bound of the slot position information is less than a preset information threshold, the brightness range, flicker frequency range, and saturation range are respectively mapped into the preset comfort boundary.
[0009] Optionally, S3 includes: Each lamp is treated as a node, and spatial adjacency and communication link are treated as edges. Node features and edge features are generated based on the number of LEDs, color channels, power limit, power supply status, power consumption, installation location, communication delay, communication bandwidth, and packet loss rate. The node features and edge features are input into the device capability graph attention network, which outputs the lamp group weights and phase offsets. Synchronization priorities are generated based on the lamp group weights and the link stability component, power availability component, and spatial influence component obtained from the node features and edge features. The device capability map is embedded based on the lamp group weight, phase offset, and synchronization priority. Furthermore, the synchronization priority is determined based on the light group weight, communication delay, packet loss rate, power supply status, power consumption, and installation location. Specifically, it includes: normalizing the communication delay and packet loss rate into a link stability component, normalizing the power supply status and power consumption into a power supply availability component, mapping the installation location into a spatial influence component, and generating a synchronization priority based on the light group weight, the link stability component, the power supply availability component, and the spatial influence component.
[0010] Optionally, S4 includes: The semantic control token is mapped to a semantic vector, and the device capability graph is embedded and mapped to a capability vector. When a music rhythm feature is obtained, the music rhythm feature is encoded into a music rhythm vector; when no music rhythm feature is obtained, the empty rhythm flag is encoded into a default rhythm vector. Encode the ambient illumination sequence and the music rhythm vector or default rhythm vector into a temporal environment vector; The semantic vector, capability vector and temporal environment vector are input into the temporal decoder of the conditional lighting effect sequence generation model, and the temporal decoder outputs candidate lighting effect sequences at preset frame intervals. The color parameters in the candidate lighting effect sequence are represented by the channel values of the corresponding color channels, the brightness curve is represented by piecewise linear control points, and the phase parameters of the lighting group are represented by the phase offset relative to a unified clock.
[0011] Optionally, S5 includes: A feasible domain is established, consisting of brightness comfort, power consumption, color gamut, number of LEDs, communication bandwidth, and synchronization delay constraints determined by communication delay and spatial topology. The brightness comfort constraint is determined based on the ambient illuminance sequence, brightness preference, and glare suppression level, while the power consumption constraint is determined based on the power limit and the color parameters and brightness curves in the candidate lighting effect sequence. The projection objective function is composed of the semantic deviation between the projected lighting effect sequence and the semantic control token, the frame parameter deviation between the projected lighting effect sequence and the candidate lighting effect sequence, and the cross-light group synchronization deviation. Solve the projection objective function within the feasible region to obtain the lighting effect control frame that satisfies the physical execution conditions of the device; Furthermore, the keyframe compressed package is encoded by the color parameter difference, brightness curve control point difference, and phase offset difference between adjacent control frames, wherein the brightness curve control point is a piecewise linear control point of the brightness curve, and the phase offset is calculated from the phase parameters of the lamp group. The local interpolation rules include linear interpolation rules and easing interpolation rules executed according to a unified clock, and the key frame transmission period is selected according to the communication bandwidth, and the number of local interpolation frames is selected according to the synchronization delay constraint.
[0012] On the other hand, the present invention also provides a display lighting effect control system based on semantic understanding, comprising: The input acquisition module is used to acquire user natural language input, ambient illuminance sequence, optional music rhythm features, and device status parameters of multiple lamps, and generates an empty rhythm flag when music rhythm features are not acquired; the semantic token generation module is used to generate semantic slots and semantic control tokens through a trained semantic parsing model; the device capability encoding module is used to construct a device capability map and generate a device capability map embedding; the lighting effect sequence generation module is used to generate candidate lighting effect sequences; and the feasible region projection module is used to perform feasible region projection correction on the candidate lighting effect sequences and output lighting effect control frames with timestamps, keyframe compressed packages, and local interpolation rules.
[0013] The beneficial effects of this invention are: 1. By parsing user natural language input into semantic slots and generating semantic control tokens, scene type, emotional atmosphere, activity stage, brightness preference, color style, dynamic rhythm and glare suppression level can be included in the subsequent lighting effect generation process, reducing the loss of intent expression caused by relying solely on keywords or fixed templates.
[0014] 2. By constructing a device capability map and embedding it using the number of LEDs, color channels, power limit, power consumption, installation location, communication delay, and packet loss rate, the candidate lighting effect sequence incorporates the capability differences of multiple LED devices during the generation stage, providing a structured basis for subsequent synchronous control and phase allocation.
[0015] 3. By performing feasible domain projection correction on the candidate lighting effect sequence before execution, and outputting lighting effect control frames with timestamps, key frame compressed packages, and local interpolation rules, constraints such as brightness comfort, power consumption, color gamut, communication bandwidth, and synchronization delay can constrain the final output, thereby improving the accessibility of lighting effect execution and the stability of multi-light linkage. Attached Figure Description
[0016] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a display lighting effect control method based on semantic understanding.
[0017] Figure 2 This is a flowchart of the feasible region projection correction step S5 of the present invention. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] refer to Figures 1-2 A semantic understanding-based method for controlling display lighting effects includes: S1. Obtain user natural language input, ambient illuminance sequence, optional music rhythm features, and device status parameters of multiple lamps. The device status parameters include the number of lamp beads, color channels, color gamut, power limit, power supply status, power level, installation location, communication delay, communication bandwidth, packet loss rate, and spatial topology. S2. The semantic parsing model trained by the user's natural language input outputs semantic slots corresponding to scene type, emotional atmosphere, activity stage, brightness preference, color style, dynamic rhythm and glare suppression level, and generates semantic control tokens based on the semantic slots. S3. Construct a device capability diagram based on the device status parameters, and encode the device capability diagram into a device capability diagram embedding; S4. Input the semantic control token, the embedded device capability map, the ambient illuminance sequence, and the rhythm input feature formed by the optional music rhythm feature or empty rhythm flag into the conditional lighting effect sequence generation model trained to generate candidate lighting effect sequences. The candidate lighting effect sequences include color parameters, brightness curves, flashing frequency, gradation mode, cycle period, and lamp group phase parameters. S5. Based on brightness comfort, power consumption, color gamut, number of LEDs, communication bandwidth, and synchronization delay constraints determined by communication delay and spatial topology, perform feasible domain projection correction on the candidate lighting effect sequence, and output lighting effect control frames with timestamps, key frame compressed packages, and local interpolation rules.
[0020] In this specific embodiment, S1 includes: A unified timeline is established using the local gateway as the time reference node. ,in Indicates the first Each sampling time is measured in milliseconds. This indicates the total number of sampling points. The gateway uses NTP time synchronization to ensure that the clock deviation with the control interface of each lamp does not exceed 5ms, and sets the sampling interval to [value missing]. To establish a unified time granularity for subsequent fusion; When the input is speech, the gateway performs speech recognition on the mono audio stream captured by the microphone to obtain multiple candidate texts. The audio stream sampling rate is set to 16kHz and the quantization bit width is set to 16 bits. The speech recognition model adopts an end-to-end structure of a Conformer encoder and a Transformer decoder. The Conformer encoder has 12 layers, 8 heads per layer for multi-head self-attention, 80-dimensional Mel filter bank features for the front-end acoustic features, a frame length of 25ms, and a frame shift of 10ms. The decoder uses beam search and sets the beam width to 5 to output 5 candidate texts and their posterior confidence. The gateway sorts these 5 candidate texts in descending order of confidence to form a candidate text set and uses it as the user's natural language input. When the input is text, the gateway retains the original text and uses the original text as the user's natural language input, so that the output form of the user's natural language input in this step is either "a single piece of original text" or "a set of candidate texts containing 5 candidate texts"; The gateway collects ambient illuminance sequences from the ambient light sensor according to the unified time axis and performs time alignment. The ambient light sensor outputs illuminance in lx. The gateway performs time alignment at each sampling time. Read the illuminance values and form an ambient illuminance sequence, so that the ambient illuminance sequence can characterize the ambient light intensity that changes over time and participate in the subsequent luminance comfort constraint calculation; When there is audio input and the audio input contains music content, the gateway extracts music rhythm features from the same audio stream and aligns them with a unified time axis. The music rhythm features are calculated from the initial intensity envelope obtained by short-time Fourier transform. The window length of the short-time Fourier transform is set to 1024 points and the window function is a Hann window. The frame shift is set to 512 points. The gateway performs beat tracking based on the initial intensity envelope to obtain the beat intensity, beat phase and beat rate at each sampling moment and form a music rhythm feature sequence. When there is no audio input or the audio input does not contain music content, the gateway generates an empty rhythm flag. And set it to 1, while setting the music rhythm feature sequence to an empty sequence to ensure the structural consistency of subsequent model inputs, where This indicates the existence of effective musical rhythmic features and This indicates the absence of a valid musical rhythmic feature; The gateway reads the device status parameters of multiple lamps one by one through the lamp control interface and encapsulates them in a standardized manner. The lamp control interface adopts a local control protocol for lamps and returns structured fields. The device status parameters for each lamp include at least the number of LEDs, color channels, color gamut, power limit, power supply status, power level, and installation location. The installation location is represented by three-dimensional coordinates in the room coordinate system and the unit is meters, so that it can be directly used for spatial topology generation in this step. The gateway obtains communication latency, bandwidth, and packet loss rate through link probing. Link probing involves periodically sending fixed-size probe packets on the lighting fixture's communication link and waiting for acknowledgments. The probe packet payload is set to 64 bytes, the sending period is set to 100ms, and the sliding window length is set to 5s, collecting data within each sliding window. Round-trip delay sample Used to calculate communication delay, communication delay Determine using the following formula: ; in This represents the communication delay within the sliding time window, expressed in milliseconds. Indicates the first Round-trip delay samples for secondary link probing, in milliseconds. This indicates the number of detected samples within the sliding time window. This represents the operation of taking the median value of the input sample set; The gateway counts the number of successfully confirmed bytes of probe packets and service control packets in each 1-second statistical period, and divides this number by the statistical period to obtain the communication bandwidth. The communication bandwidth is expressed as available load per unit time and the unit is 1. Meanwhile, the packet loss rate is calculated as the ratio of the number of unacknowledged messages to the total number of messages sent within the statistical period. ; The gateway generates a spatial topology based on the installation location and uses it together with the communication link to construct the subsequent device capability map. The spatial topology is expressed as an undirected graph with lamps as nodes and spatial adjacency as edges. Spatial adjacency is determined by calculating the Euclidean distance between the installation locations of any two lamps and comparing it with a threshold of 2m. When the distance between two lamps is not greater than 2m, an edge is established in the spatial topology and the distance of the edge is recorded as an edge attribute. Thus, at the end of this step, the user's natural language input, ambient illuminance sequence, music rhythm features or empty rhythm flag, and a set of device status parameters containing device capability fields, link fields, and spatial topology fields are output.
[0021] In this specific embodiment, S2 includes: The gateway receives natural language input from users and organizes it into a collection of text sequences. ,in Indicates the first A sequence of texts and This represents the number of text sequences, taken when the user's natural language input is a set of candidate texts. and Sort by descending order of posterior confidence score for speech recognition; when the user's natural language input is raw text, select... and The original text; The gateway constructs historical preferences from historical control records and forms a historical preference vector. The historical control records are taken from the control logs of the most recent 30 days, including the scene type, mood atmosphere, brightness setting, color style setting, and dynamic rhythm setting for each control. The gateway aggregates the control logs according to time decay weights to obtain the brightness preference center value, color style preference distribution, and rhythm preference distribution. The aggregation result is then encoded into a 32-dimensional floating-point vector as... ; The gateway collects environmental noise features from the input environment and forms an environmental noise vector. The environmental noise characteristics are calculated from the microphone signal within the most recent second and include A-weighted sound pressure level, short-time energy variance, and voice activity duty cycle. The gateway normalizes these environmental noise characteristics and encodes them into an 8-dimensional floating-point vector. ; The gateway will and The semantic parsing model trained is used to obtain semantic slot probabilities and slot position confidence intervals. The semantic parsing model consists of a text encoder and a multi-task prediction head. The text encoder is a 12-layer Transformer structure with a hidden dimension of 768, 12 attention heads, a maximum sequence length of 64, and uses WordPiece word segmentation with [CLS] representing the semantic vector of the entire sentence. A history preference vector is also included. With environmental noise vector After being projected onto a 768-dimensional network through two fully connected layers, the fused semantic representation is added to the [CLS] vector; The multi-task prediction head includes a discrete slot classification head and a continuous slot regression head. The discrete slot classification head outputs the category probability distribution for scene type, mood atmosphere, activity stage, color style, dynamic rhythm, and glare suppression level, and takes the category with the highest probability as the slot prediction result. The continuous slot regression head outputs the predicted values of the upper and lower bounds of the brightness range, flicker frequency range, and saturation range, respectively, with the brightness range being within... The normalized luminance range is represented in Hz, the flicker frequency range in Hz, and the saturation range in the HSV color space. Intradomain representation; The semantic parsing model is trained offline on a dataset with slot annotations. The training objective is a weighted sum of the cross-entropy loss of discrete slots and the Huber loss of continuous slots. The optimizer uses AdamW with a learning rate set to [value missing]. Batch size is set to 64, number of training rounds is set to 10, and Dropout ratio is set to 0.2. The gateway performs each text sequence during the inference phase. implement Monte Carlo Dropout forward computation is performed to form a slot probability sampling set, and a slot position confidence interval is output for each semantic slot. The confidence interval of a discrete slot is taken as the maximum probability of the class in the sampling set. quantiles and The quantile is used as the lower and upper bound, and the confidence interval for consecutive slots is taken from the sample set of predicted values of the upper and lower bounds of the interval. quantiles and The quantile serves as the lower and upper bounds; The gateway further calculates a lower bound for the slot position information for each semantic slot and uses it for uncertainty gating. Determine using the following formula: ; in This indicates a semantic slot index that covers scene type, emotional atmosphere, activity stage, brightness preference, color style, dynamic rhythm, and glare suppression level. Indicates slot The confidence lower bound and the range of values is Indicates the first Text sequence Slots obtained after inputting the semantic parsing model The confidence level corresponding to the maximum class probability or interval prediction, and the value range is: Indicates the number of text entries in the sequence. Indicates taking from the input set The operation of quantile values and This represents the text consistency score and its value range is [value range missing]. ,when hour Depend on The average cosine similarity of the sentence vectors obtained by the text encoder is then linearly mapped to... Get, when hour From the original text The cosine similarity between the sentence vector and the sentence vector of the "slot description text" generated from the slot prediction results is linearly mapped to... The obtained slot description text is formed by splicing together scene type tags, mood atmosphere tags, activity stage tags, color style tags, and dynamic rhythm tags in a fixed order; The gateway is based on semantic slot prediction results, slot position confidence intervals, and historical preference vectors. Generate a semantic control token, which is represented by structured fields and includes scenario encoding. Emotion coding Stage coding Brightness range Color style coding Rhythm coding Glare suppression level flicker frequency range saturation range With confidence interval set ,in and All of them use a pre-defined lookup table mapping to map the corresponding discrete slot category to a 16-bit unsigned integer code. The value is an integer from 0 to 3, and the larger the value, the stronger the glare suppression. Output from the continuous slot regression head and compared with the historical preference vector The fusion of the mid-brightness preference center values is weighted and the fusion weights are fixed at 0.7 for semantic prediction and 0.7 for historical preference. and The token is output from the continuous slot return header and written directly to it. The gateway sets the preset threshold to When the slot position corresponding to the brightness range, flicker frequency range, or saturation range is at its lower bound... Less than Comfort boundary mapping is performed, and the brightness range, flicker frequency range, and saturation range are each restricted within preset comfort boundaries, wherein the preset comfort boundaries are fixedly set as brightness boundaries. Frequency boundary With saturation boundary The restriction method is to and The upper and lower bounds are truncated according to the corresponding boundaries while keeping the order of the upper and lower bounds unchanged. Thus, at the end of this step, a semantic control token containing the encoding field, the interval field, and the confidence field is output, and the brightness, flicker, and saturation under the low-confidence semantic parsing result meet the visual comfort constraints.
[0022] In this specific embodiment, S3 includes: The gateway constructs a device capability diagram based on the status parameters of multiple lighting devices. ,in Represents a set of lighting fixture nodes with the number of nodes being [0, 1]. Each node corresponds to a lamp and is uniquely identified by its lamp chip count, color channel, color gamut, power limit, power supply status, battery level, and installation location. The set of edges is represented by the combination of spatial adjacency edges and communication link edges. Spatial adjacency edges are directly given by the spatial topology and their attributes include the Euclidean distance between the two lamp installation locations. Communication link edges are directly given by the link detection results and their attributes include communication delay, communication bandwidth, and packet loss rate. The gateway constructs a node feature vector for each node and an edge feature vector for each side. The node feature vectors are determined by the number of LEDs in intervals. The linearly normalized values and the number of color channels are divided into intervals. Linearly normalized values and color gamut by interval Normalized coverage and power upper limit by interval Linearly normalized values, binary encoding of power supply status, and power consumption by interval The linearly normalized values and the two-dimensional coordinates of the installation location according to room dimensions The coordinate components after linear normalization are concatenated to form the edge feature vector, which is formed by spatial distance according to intervals. Linearly normalized values and communication delays by interval The linearly normalized values and the communication bandwidth are divided into intervals. The linearly normalized values and the packet loss rate in intervals The results are formed by concatenating linearly normalized values. All normalizations use linear mapping and the results are truncated to [a certain value]. ; The gateway inputs node features and edge features into the device capability graph attention network to obtain the light group weight and phase offset of each node. The device capability graph attention network is a graph attention network with edge features and contains two message passing layers and three output heads. The message passing layer uses eight-head attention parallel computation and each head has a hidden dimension of 32, thus obtaining a 64-dimensional graph representation for each node. The attention score is jointly determined by the graph representation and edge features of adjacent nodes, and the activation function is LeakyReLU with a fixed negative slope of 0.2. The inter-layer dropout ratio is fixed at 0.1. The first output head is the lamp group weight head, and the 64-dimensional graph representation of each node is input into a two-layer fully connected network to obtain 4-dimensional lamp group weight logarithmic values. Then, the Softmax function is used to obtain the lamp group weight vector. To indicate the first The lamps belong to the soft-weighted distribution of 4 lamp groups. The second output head is the phase offset head, and the 64-dimensional graph representation of each node is input into a two-layer fully connected network to obtain the phase offset. and through Limit it to The third output header is used to generate intermediate synchronization features and output a node-level scalar for synchronization priority calculation, representing the phase offset relative to the uniform clock. The gateway generates three components required for synchronization priority based on the aforementioned logic and determines the synchronization priority together with the light group weights, including the link stability component. Due to nodes The average of the normalized communication latency and packet loss rate of the connected communication links, then inverted, yields the result, thus allowing low latency and low packet loss to correspond to a larger... Power availability component Determined by both power supply status and power consumption, and when supplied with mains power... Set to 1 and use the normalized charge value as the battery power when powered. Spatial influence components Obtained by the installation location mapping and through the compute node Distance to the geometric center of the room and by interval The result is obtained by inverting the linear normalization, which makes the space more central and the corresponding position larger. Meanwhile, in order to incorporate the weight of the light groups into the synchronous scheduling, the gateway takes... The largest component is used as a node Light group weight scalar To indicate the strength of its membership in the main light group; The gateway uses a fixed coefficient to... and Merge to generate synchronization priority And calculate according to the following formula: ; in Indicates the first The synchronization priority of each lamp and its value range are: The Sigmoid function is used to compress the result of a linear combination to... Represents the weight vector of the lamp group The maximum component is used to obtain the scalar weight of the lamp group, and its value range is [value range missing]. Represents the link stability component and its value range is Represents the power availability component and its value range is This represents the spatial influence component and its value range is... to For a fixed fusion coefficient and take respectively , ; Gateway based on light group weight vector Phase offset Synchronization priority Embedded device capability map The device capability graph embedding is formed using a global pooling method, and the 64-dimensional graph representation of each node is divided into... Weighted summation yields a 64-dimensional synchronous embedding, and each node's 64-dimensional graph representation is then processed according to... Weighted summation yields a 64-dimensional grouped embedding, which is then concatenated in a fixed order to obtain a 128-dimensional vector. and will and As with The consistent binding of the device-side structured output is used for subsequent lamp group phase allocation and synchronization control generation.
[0023] In this specific embodiment, S4 includes: The gateway receives the semantic control token and maps it to a semantic vector. Scene encoding in semantic control token Emotion coding Stage coding Color style coding With rhythm coding Each vector is mapped to a 64-dimensional vector using an independent trainable embedding table and then concatenated in a fixed order to form a 320-dimensional discrete semantic sub-vector. Glare suppression level After being encoded with a length of 4 using a one-hot encoding method, it is mapped to a 16-dimensional vector through a fully connected layer, representing the brightness range. flicker frequency range and saturation range Six scalars, including upper and lower bounds, are arranged in a fixed order to form a continuous semantic subvector, which is then mapped to a 128-dimensional vector through a two-layer fully connected network. The layer width of the two fully connected networks is fixed. Furthermore, the activation function is GELU, and Dropout is applied after each layer with a fixed ratio of 0.1. The set of confidence intervals... The lower bounds of the slot position information are taken and formed into a confidence vector of length 7. After passing through a fully connected layer, it is mapped to a 32-dimensional vector. The gateway concatenates the 320-dimensional discrete semantic sub-vector, the 16-dimensional glare sub-vector, the 128-dimensional continuous semantic sub-vector, and the 32-dimensional confidence sub-vector, and then performs a linear mapping to obtain a 256-dimensional semantic vector. ; Gateway receiving device capability diagram embedding And mapped to capability vectors ,in A 128-dimensional vector is mapped to a 256-dimensional capability vector through a two-layer fully connected network. The layer width of the two fully connected networks is fixed at 1. The activation function is GELU, and Dropout is applied between layers with a fixed ratio of 0.1. The gateway records the ambient illuminance sequence according to a unified time axis as follows: ,in Indicates the sampling time The ambient illuminance is measured in lx, and rhythm input features are constructed based on musical rhythm characteristics or empty rhythm markers. When musical rhythm features exist, the beat intensity, beat phase, and beat rate at each sampling moment are combined to form a 3D rhythm feature, which is then normalized according to fixed intervals and mapped to a 32-dimensional musical rhythm vector through a fully connected layer. Rhythm markers in the air When Inputting an embedding table of length 2 yields a 32-dimensional default rhythm vector. This results in each sampling moment having a 32-dimensional rhythm vector; The gateway will and The data is concatenated to form a 33-dimensional temporal input, which is then input into a temporal environment encoder to obtain a temporal environment vector. The temporal environment encoder is a 2-layer GRU with a fixed hidden dimension of 128 and a unidirectional structure to maintain causality. The GRU output is linearly mapped to obtain a 128-dimensional temporal environment vector. It is then combined with the sinusoidal position code of the sampling time index to preserve the time position information; The gateway will use semantic vectors Capability Vector With temporal environment vector Input the time-series decoder of the conditional lighting effect sequence generation model obtained from the training and execute it according to the preset frame interval. Candidate lighting effect sequences are generated frame by frame. The temporal decoder is a 6-layer Transformer decoder with a fixed hidden dimension of 512, a fixed number of attention heads of 8, and a fixed feedforward layer dimension of 2048. An autoregressive generation method is used. Frame number The frame embedding is obtained by passing the frame vector of the output frame through a linear layer. As the autoregressive input of the current frame and combined with and Together, they determine the output of the current frame, specifically expressed by the following formula: ; in Indicates the first The candidate lighting effect frame parameter vector of the frame. Indicates by parameters Deterministic timing decoder mapping, Indicates by the first The frame embedding is obtained by passing the frame candidate lighting effect frame parameter vector through a linear layer. Represents a 256-dimensional semantic vector. Represents a 256-dimensional capability vector. Indicates the first A 128-dimensional temporal environment vector for each sampling time point; The gateway will The parameters are unpacked into a set of candidate lighting effect sequences according to fixed fields, ensuring that the field semantics are consistent with subsequent steps. The color parameters are represented by channel values of up to 5 color channels and constrained using a Sigmoid function. Then, it is quantized into 8-bit channel values according to the lighting protocol. The brightness curve is represented by piecewise linear control points, and the output is displayed in each cycle. Each control point is stored in pairs with its normalized time coordinate and normalized brightness value to ensure the reconstruction of a piecewise linear curve. The flicker frequency is limited to non-negative values using Softplus and expressed in Hz with an upper limit clipped to 20Hz. The fading mode is represented by three discrete codes and obtained from the highest probability category output by Softmax. The cycle period is output using Softplus and expressed in seconds with a clipped value. The phase parameters of the lamp group are expressed as the phase offset relative to a uniform clock and for Each light group outputs a phase offset and... Scaling limit This allows for subsequent multi-lamp synchronization control by combining phase offset and synchronization priority, thereby forming a candidate lighting effect sequence arranged along a unified time axis at the end of S4 and outputting it to S5 for feasible domain projection correction.
[0024] In this specific embodiment, S5 includes: The gateway receives the candidate lighting effect sequence and represents it according to a unified timeline. ,in Indicates timestamp as The The candidate lighting effect frame parameter vector includes color parameters, brightness curve control points, flashing frequency, fading mode, cycle period, and light group phase parameters. and All are determined by S1; The gateway establishes a feasible domain composed of multiple constraints and performs projection correction for each frame. These constraints include brightness comfort constraints, power consumption constraints, color gamut constraints, LED quantity constraints, communication bandwidth constraints, and synchronization delay constraints. The brightness comfort constraints are based on the ambient illuminance sequence. Brightness range in semantic control tokens With glare suppression level Generate the upper and lower bounds of the allowed brightness at each time step and truncate the candidate brightness values. Indicates time Ambient illuminance and the unit is and The normalized brightness upper and lower bounds output from step S2 and after uncertainty gating, with a value range of [value range missing]. , The glare suppression level output in step S2 has a value of 0 to 3 and is used to tighten the upper limit of allowable brightness under the same ambient illuminance. The power consumption constraint is based on the upper limit of power for each lamp obtained in step S1 and The color parameters and brightness curves in the system estimate the instantaneous power consumption of each frame and perform scaling correction. The gateway pre-stores the channel power consumption coefficient for each lamp and substitutes the color channel value and brightness value into the coefficient in each frame to obtain the power consumption estimate. When the power consumption estimate exceeds the power limit, all color channel values and brightness values of the lamp in that frame are scaled by the same proportion until the power consumption limit is met. Based on the color gamut information obtained in step S1, the color parameters of each lamp are projected within the color gamut. The gateway pre-stores the color parameters of each lamp from the color channel space to the CIE color space. After calibrating the color space and mapping the candidate colors to chromaticity coordinates, it is determined whether they fall within the color gamut polygon of the luminaire. If they do not fall within it, the "projection to the nearest point of the color gamut polygon" is used to obtain the corrected chromaticity coordinates and then the colors are inversely mapped back to the color channel values to ensure that the output color can be achieved by the actual light emitted by the luminaire. The LED number constraint is based on the LED number obtained in step S1 to limit the dynamic change range to avoid obvious discrete flickering in devices with a small number of LEDs under high-speed changes. The gateway determines the maximum brightness step and maximum color step of each lamp based on its LED number and performs truncation on the brightness difference and color channel difference between adjacent frames, thereby ensuring that the changes between adjacent frames can be smoothly presented under a given number of LEDs and PWM resolution. The communication bandwidth constraint is based on the communication bandwidth obtained in step S1. To limit the amount of data sent and determine the keyframe transmission period, the gateway first generates an estimate of the byte size of the "candidate keyframe compressed package" using differential encoding and then uses the available link load. of As a control budget, the keyframe sending period is then set to The key frame transmission period is gradually increased in multiples of the specified number of units and the standard of "the average transmission rate of the compressed packet does not exceed the control budget" until the bandwidth constraint is met. The synchronization delay constraint determines the allowable error of cross-lamp group synchronization based on the communication delay and spatial topology obtained in step S1 and corrects the phase parameters of the lamp group. The gateway maps the communication delay to the predicted one-way delay from each lamp to the gateway and selects the synchronization reference lamp group for generating control frames in combination with the spatial topology. Then, the phase parameters of other lamp groups are jointly corrected with the phase offset output in step S3 so that the execution time deviation of any two lamp groups under the same clock does not exceed 20ms and the number of missing frames corresponding to the deviation is used as the number of local frames for subsequent interpolation frames. To obtain semantic meaning within the feasible domain Figure 1 To obtain a final sequence that deviates little from the candidate sequence and simultaneously satisfies multi-lamp synchronization, the gateway constructs a projection objective function and solves it using projection gradient descent. The projection objective function is: ; in Indicates timestamp as The After frame projection, the lighting effect control frame parameter vector and the field definitions are the same as those in the image. Consistent Indicates the first Frame candidate lighting effect frame parameter vector The L2 norm is used to measure frame parameter deviation. It represents a semantic control token and includes scene encoding, emotion encoding, stage encoding, brightness range, color style encoding, rhythm encoding, glare suppression level, flicker frequency range, saturation range, and confidence interval. It represents a device capability diagram and includes node capabilities and link attributes. Indicate semantic deviation and through judgment Do the brightness, flicker frequency, and saturation values fall within the range? Given an interval and applying a squared penalty to the out-of-bounds portion, we obtain... Indicate the synchronization deviation across lamp groups and calculate it. The expected execution deviation of each lamp group's phase under the link delay conditions in the equipment capability diagram is obtained by applying a squared penalty to the portion exceeding the synchronization allowable error. These are semantic deviation weights, frame parameter deviation weights, and synchronization deviation weights, respectively, and are fixed at... , ; Gateway As the initial value, 20 rounds of projection gradient descent iterations are performed with a fixed step size of 0.05. In each iteration, the gradient of the objective function is first calculated and updated to obtain a temporary sequence. Then, the projection operator is executed sequentially according to the above-mentioned constraints on brightness comfort, power consumption, color gamut, number of LEDs, and synchronization delay to pull the temporary sequence back into the feasible region, thereby obtaining the lighting effect control frame that meets the physical execution conditions of the device. Add a timestamp to each frame. Create lighting effect control frames with timestamps; The gateway encodes the color parameter difference, brightness curve control point difference, and phase offset difference between adjacent control frames into a keyframe compressed packet. The color parameter difference is calculated and quantized into a signed 8-bit integer according to the channel order. The brightness curve control point difference is calculated and quantized into a signed 12-bit integer according to the control point number and coordinate component order. The phase offset difference is converted from the lamp group phase parameters and quantized into a signed 12-bit integer. The keyframe compressed packet adopts the structure of "the first keyframe stores the absolute value + subsequent keyframes store the difference value" and writes the lamp group identifier, frame number range, quantization scale, and CRC check in the packet header for reconstruction at the receiving end. The gateway simultaneously generates local interpolation rules and associates them with the keyframe compressed package for distribution. The local interpolation rules include linear interpolation rules executed according to a unified clock and easing interpolation rules, and each keyframe interval is marked with a 2-bit rule code. The linear interpolation rules perform a proportional transition of color parameters and brightness values according to the time ratio within the keyframe interval. The easing interpolation rules use a preset 256-point cubic easing lookup table to perform non-linear remapping of the time ratio within the keyframe interval before performing a proportional transition to reduce the abruptness. The gateway determines the keyframe transmission period based on the communication bandwidth constraint and determines the number of local supplementary frames based on the synchronization delay constraint. Thus, at the end of step S5, the gateway outputs a timestamped lighting effect control frame, a keyframe compressed package, and local interpolation rules.
[0025] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
[0026] This invention enables the continuous use of natural language intent, device execution capabilities, and environmental conditions in the same lighting effect generation chain through the joint processing of semantic control tokens, device capability graph embedding, ambient illuminance sequence, and music rhythm features. It can transform the user's description of the atmosphere into executable controllable content such as color, brightness curve, flashing frequency, gradation mode, cycle period, and light group phase parameters.
[0027] This invention sets up an uncertainty gating unit, a device capability graph attention network, and a feasible domain projection layer, so that the brightness, flicker frequency, and saturation under low-confidence semantic results are constrained by the comfort boundary, and the weight, phase offset, and synchronization priority of multi-lamp devices participate in the generation of control frames, thereby better achieving intent matching, visual comfort, and multi-lamp synchronization stability.
Claims
1. A display lighting effect control method based on semantic understanding, characterized in that, include: S1. Obtain user natural language input, ambient illuminance sequence, optional music rhythm features, and device status parameters of multiple lamps. Device status parameters include number of LEDs, color channels, color gamut, power limit, power supply status, power level, installation location, communication delay, communication bandwidth, packet loss rate, and spatial topology. S2. Input the user's natural language input into the semantic parsing model trained to output semantic slots corresponding to scene type, emotional atmosphere, activity stage, brightness preference, color style, dynamic rhythm and glare suppression level, and generate semantic control tokens based on the semantic slots. S3. Construct a device capability diagram based on device status parameters, and encode the device capability diagram into a device capability diagram embedding; S4. Input the semantic control token, device capability map embedding, ambient illuminance sequence, and rhythm input features formed by optional music rhythm features or empty rhythm markers into the conditional lighting effect sequence generation model trained to generate candidate lighting effect sequences. The candidate lighting effect sequences include color parameters, brightness curves, flashing frequency, gradation mode, cycle period, and lamp group phase parameters. S5. Based on brightness comfort, power consumption, color gamut, number of LEDs, communication bandwidth, and synchronization delay constraints determined by communication delay and spatial topology, perform feasible domain projection correction on the candidate lighting effect sequence, and output lighting effect control frames with timestamps, key frame compressed packages, and local interpolation rules.
2. The display lighting effect control method based on semantic understanding according to claim 1, characterized in that, S1 includes: Speech recognition is performed on voice input to obtain multiple candidate texts, while the original text is retained for text input, and the multiple candidate texts or the original text is used as the user's natural language input. Ambient illumination sequences are collected according to a unified timeline. Music rhythm features are extracted when audio input is present, and empty rhythm markers are generated when no audio input is present. The system reads the number of LEDs, color channels, color gamut, power limit, power supply status, power level, and installation location from the lighting control interface. It also obtains the communication delay (represented by the median round-trip delay within a sliding time window), the communication bandwidth (represented by the available load per unit time), and the packet loss rate through link detection. The system then generates a spatial topology based on the installation location.
3. The display lighting effect control method based on semantic understanding according to claim 2, characterized in that, S2 includes: The multiple candidate texts or the original text, the historical preferences obtained from the historical control record, and the environmental noise features collected from the input environment are input into the semantic parsing model to obtain the semantic slot probability and the slot position confidence interval. The semantic control token is generated based on the semantic slot probability, slot position information interval, and historical preferences; when the lower bound of the slot position information is less than the preset information threshold, the brightness interval, flicker frequency interval, and saturation interval are respectively restricted within the preset comfort boundary.
4. The display lighting effect control method based on semantic understanding according to claim 2, characterized in that, S3 includes: taking each lamp as a node and spatial adjacency and communication link as edges, generating node features and edge features based on the number of LEDs, color channels, power limit, power supply status, power consumption, installation location, communication delay, communication bandwidth, and packet loss rate; inputting the node features and edge features into a device capability graph attention network, outputting lamp group weights and phase offsets, and generating synchronization priorities based on the lamp group weights and the link stability component, power supply availability component, and spatial influence component obtained from the node features and edge features; and generating the device capability graph embedding based on the lamp group weights, phase offsets, and synchronization priorities.
5. The display lighting effect control method based on semantic understanding according to claim 1, characterized in that, S4 includes: mapping the semantic control token to a semantic vector, and embedding the device capability map to a capability vector; encoding the music rhythm feature into a music rhythm vector when the music rhythm feature is obtained, and encoding the empty rhythm flag into a default rhythm vector when the music rhythm feature is not obtained; encoding the ambient illuminance sequence and the music rhythm vector or the default rhythm vector into a temporal environment vector; inputting the semantic vector, capability vector and temporal environment vector into the temporal decoder of the conditional lighting effect sequence generation model, and having the temporal decoder output candidate lighting effect sequences at preset frame intervals; the color parameters in the candidate lighting effect sequences are represented by the channel values of the corresponding color channels, the brightness curves are represented by piecewise linear control points, and the lamp group phase parameters are represented by the phase offset relative to a unified clock.
6. The display lighting effect control method based on semantic understanding according to claim 1, characterized in that, S5 includes: establishing a feasible region consisting of brightness comfort, power consumption, color gamut, number of LEDs, communication bandwidth, and synchronization delay constraints determined by communication delay and spatial topology, wherein the brightness comfort constraint is determined based on the ambient illuminance sequence, brightness preference, and glare suppression level, and the power consumption constraint is determined based on the power limit and the color parameters and brightness curves in the candidate lighting effect sequence; constructing a projection objective function using the semantic deviation between the projected lighting effect sequence and the semantic control token, the frame parameter deviation between the projected lighting effect sequence and the candidate lighting effect sequence, and the synchronization deviation across lighting groups; solving the projection objective function within the feasible region to obtain a lighting effect control frame that satisfies the physical execution conditions of the device.
7. The display lighting effect control method based on semantic understanding according to claim 3, characterized in that, The semantic control token includes scene encoding, emotion encoding, stage encoding, brightness range, color style encoding, rhythm encoding, glare suppression level, flicker frequency range, saturation range, and confidence range. The lower bound of the slot position confidence is determined by the quantile value of the semantic slot probability and the text consistency score. The text consistency score of multiple candidate texts is determined by the semantic consistency between multiple candidate texts, and the text consistency score of the original text is determined by the consistency between the original text and the semantic slot prediction result. The preset comfort boundary includes a brightness boundary, a frequency boundary, and a saturation boundary. When the lower bound of the slot position confidence is less than the preset confidence threshold, the brightness range, flicker frequency range, and saturation range are mapped to the preset comfort boundary, respectively.
8. The display lighting effect control method based on semantic understanding according to claim 4, characterized in that, The synchronization priority is determined based on the light group weight, communication delay, packet loss rate, power supply status, power consumption, and installation location. Specifically, it includes: normalizing the communication delay and packet loss rate into a link stability component, normalizing the power supply status and power consumption into a power supply availability component, mapping the installation location into a spatial influence component, and generating a synchronization priority based on the light group weight, the link stability component, the power supply availability component, and the spatial influence component.
9. The display lighting effect control method based on semantic understanding according to claim 6, characterized in that, The keyframe compressed package is encoded by the color parameter difference, brightness curve control point difference, and phase offset difference between adjacent control frames. The brightness curve control points are piecewise linear control points of the brightness curve, and the phase offset is calculated from the phase parameters of the lamp group. The local interpolation rules include linear interpolation rules executed according to a unified clock and easing interpolation rules. The keyframe transmission period is selected according to the communication bandwidth, and the number of local supplementary frames is selected according to the synchronization delay constraint.
10. A semantic understanding-based display lighting effect control system, used to execute the semantic understanding-based display lighting effect control method according to any one of claims 1 to 9, characterized in that, include: The input acquisition module is used to acquire user natural language input, ambient illuminance sequence, optional music rhythm features and device status parameters of multiple lamps, and generate an empty rhythm flag when the music rhythm features are not acquired; the semantic token generation module is used to generate semantic slots and semantic control tokens through the trained semantic parsing model; the device capability encoding module is used to construct a device capability map and generate a device capability map embedding. The lighting effect sequence generation module is used to generate candidate lighting effect sequences; The feasible region projection module is used to perform feasible region projection correction on the candidate lighting effect sequence and output lighting effect control frames with timestamps, key frame compressed packages, and local interpolation rules.