Light control method and system, controller and storage medium

Through real-time data acquisition and reinforcement learning model, combined with linear compensation and multi-objective weighted reward function, sensor drift and system vulnerability problems in intelligent lighting systems are solved, precise perception and adaptive balance of multi-dimensional dimming is achieved, and the system's robustness and energy consumption efficiency are improved.

CN120282349APending Publication Date: 2025-07-08BEIJING BIHAIYIJING LANDSCAPING CO LTD
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510381347.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When responding to complex lighting needs, existing intelligent lighting systems have problems such as sensor drift, rigid multi-objective dimming strategies and high system vulnerability, resulting in insufficient data stability, waste of energy consumption and safety risks.

Method used

Ambient light sensor, flow sensor and user terminals are used to collect data in real time, and the brightness and color temperature of the light is dynamically calculated through the reinforcement learning model, and combined with linear compensation, failover and multi-objective weighted reward functions, dynamic dimming and adaptive balance are achieved.

Benefits of technology

It realizes accurate perception of the entire scene, reduces sensor drift, improves system robustness and energy consumption balance, and ensures adaptive adjustment of multi-dimensional indicators and reliability of basic lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120282349A_ABST
    Figure CN120282349A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent control, and discloses a light control method and system, a controller and a storage medium, and the method comprises the steps: collecting ambient light, human traffic and user preference data through multi-sensor fusion; after normalization processing, inputting a reinforcement learning model to dynamically calculate target brightness and color temperature, and generating a control instruction to adjust a lamp; and optimizing model parameters in combination with user feedback. The system integrates a data acquisition module, an intelligent processing module, a regulation and control execution module and a feedback optimization module to form a closed-loop control chain. The controller supports PWM dimming and a DMX512 protocol, an emergency module is arranged in the controller, and illumination is maintained according to a historical mean value when communication is interrupted. Through high-precision environment perception calibration, a multi-target dynamic dimming strategy, an intelligent closed-loop learning mechanism and a distributed fault-tolerant control system, the long-standing core problems of light sensation drift, strategy adaptation stiffness, strong cold start dependence, poor cluster stability and the like in the field of intelligent illumination are systematically solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control, and particularly to a lighting control method, system, controller and storage medium. Background Art

[0002] With the rapid development of the Internet of Things and intelligent buildings, intelligent lighting systems are gradually replacing traditional lighting solutions and becoming the infrastructure for scenarios such as airports, commercial complexes, and smart campuses. The lighting requirements in such scenarios exhibit highly dynamic characteristics: ambient natural light varies drastically with time and weather (e.g., the span from strong noon sunlight to moonlight at night reaches 10 4 lux), the difference in the density of people flow between peak and trough periods exceeds 10 times, and there are frequent conflicts in the preferences of multiple users (e.g., the coexistence of office and leisure preferences in the same area).

[0003] However, existing solutions have obvious bottlenecks in dealing with such complex requirements. On the one hand, the mainstream sensor calibration methods rely on regular manual calibration or simple linear compensation, without considering temperature drift and device aging, resulting in insufficient long-term data stability. Measured data from a certain subway station shows that after 6 months of use, the sensor error accumulates up to ±15 lux, causing frequent oscillations of dimming commands.

[0004] On the other hand, most systems adopt fixed thresholds or single-objective optimization rules. For example, they simply pursue energy conservation and forcefully suppress the brightness, or operate at full load for a long time to ensure safety, making it difficult to balance multi-dimensional indicators. In a case of a certain shopping mall, the brightness is still maintained at 80% during the low-traffic period at night, resulting in an annual energy waste of more than 120,000 kWh; while during a promotional event when the number of people surges, the response delay of the safety brightness exceeds 5 minutes, posing a risk of trampling.

[0005] In addition, the system robustness in new scenario deployment and extreme conditions has become a pain point. Traditional solutions need to re-label hundreds of thousands of sets of training data for each scenario, with high costs and long cycles; once the centralized architecture encounters network interruptions or controller failures, it is prone to cause large-scale lighting failures. Summary of the Invention

[0006] Aiming at the deficiencies of the prior art, the present invention provides a lighting control method, system, controller and storage medium, which solves the problems of environmental light perception drift, rigid multi-objective dimming strategy and high system vulnerability in intelligent lighting systems.

[0007] To achieve the above objectives, the present invention is realized through the following technical solutions: A lighting control method includes the following steps: S1. Real-time collect the ambient light intensity, the number of people flow data and user preference parameters through the ambient light sensors, the number of people flow sensors and user terminals deployed in the target area; S2. Normalize and denoise the data collected in step S1, and input it into the reinforcement learning model to dynamically calculate the target light brightness and color temperature; S3. Generate a control instruction according to the output result of step S2, and transmit it to the light controller through a wireless communication protocol to adjust the brightness and color temperature of the lamp in real time; S4. Iteratively optimize the parameters of the reinforcement learning model in step S2 based on user feedback data and historical operation records.

[0008] Preferably, the pedestrian flow data in step S1 is obtained by fusing infrared sensor data and video recognition data, specifically including: When the ambient light intensity E(t)>200 lux, the pedestrian flow density D(t) is calculated as D(t)=0.3D IR (t)+0.7D Vision (t); when the ambient light intensity E(t)≤200 lux, the pedestrian flow density D(t) is calculated as D(t)=0.5D IR (t)+0.5D Vision (t); where D IR (t) is the pedestrian flow density output by the infrared sensor, and D Vision (t) is the pedestrian flow density output by the video recognition model.

[0009] Preferably, the processing of the ambient light intensity data in step S1 includes: Set the effective range [E min ,E max =[5,10 4 lux. When the collected value E(t) exceeds this range, perform linear compensation according to the formula E corrected (t)=0.98E(t)+2 lux; If the calibration fails three times in a row, switch to the backup sensor to collect data and mark the faulty area.

[0010] Preferably, the reward function R t in step S2 is a multi-objective weighted sum, and is calculated according to the following formula: Among them, λ1 = 0.5, λ2 = 0.3, λ3 = 0.2 are preset weight coefficients, is the normalized ambient light intensity, D(t) is the pedestrian flow density after fusion in step S1, R t is the reward value at time point t, L 实际 is the actual brightness value of the lamp, L 目标 is the target brightness value, and log(D(t)+1) is the non-linear adjustment term of the pedestrian flow density to the safe brightness.

[0011] Preferably, the reward function includes a user satisfaction index, an energy-saving index, and a safety brightness index, where: The user satisfaction index λ1(1 - |L 实际 - L 目标 |) is used to measure the deviation between the actual brightness and the user-preferred target value; The energy-saving index Optimizes energy consumption by normalizing the ratio of ambient light to brightness; The safety brightness index λ3log(D(t) + 1) dynamically adjusts the minimum safety brightness according to the crowd density.

[0012] The present invention also provides a lighting control system, which is applied to the lighting control method as described above, and includes: A data acquisition module, which is used to acquire ambient light, the number of people, and user preference data, and transmit the data to the intelligent processing module; an intelligent processing module, which is connected to the data acquisition module, and is used to perform normalization processing on the data and run a reinforcement learning algorithm to generate control instructions; A regulation execution module, which is connected to the intelligent processing module, and is used to send the control instructions to the lighting controller through a wireless communication protocol; a feedback optimization module, which is connected to the regulation execution module and the user terminal, and is used to optimize the algorithm parameters according to user feedback and system operation data, and feed back the updated parameters to the intelligent processing module.

[0013] The present invention also provides a controller, which is applied to the lighting control method as described above, and includes: A communication interface, which is used to receive control instructions from the lighting control system; A processing unit, which is connected to the communication interface unit, and is used to parse the instructions and generate a dimming signal; A driving circuit, which supports PWM dimming and the DMX512 protocol, and is connected to the data processing unit, and is used to drive the lamps according to the dimming signal, and support operation according to an emergency strategy when communication is interrupted; A local cache unit, which stores the historical lighting parameters of the most recent hour, and controls the lighting according to the formula when communication is interrupted, where L 应急 (t) is the target brightness value of the lamp in the emergency mode at time point t, N is the length of the historical data window, and L(t - k) is the historical brightness value of the lamp at time point t - k. The present invention also provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method as described above, and stores the following data: Ambient light intensity data E(t), fused number-of-people data D(t), user preference parameter L 目标 and T 目标 ; The policy network parameters of the reinforcement learning model and the reward function weights λ1, λ2, and λ3.

[0014] The present invention provides a lighting control method, system, controller, and storage medium. It has the following beneficial effects: 1. The present invention adopts ambient light sensor linear compensation and fault switching logic to achieve precise perception of the full scene from moonlight to strong sunlight. Compared with the defects of large zero-point drift and slow fault recovery of sensors in the prior art, the compensation residual is reduced to ±2 lux, and the fault switching delay is <10 seconds, solving the pain point of large data fluctuations in large-scale scenes such as airports and squares.

[0015] 2. The present invention realizes the adaptive balance of multi-dimensional indicators through the weighted fusion of user preferences and the non-linear adjustment of safety brightness. Traditional solutions rely on fixed thresholds or single-objective rules and cannot dynamically adapt to changes in the flow of people and preferences. This solution dynamically adjusts the reward function weights.

[0016] 3. Based on the weighted feedback mechanism of explicit scoring and implicit coverage frequency, combined with transfer learning, the present invention improves the cold start efficiency of new regions. In the prior art, model iteration depends on manual annotation or fixed data sets. This solution solves the problem of control rigidity caused by sparse data in small and medium-sized scenes through online learning and simulated data injection.

[0017] 4. Through sensor interpolation completion, local historical caching, and redundant takeover of multiple controllers, the system can still maintain 92% of the area with normal lighting when 30% of the nodes fail. Traditional centralized control schemes are prone to collapse due to single-point failures. This design combines edge computing and distributed decision-making to provide basic safety lighting even under extreme conditions such as power outages and network disconnections, and the robustness reaches industrial-grade standards. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic diagram of the method flow of the present invention; Figure 2 It is a system framework diagram of the present invention; Figure 3 It is a schematic diagram of the controller structure of the present invention; Figure 4 It is a schematic diagram of the storage medium of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0020] Please refer to the attachedFigure 1 , embodiments of the present invention provide a lighting control method, including the following steps: S1. Real-time collect ambient light intensity, pedestrian flow data, and user preference parameters through ambient light sensors, pedestrian flow sensors, and user terminals deployed in the target area; Generally, ambient light sensors are deployed at the edges, centers, and densely populated areas of the target area, such as key nodes like building facades and square entrances. The sensors use high-precision photosensitive elements (such as XYZ-500 type), with a collection frequency of once per second, a coverage range of 0 - 100,000 lux, and the effective detection range is set to min , E max = [5, 10 4 lux to adapt to a wide range of scenarios from moonlight to strong sunlight.

[0021] In this embodiment, the sensor is built-in with a temperature compensation circuit. When the ambient temperature fluctuates by more than ±10 °C, zero calibration is automatically triggered. The collected original light intensity E(t) (unit: lux) is transmitted to the central processing unit through a wireless communication module (ZigBee protocol).

[0022] Specifically, when the sensor collection value E(t) exceeds the effective range, linear compensation calibration is triggered. The calibration formula is: E corrected (t) = α·E(t) + β (α = 0.98, β = 2 lux); Among them, E corrected (t) is the compensated ambient light intensity value, the slope coefficient α = 0.98 and the intercept β = 2 lux, which are determined through the calibration experiment before the sensor leaves the factory: Calibration method: In a standard darkroom, use an adjustable light source (model LM-300) to provide stable illuminations of 5 lux, 1000 lux, 5000 lux, 10 4 lux in sequence, and record the sensor output value E raw ; Linear regression analysis: Based on the least squares method, fit E corrected = αE raw + β, where E raw is the measured light intensity value of the original output of the sensor. Finally, α = 0.98 and β = 2 lux are obtained, and the fitting residual R 2 > 0.995.

[0023] As an option, if E corrected (t) still exceeds the effective range after three consecutive calibrations, it is determined that the sensor is faulty. Switch to a backup sensor within a radius of 5 meters to collect data, and send a fault alarm signal (including the fault location coordinates, timestamp, and abnormal data record) to the central system.

[0024] Further, the effective range is 5, 10 4 lux is set based on the following: E min = 5 lux: Refer to the minimum requirement for pedestrian road lighting in the standards of the International Commission on Illumination (CIE); E max = 10 4 lux: The typical light intensity under direct sunlight is 10 5 lux, and a 1 / 10 margin is reserved for the dynamic range of the sensor to avoid saturation.

[0025] In a possible implementation, an infrared sensor (model IR-200) and a video recognition device (equipped with the YOLOv5 model) are respectively deployed at key positions such as pedestrian walkways and entrances and exits. The infrared sensor detects human movement based on the pyroelectric effect and outputs the pedestrian flow density D IR (t) (unit: person / m 2 ); The video recognition module analyzes the surveillance video at a frame rate of 30 fps and outputs the density D Vision (t) (unit: person / m 2 ). In this embodiment, the fusion algorithm dynamically adjusts the weights according to the environmental light intensity, and the specific process is as follows: When the environmental light intensity E(t) > 200 lux (daytime or strong lighting scenario), the video recognition accuracy is relatively high, and the pedestrian flow density is calculated according to the following formula: D(t) = 0.3D IR (t) + 0.7D Vision (t); When the environmental light intensity E(t) ≤ 200 lux (nighttime or low light scenario), the infrared sensor has better stability, and the fusion formula is adjusted to: D(t) = 0.5D IR (t) + 0.5D Vision (t); Specifically, the weight coefficients 0.3 and 0.7 are determined by comparing the measured data: When the light is sufficient, the video recognition accuracy can reach 93% (the false detection rate is less than 5%), while the infrared sensor is easily interfered by heat sources (accuracy 82%); Under low light, the video recognition accuracy drops to 75%, and the infrared sensor maintains above 80%. Therefore, the weights are balanced to improve the robustness.

[0026] Further, the setting of the light threshold of 200 lux refers to the indoor and outdoor lighting transition value in the CIE standard (such as the square lighting design specification GB50034-2013).

[0027] As an option, the video recognition model is trained with data augmentation for low light scenarios: The training set contains 100,000 pairs of night infrared images and visible light images; The loss function uses weighted cross-entropy where, y i is the one-hot encoding of the true label, and p i is the probability value of the i-th category predicted by the model. The weight of night samples is increased to 1.5 times; The mAP (mean average precision) of the model reaches 78.5% on the night test set, which is 12% higher than the non-enhanced version.

[0028] Users set the brightness L i (range: 0 - 5000 lux), color temperature T i (range: 2000 - 6500 K) and dynamic mode parameters (such as gradient frequency, flicker mode) through a mobile terminal (such as a mobile phone APP). After the data is encrypted by AES-256, it is uploaded to the cloud server through the HTTPS protocol, and the key management is protected by a hardware security module (HSM).

[0029] In this embodiment, the group preference is generated by a weighted average algorithm: where, L 目标 is the target brightness value after fusing the preferences of multiple users, T 目标 is the target color temperature value after fusing the preferences of multiple users, n is the total number of users participating in preference fusion in the current area, and the weight W i is calculated according to the monthly active times of users: N active,i is the active times of user i in the past 30 days (logging in and operating the APP every day is counted as 1 time); is the total sum of the active times of all m users in the current area; Example: There are 3 users in a certain area, and their active times are 15, 5, and 30 times respectively. Then the weights are W1 = 0.3, W2 = 0.1, W3 = 0.6.

[0030] In a possible implementation, the user preference data supports real-time update and historical record rollback. When multiple people are in the same area, the system preferentially matches the average preference value in the recent 30 days as the initial target value. If a new user enters, it will be dynamically adjusted according to the real-time weight. Specifically, the encrypted transmission process includes: The client generates a random symmetric key K session , and transmits it to the server after being encrypted by RSA-2048; The preference data P i ={L i ,T i} Use K session Encrypt it into ciphertext C i = AES-256(P i , K session ) The server decrypts it and stores it in the distributed database (Cassandra), and the access permission is controlled by the OAuth2.0 protocol.

[0031] After aligning the environmental light data and the pedestrian flow data through timestamps, they are input into the distributed database (such as Cassandra) and associated with the user preference parameters to generate a multi-dimensional input vector. For example, in the commercial complex scenario, when the light in a certain area suddenly drops to 50 lux and the pedestrian flow increases to 0.8 people / m², the system executes the following linkage logic: Data verification: Check whether the environmental sensor is faulty. If it is normal, enter the dynamic dimming process; Preference matching: Call the preference data of active users (monthly activity ≥ 15 times). If there are no active users in this area, use the historical average value L 目标 = 400 lux, T 目标 = 3500; Exception handling: If the sensor is faulty, interpolate and complete based on the data of adjacent nodes. The interpolation formula is: Among them, E interp (t) is the interpolated and completed environmental light intensity value calculated at time point t. k is the number of adjacent sensors participating in the interpolation calculation. E j (t) is the environmental light intensity value collected by the j-th adjacent sensor at time point t, and j is the serial number index of the adjacent sensor.

[0032] Furthermore, the fault marking mechanism is linked with the subsequent regulation execution module: when a certain sensor is marked as faulty, the system automatically shields its data stream and triggers the edge node to reallocate the calculation task to ensure data integrity.

[0033] S2. Normalize and denoise the data collected in step S1, and input it into the reinforcement learning model to dynamically calculate the target light brightness and color temperature; Generally, the raw data collected in step S1 (environmental light E(t), pedestrian flow density D(t), user preferences L 目标 and T 目标 ) needs to be standardized to adapt to the input requirements of the algorithm. In this embodiment, the data preprocessing includes the following operations: Environmental light normalization: Among them, E norm (t) is the normalized environmental light intensity value, Emin = 5 lux is the minimum effective light intensity value after sensor calibration, E max = 10 4 lux is the maximum effective light intensity value after sensor calibration; E norm (t) ∈ [0, 1] to eliminate the dimension difference; E min is set strictly consistent with the sensor calibration range as E max .

[0034] Normalization of pedestrian flow density: Among them, D norm (t) is the normalized pedestrian flow density value, D capacity = 0.8 person / m 2 is the maximum safe pedestrian flow density threshold per unit area; D capacity is set according to the "Safety Code for Crowd Density in Public Places" (GB / T 28921 - 2012); The output range D norm (t) ∈ [0, 1].

[0035] Denoising processing: Apply a 5 - minute sliding window mean filter to E(t) and D(t) respectively, and the data within the window is weighted by time for calculation: Among them, is the smoothed value or predicted value at time t, and E(t - k) is the original observed value of the time series at time t - k; The attenuation coefficient α k suppresses sudden noise (such as insects blocking the sensor) and retains the trend change.

[0036] In a possible implementation, the Proximal Policy Optimization (PPO) algorithm is used to construct a dynamic dimming model, and its state space, action space, and reward function are designed as follows: State space: Among them, E norm (t) is the normalized ambient light intensity value, D norm (t) is the normalized pedestrian flow density value.

[0037] The input dimension is 4, corresponding to the normalized ambient light, pedestrian flow density, and user preference target value respectively.

[0038] Action space: A t= [ΔL, ΔT] (ΔL ∈ [-50, 50] lux, ΔT ∈ [-500, 500] K); The limit values of the brightness adjustment amount ΔL and the color temperature adjustment amount ΔT are set according to the lamp response speed.

[0039] Reward function: where λ1 = 0.5, λ2 = 0.3, λ3 = 0.2 are preset weight coefficients, is the normalized ambient light intensity, D(t) is the pedestrian flow density after fusion in step S1, R t is the reward value at time point t, L 实际 is the actual brightness value of the lamp, L 目标 is the target brightness value, log(D(t) + 1) is the non-linear adjustment term of the pedestrian flow density to the safety brightness.

[0040] User satisfaction index R user : where: the smaller the absolute deviation between the actual brightness and the target value, the higher the score, and the maximum deviation is normalized to L max (such as 5000 lux); Example: If L 目标 = 400 lux, L 实际 = 380 lux, then R user = 0.5×(1 - 20 / 5000) ≈ 0.498.

[0041] Energy-saving index R energy : Technical meaning: The stronger the ambient light E norm (t) and the lower the actual brightness L 实际 are, the more significant the energy-saving effect; Constraint condition: When E norm < 0.1 (i.e., E(t) < 1000 lux), force L 实际 ≥ 200 lux to avoid insufficient lighting caused by excessive energy saving.

[0042] Safety brightness index R safe : R safe = λ3log(D(t) + 1); Technical meaning: When the pedestrian flow density D(t) increases, the safety brightness requirement increases gently according to the logarithmic function, avoiding a sharp increase in energy consumption caused by linear growth; Parameter verification: When D(t) increases from 0.1 to 0.8 person / m 2When R safe increases from 0.041 to 0.207, the increase in brightness demand is controllable.

[0043] In some embodiments, the model training adopts a distributed architecture, and the specific process includes: Experience replay pool: Store historical state-action-reward tuples (S t , A t , R t , S t+1 ), with a capacity of 100,000, and high-reward samples are preferentially retained; Policy update: Sample batch data (batchsize = 256) every 1000 steps, calculate the advantage function A t = R t - V(S t ), and update the policy network parameters θ: where θ old is the policy network parameter before update, θ new is the policy network parameter after update, η is the learning rate, r t (θ) is the importance sampling ratio, A t is the advantage function, clip(r t (θ), 0.8, 1.2) is the clipping function; Learning rate: η = 0.001, referring to the default setting of the PPO algorithm; Clipping ratio: 0.8 - 1.2, to prevent the policy update from being too large.

[0044] Model deployment: The trained policy network is converted into a TensorRT engine and deployed to an edge computing node (such as NVIDIA Jetson Xavier), and the inference latency is less than 10ms.

[0045] S3. Generate a control instruction according to the output result of step S2, and transmit it to the lighting controller through a wireless communication protocol to adjust the brightness and color temperature of the lamp in real time; Generally, the action space output by step S2 is the brightness adjustment amount ΔL and the color temperature adjustment amount ΔT, and they need to be converted into drive signals that can be parsed by the lamp. In this embodiment, the generation logic of the control instruction includes the following core operations: Absolute brightness L new Calculation: L new = L old + ΔL (L new ∈ [0, L max , L max = 5000 lux); Out-of-bounds processing: If L new< 0, force it to be set to 0; if L new > L max , force it to be set to L max .

[0046] Example: If the current brightness L old = 400 lux and ΔL = +50, then L new = 450 lux.

[0047] Absolute color temperature T new Calculation: T new = T old + ΔT (T new ∈ [T min , T max , T min = 2000K, T max = 6500K) Constraint rule: When the color temperature exceeds the range, truncate it according to the nearest principle (e.g., T new = 6800K → 6500K} = 6500K).

[0048] Control signal encoding: PWM dimming signal: Duty cycle Frequency f = 1 kHz, accuracy 0.1%; DMX512 protocol signal: Channel value Value range [0, 255].

[0049] As an option, the control instruction is sent to the lamp controller through the wireless communication module to ensure low latency and high reliability.

[0050] In this embodiment, the transmission protocol and mechanism are designed as follows: Communication protocol: Adopt ZigBee3.0 protocol, operating frequency band 2.4 GHz, transmission power 10 dBm, support AES-128 encryption; Data frame format: Brightness instruction: 16-bit integer, representing L new (e.g., 450 lux is encoded as 0x01C2); Color temperature instruction: 16-bit integer, representing T new (e.g., 3500K is encoded as 0x0DAC); Checksum: CRC-16 check, polynomial 0x8005, covering the entire instruction field.

[0051] Specifically, the single communication delay is controlled within 50 ms, and the packet loss rate is less than 0.1%, meeting the real-time regulation requirements. For example, in the commercial complex scenario, the central controller broadcasts instructions to 10 lamps in area A to ensure that all nodes respond synchronously within 100 ms.

[0052] In a possible implementation, after receiving an instruction, the lighting controller adjusts the state of the lighting fixture through a driving circuit and feeds back the actual parameters to the central system. In this embodiment, the execution process includes: Instruction parsing: Verify the integrity of the frame. If a CRC error occurs, request retransmission; Parse L new and T new , and convert them into PWM duty cycle and DMX channel values.

[0053] Driver output: PWM dimming: Generate a square wave signal with a duty cycle of δ through a timer to drive the LED current to change linearly; DMX512 control: Write C DMX to the preset channel (such as channel 1), and the color temperature driving chip (such as TLC5973) adjusts the RGBW mixing ratio according to the value.

[0054] Status feedback: Collect the actual brightness L 实际 (feedback through a photoresistor) and the color temperature T 实际 (feedback through a color temperature sensor); The feedback data is encapsulated according to the protocol and uploaded to update the state space S of the reinforcement learning model t .

[0055] When a communication interruption or controller failure occurs, the system activates an emergency strategy to maintain the basic lighting function. In this embodiment, the emergency logic includes: Local cache execution: The controller calls the historical brightness data of the last hour and calculates according to the formula: Calculate the average brightness and keep the color temperature at a safe value T default = 4000K, where L 应急 (t) is the target brightness value of the lighting in the emergency mode at time point t, N is the length of the historical data window, and L(t - k) is the historical brightness value of the lighting at time point t - k.

[0056] Fault marking and isolation: If communication fails three times in a row, the controller sends a fault alarm to adjacent nodes, the system automatically masks the data of this node, and a redundant controller takes over the task.

[0057] S4. Based on the user feedback data and historical operation records, iteratively optimize the parameters of the reinforcement learning model in step S2.

[0058] Generally, user feedback data includes two categories: explicit ratings and implicit behaviors. Historical operation records cover parameters such as the actual brightness, color temperature, energy consumption of the lamp, and environmental status. In this embodiment, the data collection and processing mechanism is as follows: Explicit feedback: The user rates the current lighting effect Q ∈ [1, 5] through the mobile APP, and the rating rule is: Implicit feedback: The frequency F (times / day) of the user manually overriding the automatic settings is counted, which reflects the dissatisfaction with the system's recommendation strategy.

[0059] Historical records: Store the S t ,A t ,R t ,S t+1 quadruple to the distributed database (Cassandra), and the retention period is 90 days.

[0060] As an option, the user feedback data is encrypted and transmitted through HTTPS, and after being associated with the historical operation records, weighted training samples are generated. For example, a certain sample is expressed as: Sample = (S t ,A t ,R t +αQ+βF,S t+1 )(α = 0.1, β = -0.05); Among them, α and β are feedback weight coefficients, which are determined through A / B testing.

[0061] In a possible implementation, an optimization strategy that combines offline batch processing and online incremental learning is adopted, and the update period is the low-load period in the early morning every day.

[0062] In this embodiment, the optimization process includes the following core steps: Experience replay pool update: Randomly sample N = 10 5 samples from the historical records, and divide the training set (80%) and the validation set (20%) according to the time stamp; Eliminate abnormal samples (such as R t > 3×avg or ΔL> 100 lux).

[0063] Policy network update: Based on the PPO algorithm, the policy loss function is defined as: Among them, r t (θ) is the importance sampling ratio, and clip(r t (θ), 0.8, 1.2) is the clipping function.

[0064] The value function loss is the mean squared error: where the discount factor γ = 0.99, and V θ (s t ) is the state value function, is the value function estimate of state s under the old policy, and fixed parameters are used to prevent overfitting. t+1

[0065] Dynamic weight adjustment: The user satisfaction weight λ1 is dynamically adjusted according to the feedback score: where, and are the user satisfaction weights before and after update, ∑Q i is the sum of all user ratings during the statistical period, N Q is the number of valid rating samples, and η is the learning rate; If then it is truncated to 0.7 to balance multi-objective optimization.

[0066] When deploying in a new area or when user feedback is sparse, the system adopts transfer learning and simulated data injection strategies.

[0067] Specifically: Transfer learning: Load the pre-trained model parameters of a similar area (such as a commercial complex), freeze the underlying network (feature extraction layer), and only fine-tune the top-level policy network; Simulated data generation: Synthesize virtual samples based on historical mean and standard deviation: where, μ S , σ S are the statistics of each dimension of the state space, obtained by statistics from historical data s t , is the normal distribution, generating virtual state S that conforms to the historical distribution characteristics sim , is the old policy network, generating reasonable action A based on the virtual state sim .

[0068] After the model is updated, the performance improvement is verified through offline evaluation and online A / B testing. In this embodiment, the evaluation metrics include: User satisfaction improvement rate: Energy saving rate: Policy stability: Statistically calculate the variances of the action spaces ΔL and ΔT, ensuring that Var(ΔL) < 50 and Var(ΔT) < 500. ​

[0069] Among them, Q new , Q old is the comparison value of user ratings in the same area before and after optimization, and E new , E old is the lamp energy consumption in the same time period before and after optimization.

[0070] Example: After a certain update, ρ user = 8.5%, ρ energy = 3.2%, the action variance decreased by 12%, and it was determined that the optimization was effective and fully deployed.

[0071] A lighting control system described below can be correspondingly referred to with a lighting control method described above.

[0072] Please refer to the appendix Figure 2 , a lighting control system, applied to the lighting control method as described above, includes: A data acquisition module, used to collect ambient light, pedestrian flow, and user preference data, and transmit the data to the intelligent processing module; an intelligent processing module, connected to the data acquisition module, used to perform normalization processing on the data and run a reinforcement learning algorithm to generate control instructions; A regulation execution module, connected to the intelligent processing module, used to send the control instructions to the lighting controller through a wireless communication protocol; a feedback optimization module, connected to the regulation execution module and the user terminal, used to optimize the algorithm parameters according to user feedback and system operation data, and feed the updated parameters back to the intelligent processing module The system of this embodiment can be used to execute the method embodiment above, and its principle and technical effects are similar, so they will not be elaborated here.

[0073] Please refer to the appendix Figure 3 , the present invention also provides a controller, applied to the lighting control method as described above, including: a communication interface, used to receive control instructions from the lighting control system; A processing unit, connected to the communication interface unit, used to parse the instructions and generate a dimming signal; A driving circuit, supporting PWM dimming and DMX512 protocol, which is connected to the data processing unit, used to drive the lamps according to the dimming signal, and support running according to an emergency strategy when the communication is interrupted; A local cache unit, storing historical lighting parameters for the most recent hour, and controlling the lights according to the formula when the communication is interrupted, where L 应急 (t) is the target brightness value of the lights in the emergency mode at time point t, N is the length of the historical data window, and L(t - k) is the historical brightness value of the lights at time point t - k.

[0074] Please refer to the appendix Figure 4, the present invention also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the above method and stores the following data: Environmental light intensity data E(t), fused pedestrian flow data D(t), user preference parameter L 目标 and T 目标 ; The policy network parameters of the reinforcement learning model and the reward function weights λ1, λ2 and λ3.

[0075] Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (Static Random Access Memory, abbreviated as SRAM), electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM), erasable programmable read-only memory (Erasable Programmable Read-Only Memory, abbreviated as EPROM), programmable read-only memory (Programmable Red-Only Memory, abbreviated as PROM), read-only memory (Read-Only Memory, abbreviated as ROM), magnetic memory, flash memory, magnetic disk or optical disc.

[0076] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A lighting control method, characterized in that, Including the following steps: S1. Real-time collect ambient light intensity, pedestrian flow data, and user preference parameters through an ambient light sensor, a pedestrian flow sensor, and a user terminal deployed in the target area; S2. Normalize and denoise the data collected in step S1, and input it into the reinforcement learning model to dynamically calculate the target light brightness and color temperature; S3. Generate a control instruction according to the output result of step S2, and transmit it to the light controller through a wireless communication protocol to adjust the brightness and color temperature of the lamp in real time; S4. Based on the user feedback data and historical operation records, iteratively optimize the parameters of the reinforcement learning model in step S2.

2. The lighting control method according to claim 1, characterized in that In step S1, the pedestrian flow data is obtained by fusing infrared sensor data and video recognition data, specifically including: When the environmental light intensity E(t) > 200 lux, the pedestrian flow density D(t) is calculated according to D(t) = 0.3D IR (t) + 0.7D Vision (t); When the environmental light intensity E(t) ≤ 200 lux, the pedestrian flow density D(t) is calculated as D(t) = 0.5D IR (t) + 0.5D Vision (t); Among them, D IR (t) is the pedestrian flow density output by the infrared sensor, and D Vision (t) is the pedestrian flow density output by the video recognition model.

3. A lighting control method according to claim 1, characterized in that, The processing of the ambient light intensity data in step S1 includes: Set the effective range min , E max = [5, 10 4 lux. When the collected value E(t) exceeds this range, perform linear compensation according to the formula E corrected (t) = 0.98E(t) + 2 lux; If the calibration fails three times in a row, switch to the backup sensor to collect data and mark the fault area.

4. A lighting control method according to claim 1, characterized in that, The reward function R of the reinforcement learning model in step S2 t is a weighted sum of multiple objectives and is calculated according to the following formula: Among them, λ1 = 0.5, λ2 = 0.3, and λ3 = 0.2 are preset weight coefficients. is the normalized ambient light intensity, D(t) is the density of the flow of people after fusion in step S1, and R t is the reward value at time point t, and L 实际 is the actual brightness value of the lamp, and L 目标 is the target brightness value, and log(D(t)+1) is the non-linear adjustment term of the flow density of people to the safety brightness.

5. A lighting control method according to claim 4, characterized in that, The reward function includes user satisfaction indicators, energy-saving indicators, and safe brightness indicators, where: The user satisfaction index λ1(1 - |L 实际 - L 目标 |) is used to measure the deviation between the actual brightness and the user-preferred target value; Energy saving means Optimizing energy consumption by normalizing the ratio of ambient light to brightness; The safe brightness indicator λ3log(D(t)+1) dynamically adjusts the minimum safe brightness according to the pedestrian flow density.

6. A lighting control system, characterized in that, Applied to the lighting control method according to any one of claims 1-5, including: A data acquisition module for collecting ambient light, pedestrian flow, and user preference data, and transmitting the data to the intelligent processing module; an intelligent processing module, connected to the data acquisition module, for normalizing the data and running a reinforcement learning algorithm to generate a control instruction; A regulation execution module, connected to the intelligent processing module, for sending the control instruction to the light controller through a wireless communication protocol; a feedback optimization module, connected to the regulation execution module and the user terminal, for optimizing the algorithm parameters according to user feedback and system operation data, and feeding back the updated parameters to the intelligent processing module.

7. A controller, characterized in that, Applied to the lighting control method according to any one of claims 1-5, including: A communication interface for receiving a control instruction from the lighting control system; A processing unit, connected to the communication interface unit, for parsing the instruction and generating a dimming signal; A drive circuit, supporting PWM dimming and DMX512 protocol, connected to the data processing unit, for driving the lamp according to the dimming signal, and supporting operation according to the emergency strategy when the communication is interrupted; Local cache unit, which stores historical lighting parameters for the most recent hour and controls the lights according to the formula when communication is interrupted. Among them, L 应急 (t) is the target brightness value of the light in the emergency mode at time point t, N is the length of the historical data window, and L(t - k) is the historical brightness value of the light at time point t - k.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-5, and stores the following data: Environmental light intensity data E(t), fused pedestrian flow data D(t), user preference parameter L 目标 and T 目标 ; The policy network parameters of the reinforcement learning model and the reward function weights λ1, λ2, and λ3.

Citation Information

Cited By

  • Intelligent lighting decision optimization method based on big data

    CN120525021A

  • Intelligent lighting decision optimization method based on big data

    CN120525021B

  • LED lamp energy-saving control method and system

    CN120692728A

  • Illumination regulation and control method, adjustable illumination equipment and goods shelf

    CN120769406A

  • Scene-based LED lamp strip adaptive control method, device and equipment

    CN120935903A