A Dynamic Energy Consumption Optimization Method and System for Long-Cascaded LED Light Strips Based on Deep Reinforcement Learning

CN122579375APending Publication Date: 2026-08-14MYNICE OPTOELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现有技术多采用固定功率调节策略,未充分融合多维度运行参数与场景化能耗约束,难以动态适配不同应用场景的需求,导致能耗优化效果不佳,无法实现能效与照明需求的动态平衡

Benefits of technology

[0012]相比现有技术,本发明提供的有益效果包括:采用本发明公开的一种基于深度强化学习的长级联LED灯带动态能耗优化方法及系统,包括:首先采集灯带各单元运行感知信息,结合当前应用场景及能耗约束生成调控指令描述;通过深度强化学习构建的能耗决策模型提取灯带状态并生成功率控制决策,下发至分段控制单元执行调节;采集调节后状态反馈数据,对比约束目标生成强化学习奖励信号,将相关数据作为训练样本迭代更新模型。本方法解决了现有固定策略适配性差的问题,实现能耗动态精准优化,有效提升长级联LED灯带能效水平。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122579375A_ABST
    Figure CN122579375A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for dynamic energy consumption optimization of long cascaded LED light strips based on deep reinforcement learning. The method includes: firstly, collecting operational perception information from each unit of the light strip and generating control command descriptions based on the current application scenario and energy consumption constraints; secondly, extracting the light strip state and generating power control decisions through an energy consumption decision model constructed using deep reinforcement learning, and issuing these decisions to the segmented control units for adjustment; and thirdly, collecting state feedback data after adjustment, comparing it with the constraint target to generate reinforcement learning reward signals, and using the relevant data as training samples to iteratively update the model. This method solves the problem of poor adaptability of existing fixed strategies, achieves precise dynamic energy consumption optimization, and effectively improves the energy efficiency of long cascaded LED light strips.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a method and system for dynamic energy consumption optimization of long cascaded LED light strips based on deep reinforcement learning. Background Technology

[0002] Long-chain LED light strips are widely used in landscape lighting, interior decoration, and other fields due to their wide lighting range and segmented controllability. Existing technologies mostly adopt fixed power adjustment strategies, which do not fully integrate multi-dimensional operating parameters and scenario-based energy consumption constraints. This makes it difficult to dynamically adapt to the needs of different application scenarios, resulting in poor energy consumption optimization and an inability to achieve a dynamic balance between energy efficiency and lighting requirements. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for dynamic energy consumption optimization of long cascaded LED light strips based on deep reinforcement learning.

[0004] In a first aspect, embodiments of the present invention provide a method for dynamic energy consumption optimization of long cascaded LED light strips based on deep reinforcement learning, comprising:

[0005] The operation sensing information of the target long cascaded LED light strip is collected. The operation sensing information includes the real-time current, operating voltage, junction temperature, output brightness value of each series unit of the light strip, as well as the illuminance and ambient temperature of the environment in which the light strip is located.

[0006] Based on the current application scenario and preset energy consumption constraint target of the target long cascaded LED light strip, a corresponding energy consumption control instruction description is generated;

[0007] The operational sensing information and the energy consumption control instruction description are loaded into the energy consumption decision model to extract the light strip status and generate energy consumption control decisions, thereby obtaining the energy consumption adjustment analysis results and power control decision results of the operational sensing information; the energy consumption control instruction description is used to provide the decision guidance information required by the energy consumption decision model to perform light strip status extraction and energy consumption control decision generation.

[0008] The power control decision results are sent to the segmented control unit of the target long cascaded LED light strip to execute the corresponding power adjustment operation;

[0009] Collect feedback data on the operating status of the light strip after the power adjustment operation is executed, and compare the feedback data with the constraint target in the description of the energy consumption control instruction to generate a reinforcement learning reward signal;

[0010] The operational perception information, energy consumption regulation command description, power control decision results, feedback data, and reinforcement learning reward signals are stored as training samples for reinforcement learning iterative updates of the energy consumption decision model; wherein, the energy consumption decision model is constructed based on a deep reinforcement learning algorithm.

[0011] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being used to perform the method described in the first aspect.

[0012] Compared to existing technologies, the beneficial effects provided by this invention include: The method and system for dynamic energy consumption optimization of long cascaded LED light strips based on deep reinforcement learning, as disclosed in this invention, include: firstly, collecting operational perception information from each unit of the light strip, and generating control command descriptions based on the current application scenario and energy consumption constraints; extracting the light strip state and generating power control decisions through an energy consumption decision model constructed using deep reinforcement learning, and issuing these decisions to the segmented control units for adjustment; collecting state feedback data after adjustment, comparing it with the constraint target to generate reinforcement learning reward signals, and using the relevant data as training samples to iteratively update the model. This method solves the problem of poor adaptability of existing fixed strategies, achieves precise dynamic energy consumption optimization, and effectively improves the energy efficiency level of long cascaded LED light strips. Attached Figure Description

[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 A flowchart illustrating the steps of the dynamic energy consumption optimization method for long cascaded LED light strips based on deep reinforcement learning provided in an embodiment of the present invention.

[0015] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0016] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0017] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0018] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the dynamic energy consumption optimization method for long cascaded LED light strips based on deep reinforcement learning provided in this embodiment. The following is a detailed description of this dynamic energy consumption optimization method for long cascaded LED light strips based on deep reinforcement learning.

[0019] Step S201: Collect the operation sensing information of the target long cascaded LED light strip. The operation sensing information includes the real-time current, operating voltage, junction temperature, output brightness value of each series unit of the light strip, as well as the light intensity and ambient temperature of the environment where the light strip is located.

[0020] Step S202: Generate a corresponding energy consumption control instruction description based on the current application scenario of the target long cascaded LED light strip and the preset energy consumption constraint target;

[0021] Step S203: Load the operation perception information and the energy consumption control instruction description into the energy consumption decision model to extract the light strip status and generate energy consumption control decisions, thereby obtaining the energy consumption adjustment analysis result and power control decision result of the operation perception information; the energy consumption control instruction description is used to provide the decision guidance information required for the energy consumption decision model to perform light strip status extraction and energy consumption control decision generation.

[0022] Step S204: Send the power control decision result to the segmented control unit of the target long cascaded LED light strip and execute the corresponding power adjustment operation;

[0023] Step S205: Collect the operating status feedback data of the light strip after the power adjustment operation is executed, and compare the feedback data with the constraint target in the description of the energy consumption control instruction to generate a reinforcement learning reward signal;

[0024] Step S206: Store the operation perception information, energy consumption control instruction description, power control decision result, feedback data and reinforcement learning reward signal as training samples for reinforcement learning iterative updates of the energy consumption decision model; wherein, the energy consumption decision model is constructed based on a deep reinforcement learning algorithm.

[0025] In this embodiment of the invention, for example, the server periodically or event-triggeredly acquires comprehensive operational awareness information from the hardware layer of the target long cascaded LED light strip through its data acquisition interface. Specifically, the server communicates with the segmented control unit of each LED series unit (or module) to read its real-time electrical and thermal status parameters. For example, for a long cascaded light strip consisting of 120 LED modules connected in series for lighting large indoor stadiums, the server reads the real-time operating current (e.g., 18.5mA) and operating voltage (e.g., 3.2V) from the driver chip register of each module. Simultaneously, the server obtains an estimated junction temperature (e.g., 55°C) from the temperature sensor integrated on each module (e.g., a thermistor or digital temperature sensor). Furthermore, each module is equipped with a photosensor to detect its actual output luminous efficacy (in lux or cd / m²), which reflects the true luminous efficiency of the LED chip under the current driving and temperature conditions.

[0026] In addition to the status of the LED strip itself, the server also collects external environmental parameters from environmental monitoring nodes deployed in the LED strip's illumination area. For example, within the same venue, multiple light sensors measure the ambient background illuminance (e.g., 200 lux), and multiple temperature sensors measure the ambient air temperature (e.g., 25°C). The server aggregates, timestamps, and formats all this scattered data—current, voltage, junction temperature, brightness of 120 modules, and multiple ambient light and temperature readings—to form a complete, structured operational awareness information matrix. This matrix characterizes the overall operational status of the entire long-cascaded LED strip system and its surrounding environment at a specific moment, serving as the data foundation for subsequent intelligent decision-making.

[0027] The server maintains a scenario policy library internally, which defines different application scenarios (such as "full power display mode," "energy-saving maintenance mode," "dynamic scenario mode," and "emergency lighting mode") and their corresponding preset energy consumption constraints. The server determines the current application scenario based on the system clock, external scheduling commands, or the results of automatic analysis of operational perception information. For example, in intelligent building management, when it reaches 6 pm on a weekday, the system automatically switches to "after-get off work energy-saving mode," whose preset energy consumption constraint is "to minimize total power consumption while ensuring a minimum safe illuminance of 100 lux in the corridor."

[0028] Based on this, the server generates a machine-readable, structured energy consumption control instruction description. This description is not a simple target value, but contains multi-level task definitions. For example, for the above scenario, the generated instruction description could be encoded as: {"Scenario":"post_work_saving","Core Constraints":{"Type":"Brightness Lower Limit","Value":100","Unit":"lux"}","Optimization Objective":"Minimize Total Power","Additional Requests":["Evaluate the Current Thermal Equilibrium Status of Each Unit","Output Brightness Adjustment Analysis Report"]}. This instruction description clearly defines the task boundary of the decision (must satisfy 100 lux), the optimization direction (minimum power), and the expected output of the decision process (thermal evaluation and analysis report), providing clear decision guidance information and value judgment basis for subsequent deep reinforcement learning models.

[0029] The server loads the operational perception information matrix collected in step S110 and the energy consumption control instruction description vector generated in step S120 into the energy consumption decision model already deployed in the server's memory. This model is a neural network built based on deep reinforcement learning algorithms (such as DDPG, PPO, or SAC), and it contains sub-modules such as a state representation network, a feature transformation network, a decision generation network, and a control decision network.

[0030] First, the state representation network receives the operational sensing information matrix. This network, after training, is able to extract essential light strip state features from high-dimensional, redundant raw data. For example, it can output a feature vector containing abstract features such as "system average brightness," "brightness uniformity index," "difference between maximum and average junction temperature," "total input power estimate," and "ambient light interference." These features are more compact than the raw data and better reflect the overall health and energy efficiency of the system.

[0031] Next, the feature transformation network further processes the aforementioned state features. Due to the electrical and thermal coupling between the units of the long cascaded LED light strip, its features exhibit sequence dependencies. This network (e.g., incorporating LSTM or Transformer structures) models the feature sequence, captures the mutual influence of preceding and subsequent unit states, and transforms the features into a representation domain more suitable for policy decision-making. For example, it can output a feature sequence that enhances spatial correlation information, identifying which regions have high module junction temperatures forming "hot spots," and which regions have brightness that can be preferentially reduced after meeting constraints—the "redundant areas."

[0032] Then, the decision generation network receives the transformation result features from the feature transformation network and the energy consumption regulation instruction description. The instruction description serves as conditional information, guiding the network to perform decision deduction. During the forward propagation process within the network, not only are the final decision features output, but the activation values ​​of its hidden layers are also used to generate energy consumption regulation analysis results. For example, when processing the "minimize total power" instruction, its internal layers can calculate that if the current of all modules is reduced by 5%, the total system brightness is expected to decrease from the current 120 lux to 105 lux (still above the 100 lux constraint), and the total power can decrease by 8%; if only the current of the identified "redundant area" modules is reduced by 10%, while other areas remain unchanged, the total brightness is expected to decrease to 108 lux, and the total power can decrease by 4.5%. These deduction processes are output in the form of "energy consumption regulation analysis results," such as {"Strategy 1":{"Expected Brightness":105, "Expected Energy Saving Rate":8%}, "Strategy 2":{"Expected Brightness":108, "Expected Energy Saving Rate":4.5%}}, providing a basis for decision-making for maintenance personnel.

[0033] Simultaneously, the decision generation network outputs a final energy consumption decision feature vector, which contains the specific control intent. The control decision network receives this feature vector and decodes it into a specific power control decision result that can be directly sent to the hardware for execution. This result is typically an array of control instructions for each LED series unit. For example, [{"unit_id":1,"pwm_duty":0.85},{"unit_id":2,"pwm_duty":0.87},...,{"unit_id":120,"pwm_duty":0.82}] indicates a suggestion to adjust the PWM duty cycle of each unit to a new target value.

[0034] The server, through its control command sending interface, sends the power control decision result array generated in step S130 to each segment control unit of the target long cascaded LED light strip according to a predetermined communication protocol (such as DMX512, DALI, or a proprietary protocol based on TCP / IP). Each segment control unit (typically a microcontroller or intelligent driver chip) immediately executes a power adjustment operation locally upon receiving its own control command. For example, the control unit with unit ID 1 adjusts the PWM duty cycle of its driven LED module from 0.90 to 0.85, thereby reducing the module's average operating current and power consumption. All units perform the adjustment almost synchronously, changing the operating state of the entire light strip.

[0035] After the power adjustment operation is executed and a short stabilization delay (e.g., 500 milliseconds) has elapsed, the server initiates another round of operational sensing information collection (same as step S110) to obtain feedback data on the operating status of the light strip under the new control strategy. The server then compares and analyzes this feedback data with the constraint targets described in the original energy consumption control instruction.

[0036] For example, the server calculates the actual average brightness in the feedback data to be 106 lux, satisfying the constraint of "not less than 100 lux"; the calculated actual total power consumption is 920W, a decrease of 80W from the previous 1000W. Based on a preset reward function, the server generates a reinforcement learning reward signal. The reward function can be designed to comprehensively consider target compliance and energy-saving effects, for example: Reward = (Brightness constraint satisfied? Basic reward: Large penalty) + Power saving ratio coefficient * Power saving rate. In this example, satisfying the brightness constraint yields a basic reward of +1, and saving 8% of power yields an additional reward of +0.8, for a total reward signal of +1.8. If the feedback brightness is only 95 lux, a large penalty of -5 can be triggered, resulting in a negative total reward even with reduced power consumption. This reward signal is a direct, quantitative evaluation of the quality of the decision made in step S130.

[0037] The server stores all key data generated in this optimization loop as a complete experience tuple. This tuple includes: state (s), initial operational awareness information; instruction (c), a description of energy consumption control instructions; action (a), the issued power control decision result; new state (s'), feedback data after power adjustment; and reward (r), the calculated reinforcement learning reward signal. Specifically, the reinforcement learning reward signal consists of a basic reward for satisfying brightness constraints, an energy-saving reward, and a penalty term, and its specific mathematical expression is as follows:

[0038]

[0039] in, As an indicator function, when the system average brightness Not lower than the target brightness The value is 1 if it is active, and 0 otherwise. The base reward value is set to +2.0. and These represent the total system power at the previous moment and the current moment, respectively. The energy-saving reward coefficient is set to +5.0 to amplify the reward effect of the energy-saving ratio. This is the highest junction temperature among all current LED modules. The junction temperature safety threshold is set to 85°C. The overheat penalty coefficient is set to 0.1, applying a linear penalty to junction temperatures approaching or exceeding a threshold. This reward function is designed to guide the agent to prioritize reducing system power consumption while meeting minimum brightness requirements, and simultaneously preventing any module from overheating.

[0040] The server stores this experience tuple in a database or in-memory queue called an "experience replay buffer." Once the buffer has accumulated enough samples, the server's model training module randomly samples a batch of samples as training data to perform a round of reinforcement learning iterations on the energy consumption decision model (mainly its core components such as the decision generation network). For example, when using the DQN algorithm, the temporal difference error between the current network's Q-value prediction for (s, c) and the target Q-value based on (s', r) is calculated. This error is minimized through backpropagation, thereby updating the network parameters and enabling the model to learn to choose action 'a' that yields higher long-term cumulative rewards in similar (s, c) states. Through continuous online learning or periodic offline batch learning, the energy consumption decision model can adapt to different light strip characteristics, aging levels, and environmental changes, and its energy consumption optimization strategy continuously improves over time.

[0041] In summary, this invention, led by a server system, achieves adaptive and refined optimization of the dynamic energy consumption of long-cascaded LED light strips through a complete closed loop of real-time perception, scenario-based instruction generation, deep reinforcement learning model decision-making, precise control execution, effect evaluation, and feedback learning. This method not only improves energy efficiency but also enhances the interpretability of the system by outputting adjustment analysis results and ensures the long-term applicability and optimality of the strategy through continuous learning.

[0042] In this embodiment of the invention, the energy consumption decision model includes a state representation network, a feature transformation network, a decision generation network, and a control determination network set in a sequential manner. The state representation network is used to generate light strip state features by taking the operation perception information as input. The feature transformation network is used to convert the light strip state features into transformation result features in the decision feature representation domain. The decision generation network is constructed based on a reinforcement learning decision network trained on an initial policy and is used to output energy consumption decision features and energy consumption adjustment analysis results based on the transformation result features and the energy consumption control instructions. The control determination network is used to perform decision execution based on the energy consumption decision features to obtain the power control decision result.

[0043] The energy consumption decision model is obtained by combining the light strip operating status and energy consumption feedback data to perform phased training on the state representation network, feature transformation network and decision generation network of the initial decision model to perform light strip state feature and feature association constraints, and by combining the energy consumption control task dataset to perform phased training on the initial decision model to perform light strip state extraction and energy consumption control decision generation.

[0044] In an embodiment of the present invention, for example, the energy consumption decision model deployed in the server consists of four sequentially related neural network sub-modules, forming a data processing pipeline from perception to execution.

[0045] First, the state representation network directly processes the raw operational awareness information collected by the server. For example, when the server receives an array containing current, voltage, junction temperature, and brightness data for 256 series-connected LED modules, as well as ambient light and temperature, the network compresses and abstracts this data through convolutional and fully connected layers. Its output is no longer 256 independent values, but a comprehensive low-dimensional feature vector, such as [average brightness: 5200 nits, brightness variance: 150, maximum junction temperature: 67°C, system instantaneous efficiency: 0.82], which constitutes an essential description of the overall operating state of the LED strip.

[0046] Next, the feature transformation network receives this state feature. Due to the transitivity of electrical and thermal effects in cascaded LEDs, this network uses a recurrent neural network (RNN) structure to model the spatial dependencies of features. For example, it analyzes the junction temperature feature sequence, identifies that due to poor heat dissipation, the junction temperature of subsequent modules, starting from module 120, shows a progressively increasing trend, and encodes this spatial correlation pattern into the transformation result feature, providing context for decision-making.

[0047] Then, the decision generation network simultaneously receives the aforementioned transformation result features and the energy consumption regulation instruction description generated by the server (e.g., {Objective: "Peak power consumption not exceeding 2kW", Constraint: "Overall brightness maintained above 90%"}). This module, built based on a reinforcement learning policy network, uses the former as an environmental state representation and the latter as a task objective to perform policy inference. Its internal computation process is partially extracted, generating human-readable energy consumption regulation analysis results, such as: "It is recommended to reduce the driving current in the high-temperature trend zone (module 120-150) by 8%, which can reduce the risk of hotspots, and the total power consumption is expected to drop to 1.95kW, with a brightness loss of 4%." Simultaneously, the network outputs an abstract energy consumption decision feature vector, encoding the control intent of "mildly suppressing current in a specific sequence region."

[0048] Finally, the control decision network acts as a dedicated decoder, converting the abstract energy consumption decision feature vector into precise, deployable hardware instructions. For example, it decodes the decision features into a set of 256 specific pulse width modulation (PWM) duty cycle values, generating the final power control decision result: [0, 0.98, ..., 0.92 (number 120), 0.91, ..., 0.90 (number 150), ..., 255, 0.98].

[0049] The model's training is completed by the server in two phases. In the first phase, the server pre-trains the initial model using historically accumulated "light strip operation status and energy consumption feedback data." During this phase, the parameters of the decision generation network are fixed, and the state representation network and feature transformation network are trained to accurately predict energy consumption results under given states. This optimizes their ability to extract and transform features, thereby mastering the core representations and associated constraints of the light strip states. In the second phase, the server trains using an "energy consumption control task dataset" containing specific control task instructions and expected feedback. Here, the parameters of the first two networks are fixed, and the decision generation network and control decision network are optimized, enabling them to learn to directly output control instructions that efficiently complete the specified control tasks based on high-quality features. Through this phased training strategy, each submodule is gradually shaped into a reliable unit specializing in its respective area.

[0050] Specifically, in this embodiment, the specific architecture of each sub-network of the energy consumption decision model is as follows:

[0051] State representation network: A three-layer multilayer perceptron is used. The input layer has a dimension of N×6 (N is the number of serial units, and there are 6 features: current, voltage, junction temperature, brightness, ambient light, and ambient temperature). First, a fully connected layer maps the 6-dimensional features of each unit to 16-dimensional features using the ReLU activation function. Then, another fully connected layer compresses the 16-dimensional features into a 4-dimensional unit-level feature vector. Finally, global average pooling is performed on the 4-dimensional features of all units to output a fixed 128-dimensional global state feature vector.

[0052] Feature Transformation Network: A single-layer LSTM network is used. The spatially ordered 4-dimensional sequence of unit-level features output from the state representation network is taken as input. The hidden layer dimension of the LSTM is 64. The hidden state at the last time step of the LSTM is taken as the transformed feature, with a dimension of 64.

[0053] Decision generation network: A deep deterministic policy gradient algorithm network based on the Actor-Critic framework is employed. Wherein:

[0054] The Actor network (policy network) is a concatenation of 64-dimensional transformation result features and the embedded vector of energy consumption control instructions (the instructions are encoded into a 16-dimensional vector through a word embedding layer) (totaling 80 dimensions). The network contains two fully connected hidden layers (dimensions 128 and 64, respectively, using ReLU activation), and an N-dimensional output layer using the Tanh activation function. The output range is [-1, 1], representing the relative adjustment of the PWM duty cycle of the N units.

[0055] Energy consumption regulation analysis results: The activation values ​​of the first hidden layer (128 dimensions) of the Actor network are mapped to three interpretable scalars—the predicted average system brightness, total system power, and highest junction temperature—through an independent, lightweight fully connected analytical layer (output dimension 3).

[0056] Control decision network: This is a deterministic decoding layer. Its input is the N-dimensional relative adjustment value output by the Actor network. This layer performs the following operations: .in, This is the current duty cycle vector. To control the step size, it is set to 0.05. The clamp function limits the result to the physically feasible range of [0.1, 1.0] and finally outputs the target PWM value of N units.

[0057] In this embodiment of the invention, the step of loading the operation perception information and the energy consumption control instruction description into the energy consumption decision model to extract the light strip status and generate energy consumption control decisions, thereby obtaining the energy consumption adjustment analysis result and power control decision result of the operation perception information, can be implemented through the following example.

[0058] The operational perception information is loaded into the state representation network for state feature parsing to obtain the state features of the light strip;

[0059] The light strip state features are loaded into the feature transformation network for decision transformation, so as to transform the light strip state into the decision feature representation domain of the decision generation network and obtain the transformation result features;

[0060] The transformation result features and the energy consumption regulation instruction description are loaded into the decision generation network for decision inference to obtain the energy consumption decision features and the energy consumption regulation analysis results. The energy consumption regulation analysis results are obtained by feature transformation of the energy consumption decision features based on the hidden layer of the decision generation network.

[0061] The energy consumption decision features are loaded into the control decision network for decision-making and execution to obtain the power control decision result.

[0062] In this embodiment of the invention, the server performs the following specific process for decision generation: First, the server collects operational perception information, such as a real-time operating current array [20.1, 20.0, ..., 19.8] mA, an operating voltage array [3.15, 3.14, ..., 3.16] V, a junction temperature array [45, 46, ..., 52] °C, a brightness array, and ambient light intensity of 300 lux and ambient temperature of 28 °C, and organizes it into a standardized multidimensional array. The server inputs this array into the state characterization network. This network parses and fuses these high-dimensional raw data through a series of convolutional and fully connected layers, outputting a highly condensed LED strip state feature vector, for example [total power: 158.3 W, average junction temperature: 48.7 °C, brightness uniformity coefficient: 0.94, ambient light influence factor: 0.2]. This feature vector discards redundant details and summarizes the overall energy efficiency and thermal state of the system.

[0063] Next, the server inputs the aforementioned LED strip state feature vectors into the feature transformation network. The network's built-in sequence modeling unit (such as a Long Short-Term Memory (LSTM) network) analyzes the spatial sequence patterns inherent in the features. For example, it identifies the trend reflected by the junction temperature feature: starting from module 180, due to its location at the end of the heat dissipation duct, the junction temperature exhibits a stable, gradual increase. The network models and encodes this spatial relationship, outputting a transformed feature that enhances the contextual information. This feature is now within the representation domain desired by the decision generation network, facilitating its understanding of the spatial distribution and relationships of the states.

[0064] Then, the server loads the transformed feature along with a structured energy consumption control instruction description (e.g., {"Scenario": "Midday dynamic dimming", "Requirement": "Smoothly reduce the total system power consumption by 15% within 10 minutes, with brightness fluctuation rate less than 5%"}) into the decision generation network. This network guides strategy deduction based on the instruction. During its internal forward propagation, the network not only generates an abstract energy consumption decision feature vector representing the control intent, but also decodes the activation values ​​of a specific hidden layer to generate an energy consumption adjustment analysis result in natural language format. For example, the analysis result might be: "The plan is to adopt a regional gradient dimming strategy. In the first stage (first 5 minutes), the current in the low-temperature zone (modules 1-120) will be reduced by 2%, while the current in the high-temperature zone (modules 180-256) will remain unchanged, with an expected power consumption decrease of 7% and brightness fluctuation of 2%. The overall goal will be achieved after the second stage adjustment."

[0065] Finally, the server inputs the energy consumption decision feature vector into the control decision network. This network acts as a decoder, mapping the abstract features into precise executable parameters. For example, based on the intention in the decision features to "apply small, gradual current reductions in the low-temperature region," it outputs an array containing 256 specific PWM duty cycle settings, which is the final power control decision result: [{"id":1,"pwm_target":0.97},...,{"id":200,"pwm_target":1.00},...]. This result is then sent to each segment control unit for execution.

[0066] In this embodiment of the invention, the feature transformation network includes a first sequence modeling subnetwork and a second sequence modeling subnetwork. The step of loading the light strip state features into the feature transformation network for decision transformation, so as to transform the light strip state into the decision feature representation domain of the decision generation network and obtain the transformation result features, can be implemented through the following example.

[0067] The light strip state features are loaded into the first sequence modeling subnetwork for unit-level feature representation, so as to transform the light strip state into a continuous state representation domain and obtain a unit state representation vector.

[0068] The unit state representation vector is loaded into the second sequence modeling subnetwork to model the state association relationship, thereby obtaining the transformation result feature.

[0069] In an embodiment of the present invention, the specific workflow of the feature transformation network in the server is as follows:

[0070] The server inputs the LED strip state features output by the state characterization network, such as a feature sequence containing dimensions like "junction temperature of each module" and "brightness efficiency of each module," into the first sequence modeling sub-network. This sub-network typically consists of fully connected layers or one-dimensional convolutional layers, and its core function is to perform independent deep feature extraction and representation for each LED module unit. For example, for the i-th module, this sub-network maps its corresponding original state features (such as [junction temperature: 50°C, brightness efficiency: 0.85]) into a higher-dimensional, more information-rich unit state representation vector. This process transforms the discrete, potentially non-linear states of each module into a continuous, smooth, and more expressive feature representation domain, laying the foundation for subsequent analysis of the relationships between units.

[0071] Subsequently, the server inputs the sequence of unit state representation vectors for all modules of the entire light strip into the second sequence modeling sub-network in physical connection order. This sub-network typically employs a recurrent neural network (such as LSTM) or a self-attention mechanism (Transformer encoding layer) specifically for modeling the dynamic correlations and dependencies between the states of successive units in the sequence. For example, the network analyzes the sequence of unit state representation vectors to identify the trend of gradually increasing "heat accumulation component" in the feature vectors of subsequent modules starting from module 180 due to thermal diffusion effects; or to identify the pattern of gradually decreasing "effective voltage component" from the power supply end to the end module due to voltage drop in the power supply line. By modeling and encoding these spatial correlations, the second sequence modeling sub-network finally outputs the transformation result features. These features not only contain the independent state of each unit but also embed the contextual information of the entire cascaded link, making them fully adaptable to the representation domain required by the decision generation network for policy inference.

[0072] In this embodiment of the invention, the state representation network is obtained by performing phased training on a preset energy consumption control instruction feature parsing network and an operating state feature parsing network based on a state dependency modeling mechanism, based on the correspondence between the operating state of the light strip and the energy consumption control target.

[0073] In this embodiment of the invention, for example, the server trains the state representation network by performing state consistency determination in stages. This network consists of two sub-modules: an energy consumption control instruction feature parsing network and an operating state feature parsing network.

[0074] In the first stage, the server loads a historical dataset containing multiple sets of "LED strip operating status samples," "corresponding energy consumption control instructions," and "reference control targets" to be achieved under each instruction. For example, one set of data might consist of: status samples (current, voltage, and temperature of 256 modules), instructions ("reduce total power consumption by 10%), and a reference target ("the adjusted total power value should be 144W"). The server inputs the same instruction into the instruction feature parsing network and the corresponding status samples into the operating status feature parsing network. The instruction network outputs the semantic feature vector of the instruction (e.g., expressing "global power consumption reduction of 10%)," and the status network outputs the status feature vector of the sample. The training objective is to align these two feature vectors in the semantic space, i.e., by calculating a consistency loss (e.g., cosine similarity loss) and backpropagating to update the parameters of both networks, enabling the networks to learn to extract features strongly correlated with the control target from the states.

[0075] In the second stage, the server fixes the parameters of the pre-trained runtime state feature parsing network, making it a stable feature extractor. During subsequent overall model training, this network will continuously provide the decision generation network with high-quality, control-target-aligned light strip state feature representations.

[0076] In this embodiment of the invention, the method further includes:

[0077] Acquire the operating status and energy consumption feedback data of the light strips across operating scenarios, the operating status and energy consumption feedback data of the light strips in the target operating scenario, and the control feedback sample set;

[0078] Based on the light strip operation status and energy consumption feedback data across the operation scenarios and the light strip operation status and energy consumption feedback data of the target operation scenario, the state representation network, feature transformation network and decision generation network of the initial decision model are trained in the first stage of light strip state features and feature association constraints. During the training process, the network parameters of the decision generation network are fixed and the network parameters of the state representation network and the feature transformation network are optimized and updated until the first optimization convergence state is reached.

[0079] Based on the energy consumption control task dataset formed by the light strip operation status and energy consumption feedback data of the target operation scenario and the control feedback sample set, the initial decision model that meets the first optimization convergence state is trained in the second stage of light strip state extraction and energy consumption control decision generation. During the training process, the network parameters of the state representation network and the feature transformation network are fixed and the network parameters of the decision generation network and the control judgment network are optimized and updated until the second optimization convergence state is reached.

[0080] The initial decision model that conforms to the second optimization convergence state is determined as the energy consumption decision model.

[0081] In this embodiment of the invention, for example, the server first acquires three types of data from a historical database and an experimental environment. The first type is the operating status and energy consumption feedback data of the light strip across operating scenarios. For example, it collects the operating logs of the same model of light strip under two completely different scenarios: "indoor gallery lighting" and "outdoor architectural outline decoration". Each data entry may include: a status sample (electrical parameters of each module), a control command at that time (such as "maintain the color temperature of the artistic lighting"), and a reference value of the actual control result achieved after executing the command (such as the final average color temperature value). The second type is the operating status and energy consumption feedback data of the light strip in the target operating scenario. For example, data collected specifically for the target scenario of "urban tunnel lighting" also includes status samples, commands (such as "increase the brightness of the entrance section by 20%), and a reference value of the result. The third type is a set of control feedback samples. This part of the data may include ideal control cases that are manually labeled or generated through high-precision simulation, such as a clear control command ("synchronously set the brightness of modules 50 to 100 to 80% of the standard value") and its corresponding error-free control feedback true value.

[0082] In the first phase of training, the server uses the first two types of data (cross-scenario and target scenario data) to train the initial decision model. The core objective of this phase is to enable the state representation network and feature transformation network to learn to extract state features that are strongly correlated with the energy consumption control target and have good generalization ability, and to understand the correlation constraints between features. To achieve this objective, the server fixes the network parameters of the decision generation network, preventing it from participating in updates and treating it only as a fixed "feature estimator".

[0083] During training, the server takes a single data point, such as one from an outdoor scene: a state sample (at which the junction temperatures of various modules were generally high due to the high ambient temperature), a control instruction ("maintain the brightness of the marker under strong sunlight interference"), and a result reference value (the final maintained brightness value). The server inputs the state sample into the state representation network to obtain the light strip state features. This feature is then input into the feature transformation network to obtain the transformed feature. Next, the server inputs the transformed feature along with the control instruction into a fixed decision generation network. Based on its current fixed strategy, the decision generation network outputs a predicted energy consumption adjustment result (e.g., it predicts the brightness value achievable through a certain adjustment). The server calculates the deviation (e.g., mean squared error) between this predicted result and the actual result reference value. This deviation is the first optimization deviation. Based on this deviation, the server updates only the network parameters of the state representation network and the feature transformation network using a backpropagation algorithm. Through repeated training on a large amount of such data, the state representation network and the feature transformation network are forced to learn how to transform the original state into a feature representation that enables the subsequent fixed decision network to make accurate predictions. This process continues until the model's prediction bias on the validation set no longer decreases significantly, reaching the first optimization convergence state. At this point, the first two networks are able to extract high-quality, goal-oriented light strip state features.

[0084] In the second stage, the server fixes the network parameters of the state representation network and feature transformation network, which have reached the first optimization convergence state. Subsequently, it constructs an energy consumption control task dataset, which is mainly composed of data from the target operating scenario (ensuring scenario specificity) and a set of control feedback samples (providing accurate examples).

[0085] During this training phase, the server takes a task data set, such as data from a tunnel lighting scenario: a state sample (low luminance sensor reading at the tunnel entrance) and a control command ("Increase the illuminance at the entrance to the standard value"). The server processes the state sample using a fixed state representation network and feature transformation network to obtain the transformation result features. Then, this feature, along with the control command, is input into the decision generation network. At this point, the parameters of the decision generation network are updatable. The network performs decision deduction and outputs energy consumption decision features and energy consumption adjustment analysis results. Simultaneously, the control decision network receives the energy consumption decision features and outputs specific power control decision results (such as a PWM command array).

[0086] For this data, the server can find the corresponding ground truth value for regulation feedback from the control feedback sample set (e.g., a set of precise PWM values ​​that achieve the target brightness). The server compares the power control decision result output by the network with this ground truth value and calculates the deviation. Simultaneously, it compares the energy consumption regulation analysis result output by the network with the possible second regulation target reference description in the data. Combining these deviations, a second optimization deviation is generated. Based on this deviation, the server updates only the network parameters of the decision generation network and the control decision network using the backpropagation algorithm. Through training on a large amount of task data, the decision generation network learns how to generate the correct control strategy based on high-quality state features and explicit instructions, while the control decision network learns how to accurately decode this strategy into hardware instructions. When the model's performance on the task dataset stabilizes, the second optimization convergence state is reached.

[0087] Finally, the server deploys the initial decision model that meets the second optimal convergence state as a usable energy consumption decision model for online dynamic optimization.

[0088] In this embodiment of the invention, the light strip operating status and energy consumption feedback data corresponding to the cross-operation scenario includes a first light strip operating status sample, a first energy consumption control instruction information, and a first control target reference description corresponding to the first light strip operating status sample. The light strip operating status and energy consumption feedback data corresponding to the target operating scenario includes a second light strip operating status sample, a second energy consumption control instruction information, and a second control target reference description corresponding to the second light strip operating status sample. The first energy consumption control instruction information and the second energy consumption control instruction information are used to form the decision guidance information required by the decision generation network when making decision inferences. The first control target reference description is a control result reference value corresponding to the control state generation based on the first energy consumption control instruction information, and the second control target reference description is a control result reference value corresponding to the control state generation based on the second energy consumption control instruction information. The first stage of training is implemented through the following process and can be executed through the following examples.

[0089] Using the first light strip operating state sample or the second light strip operating state sample as input to the state representation network, state feature parsing is performed to obtain the first sample light strip state feature corresponding to the first light strip operating state sample or the second sample light strip state feature corresponding to the second light strip operating state sample.

[0090] The first sample light strip state feature or the second sample light strip state feature is used as the input of the feature transformation network to perform decision transformation, so as to transform the first sample light strip state feature and the second sample light strip state to the decision feature representation domain of the decision generation network respectively, and obtain the first transformation result feature and the second transformation result feature;

[0091] The first sample adjustment analysis result and the second sample adjustment analysis result are obtained by using the first conversion result feature and the first energy consumption control instruction information as the associated constraint unit, or the second conversion result feature and the second energy consumption control instruction information as the input of the decision generation network.

[0092] The first optimization deviation amount is determined based on the deviation between the first sample adjustment analysis result and the first control target reference description, and the deviation between the second sample adjustment analysis result and the second control target reference description;

[0093] The network parameters of the decision generation network are fixed, and the state representation network and the feature transformation network are trained according to the first optimization bias until the first optimization convergence state is reached.

[0094] In an embodiment of the invention, exemplarily, the server first reads a training sample across operating scenarios from storage. This sample specifically includes: a first LED strip operating status sample, such as a snapshot of instantaneous data from an "indoor art gallery lighting system," recording an array of the operating states of 128 LED modules at a specific moment, including the real-time current (approximately 18mA), operating voltage (3.2V), junction temperature (ranging from 42°C to 48°C), output brightness value (used to illuminate paintings, averaging approximately 800 nits), ambient light intensity (50 lux from auxiliary lighting in the exhibition hall), and ambient temperature (constant 22°C) for each module. The first energy consumption control instruction information, i.e., the control command issued by the system at that time, such as a structured instruction: {"Operation": "Dynamically Maintain", "Target Parameter": "Painting Area Color Temperature", "Target Value": 3500K, "Priority": "High"}. The first control target reference description is the actual result measured and recorded by a high-precision color temperature sensor after the instruction is executed, for example, {“Status Achieved”: “Successful”, “Final Average Color Temperature”: 3498K, “Stability Index”: 0.95}. This reference description is the true value that the network learning needs to approximate.

[0095] The server inputs this sample into the training process. The first step involves loading the first LED strip's operating state sample (i.e., the original numerical arrays of current, voltage, temperature, etc.) into the input layer of the state representation network. The network performs forward propagation calculations, parsing, filtering, and fusing this high-dimensional, heterogeneous raw data through its internal weight parameters. Its output is no longer 128 independent data points, but a condensed, low-dimensional feature vector of the first sample LED strip's state. For example, this feature vector can be quantified as [System average junction temperature: 45.3°C, color temperature deviation coefficient: 0.02, ambient light interference level: low, power supply stability index: 0.98]. This feature aims to capture the essential system state most relevant to the control objective of "maintaining color temperature."

[0096] The second step involves the server inputting the first sample LED strip state feature vector into the feature transformation network. This network further processes this feature, focusing on modeling the relationships between feature elements and along the spatial sequence of the LED strip. For example, through its internal recurrent connections, it can analyze the small but crucial thermal coupling effect between the junction temperature of the "painting area" (corresponding to a specific module sequence) and the junction temperature of the "non-painting area," an effect that affects color temperature consistency. The network encodes this correlation information and outputs the first transformation result feature. This feature is now in a representation domain more aligned with the decision logic; it is no longer merely a static description of the state but includes contextual information about the dynamic relationships between states.

[0097] Third, the server constructs an associative constraint unit. It concatenates or fuses the first transformation result features with the first energy consumption control instruction information (i.e., the instruction to "dynamically maintain color temperature 3500K"), forming a joint representation that includes the "current environmental state" and the "target task." Then, the server inputs this joint representation into the decision generation network. At this point, all parameters of the decision generation network are fixed, and its weights remain unchanged throughout the current training iteration. It acts like a fixed function, performing a forward inference based on the input features and instructions, and outputting a first sample adjustment analysis result. This result is the network's prediction of the outcome after executing the instruction based on its current fixed policy; for example, it could output: {"Predicted color temperature": 3480K, "Predicted stability": 0.93, "Estimated main interference source": "thermal fluctuation in region 3"}.

[0098] The fourth step involves the server calculating the loss. It compares the first sample adjustment analysis result (predicted value) output by the decision generation network with the first adjustment target reference description (true value) provided in the sample, item by item. For example, it calculates the difference between the predicted color temperature of 3480K and the true color temperature of 3498K, and the difference between the predicted stability of 0.93 and the true stability of 0.95. The server uses a preset loss function (such as the mean squared error function) to synthesize these deviations into a scalar value, namely the first optimization deviation. For example, the calculated loss value is 15.2.

[0099] Fifth, the server performs backpropagation and parameter updates. Since the network parameters of the decision generation network are fixed, the gradients calculated during backpropagation are not used to update its parameters as they flow through the network. The gradients continue to backpropagate to its upstream feature transformation network and state representation network. Based on the calculated first optimization bias (loss value 15.2) and a gradient descent algorithm (such as the Adam optimizer), the server only updates the weight parameters of these two networks. The goal of optimization is to adjust the parameters of these two networks so that they can extract better features for the fixed decision network next time, thereby making the predictions made by the fixed decision network (adjusting the analytical results) closer to the true reference target.

[0100] The above process iterates over a large number of samples across operating scenarios (such as data from different scenarios like outdoor decorative lighting and shopping mall basic lighting) and samples from the target operating scenario (e.g., second light strip operating status samples, second energy consumption control command information, and second control target reference descriptions collected specifically for the "underground parking lot lighting" scenario). The server continuously uses the prediction error of the fixed decision network to correct the state representation and feature transformation networks. As training progresses, the state representation network learns to ignore noise irrelevant to various control targets and extracts key state factors common to multiple scenarios; the feature transformation network learns to more accurately characterize the physical constraints between states. When the training loss no longer decreases significantly on the validation set and the fluctuations tend to stabilize, the server determines that the model has reached the first optimal convergence state. At this point, the first two networks have become a reliable preprocessing module that can provide high-quality, goal-oriented feature representations for subsequent decisions.

[0101] In this embodiment of the invention, the light strip operating status and energy consumption feedback data of the target operating scenario also include the status type identifier of the second light strip operating status sample. The control feedback sample set includes sample control instructions and control feedback truth values ​​corresponding to the sample control instructions. The second stage training is implemented through the following process and can be executed through the following example.

[0102] Using the second light strip operating state sample as the input of the state representation network that conforms to the first optimized convergence state, state feature analysis is performed to obtain the third sample light strip state feature;

[0103] The third sample light strip state features are used as input to the feature transformation network that conforms to the first optimized convergence state to perform decision transformation, and the third transformation result features are obtained.

[0104] Using the associated constraint unit formed by the second conversion result features and the second energy consumption control instruction information, or the sample control instruction as the input of the decision generation network, decision deduction is performed to obtain the third sample energy consumption decision features, the third sample adjustment analysis results, and the control feedback conclusion.

[0105] The energy consumption decision features of the third sample are loaded into the control and decision network for decision execution to obtain the sample decision result;

[0106] The second optimization deviation is determined based on the third sample adjustment analysis result, the second adjustment target reference description, the adjustment feedback conclusion, the adjustment feedback truth value, the sample judgment result, and the state type identifier;

[0107] The network parameters of the state representation network and the feature transformation network are fixed, and the decision generation network and the control decision network are trained according to the second optimization deviation until the second optimization convergence state is reached.

[0108] In an embodiment of the present invention, the specific process of the server performing the second stage of training is as follows: this stage focuses on training the decision generation network and the control decision network to perform specific control tasks based on the fixed preceding network.

[0109] The server first loads a sample data set from the target operating scenario. This data includes, in addition to samples of the second light strip's operating status (e.g., an array recording the operating parameters of all modules of the underground parking lot light strip during the evening rush hour), second energy consumption control command information (e.g., {"Task": "Emergency Brightness Up", "Range": "B Zone Lane", "Target Brightness Increase": 40%}), and a second control target reference description (e.g., records showing a 38% increase in average brightness in the area after execution), a status type identifier. This identifier is used to label specific attributes of the sample, such as "normal_operation" or "voltage_sag_recovery," providing additional contextual information for training.

[0110] At the same time, the server matches or selects relevant samples from the control feedback sample set, which contains explicit sample control instructions (e.g., "set the PWM duty cycle of modules 201 to 250 to 0.95") and their corresponding, precise control feedback truth values ​​(e.g., a set of values ​​describing the expected brightness and power consumption of each module after executing the instruction).

[0111] At the start of training, the server inputs the second light strip operating status sample into the state representation network, which is already fixed and in the first optimization convergence state. This network performs forward computation based on its trained parameters and outputs the third sample light strip state features. For example, for the parking lot's evening rush hour state, it can output a feature vector [Average current load in region B: High, power supply voltage fluctuation index: 0.05, overall heat accumulation trend: Smooth].

[0112] Next, this feature is fed into a feature transformation network that is also fixed. This network performs sequence association modeling on it and outputs a third transformation result feature. For example, this feature could encode the comprehensive information that "the brightness of the lane module in area B is currently uniform, but the power supply voltage fluctuates slightly."

[0113] Subsequently, the server performs decision-making simulations. It has two data input methods: one is to combine the third conversion result features with the second energy consumption control instruction information ("Emergency brightening of area B by 40%) into an associated constraint unit and input it into the decision generation network; the other is to directly input the sample control instruction from the control feedback sample set ("Set module 201-250 PWM to 0.95") as the instruction condition, along with the third conversion result features. The decision generation network (whose parameters are trainable at this point) performs strategy calculations and outputs three parts: an abstract third-sample energy consumption decision feature vector; a third-sample adjustment analysis result for human analysis (e.g., "It is recommended to adopt a two-stage voltage boosting strategy, first increasing the power supply voltage to compensate for line loss, and then fine-tuning the PWM"); and a control feedback conclusion (e.g., "The probability of achieving the predicted target is 92%").

[0114] The server then inputs the energy consumption decision feature vector of the third sample into the control decision network. This network decodes it into specific hardware operation instructions, i.e., the sample decision result. For example, it outputs a PWM instruction array: ...,{"id":201,"pwm":0.94},...,{"id":250,"pwm":0.95}).

[0115] Finally, the server calculates the second optimization bias. This is a comprehensive loss determined by multiple biases:

[0116] Analysis result deviation: Compare the predicted value in the third sample adjustment analysis result (e.g., predicted brightness increase of 42%) with the actual value in the second adjustment target reference description (actual increase of 38%).

[0117] Feedback conclusion bias: Compare the control feedback conclusion (predicted probability 92%) with the actual situation (if the sample identifier shows that the task was successful, the true probability should be 1).

[0118] Control action deviation: The difference between the sample judgment result (PWM array output by the network) and the true value of the control feedback (ideal PWM array given in the set).

[0119] State identification constraints: Specific constraints are introduced based on the state type identifier. For example, if the identifier is "voltage_sag_recovery", an extra term is added to the loss function to penalize aggressive control strategies that could lead to further voltage instability.

[0120] The server aggregates all the above deviations and obtains the final second-optimization deviation through weighted summation. During backpropagation, the server keeps the network parameters of the state representation network and feature transformation network fixed, and updates and optimizes the parameters of the decision generation network and control decision network only based on the gradient calculated from this deviation. This process is iterated repeatedly on the large energy consumption control task dataset until the model performance stabilizes and reaches the second-optimization convergence state, thus completing the training of the entire model.

[0121] In this embodiment of the invention, both the first energy consumption control instruction information and the second energy consumption control instruction information include multi-level control requests, which include parameter description requests, state evaluation requests, and strategy deduction requests.

[0122] The parameter description request is used to trigger the decision generation network to extract and quantify parameters. The first or second control target reference description corresponding to the parameter description request has a deterministic quantitative reference. The state evaluation request is used to trigger the decision generation network to perform a comprehensive state evaluation. The first or second control target reference description corresponding to the state evaluation request has a non-quantitative evaluation reference. The strategy deduction request is used to trigger the decision generation network to derive and generate a strategy for the light strip state.

[0123] In this embodiment of the invention, the energy consumption control instruction description received and processed by the server is, for example, a structured information body containing multi-level task definitions. Its core consists of three levels of control requests. Each request triggers a decision generation network to execute a specific type of inference task and corresponds to different forms of control target reference descriptions, thereby guiding the model's training and decision-making.

[0124] Parameter description requests are used to trigger the decision generation network to accurately extract and report quantitative parameters of the current operating state. When the server generates or receives such instructions, its goal is to calculate specific system indicators from complex sensory information. For example, when monitoring the lighting strips in the atrium of a large shopping mall, the server can generate an instruction: {"Request Type": "Parameter Description", "Request Content": "Report the current total real-time power consumption and average junction temperature of the entire system"}. This instruction is loaded into the decision generation network as the first or second energy consumption control instruction. When processing this instruction, the network drives its internal computing units to perform quantification operations similar to "summation" and "average" based on the input state characteristics, and outputs the first or second sample adjustment analysis result, such as {"Total Power Consumption": 1250W, "Average Junction Temperature": 62°C}. During the training phase, the reference description of the first or second control target corresponding to this instruction is reference data with definite values ​​measured by high-precision meters and temperature sensors, such as {"Reference Total Power Consumption": 1247W, "Reference Average Junction Temperature": 61.8°C}. The server trains the network to perform accurate numerical extraction and description by calculating the mean square error between the network output and the reference data.

[0125] State assessment requests trigger the decision generation network to perform a comprehensive, non-quantitative assessment and judgment of the system state. These requests focus on the qualitative characteristics or health of the state. For example, for the same shopping mall light strip, the server might issue a maintenance instruction: {"Request Type": "State Assessment", "Request Content": "Assess the thermal equilibrium state level between each series unit"}. When processing this instruction, the decision generation network needs to analyze the distribution, gradient, and stability of the junction temperatures of all modules, rather than simply calculating averages. It outputs a conclusion based on learned assessment criteria, such as {"Thermal Equilibrium State": "Slightly Uneven", "Main Hotspot Area": ​​"Southeast corner module sequence 80-95", "Recommendation": "Check the heat dissipation of this area"}. During the training phase, the corresponding control target reference description for such instructions can be a non-quantitative assessment reference labeled by domain experts based on complete data, such as {"Expert Assessment Result": "Local overheating risk exists"} or a label indicating the level (e.g., "Level 2 Warning"). The server trains the network to form reliable evaluation logic by comparing the evaluation conclusions output by the network with the references annotated by experts, using classification loss or ranking loss.

[0126] The strategy deduction request is the highest-level request, directly triggering the decision generation network to deduce and generate future-state-oriented strategies. This is the core of the model's dynamic energy consumption optimization. For example, in a nighttime energy-saving scenario, the server generates the following instruction: {"Request Type": "Strategy Deduction", "Request Content": "Deduce a progressive dimming strategy that can reduce total energy consumption by 15% within 30 minutes, while ensuring that the brightness of the main road lighting is not lower than the minimum standard"}. Upon receiving this instruction, the decision generation network needs to perform multi-step, forward-looking planning based on the current brightness, power consumption, ambient light, and other state characteristics. It outputs specific strategy deduction results as part of the adjustment analysis results, for example: {"Deduction Strategy": "First stage (0-10 minutes): Reduce the brightness of non-core areas by 8%; Second stage (10-25 minutes): Fine-tune the core area based on sensor feedback, with an estimated total energy saving of 16.2%"}. During training, the corresponding reference description can be a successful control sequence example or the optimal strategy result calculated through advanced simulation. The server trains the network to generate control strategies that satisfy constraints and are energy efficient by using reinforcement learning rewards or imitation learning.

[0127] By integrating these three levels of requests, the server's control instructions can precisely guide the decision generation network to execute the complete cognitive chain from data extraction and state diagnosis to policy generation, thereby achieving highly intelligent and targeted dynamic energy consumption optimization.

[0128] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned dynamic energy consumption optimization method for long cascaded LED light strips based on deep reinforcement learning. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0129] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. A dynamic energy consumption optimization method for long cascaded LED light strips based on deep reinforcement learning, characterized in that, include: The operation sensing information of the target long cascaded LED light strip is collected. The operation sensing information includes the real-time current, operating voltage, junction temperature, output brightness value of each series unit of the light strip, as well as the illuminance and ambient temperature of the environment in which the light strip is located. Based on the current application scenario and preset energy consumption constraint target of the target long cascaded LED light strip, a corresponding energy consumption control instruction description is generated; The operational sensing information and the energy consumption control instruction description are loaded into the energy consumption decision model to extract the light strip status and generate energy consumption control decisions, thereby obtaining the energy consumption adjustment analysis results and power control decision results of the operational sensing information; the energy consumption control instruction description is used to provide the decision guidance information required by the energy consumption decision model to perform light strip status extraction and energy consumption control decision generation. The power control decision results are sent to the segmented control unit of the target long cascaded LED light strip to execute the corresponding power adjustment operation; Collect feedback data on the operating status of the light strip after the power adjustment operation is executed, and compare the feedback data with the constraint target in the description of the energy consumption control instruction to generate a reinforcement learning reward signal; The operational perception information, energy consumption regulation command description, power control decision results, feedback data, and reinforcement learning reward signals are stored as training samples for reinforcement learning iterative updates of the energy consumption decision model; wherein, the energy consumption decision model is constructed based on a deep reinforcement learning algorithm.

2. The method according to claim 1, characterized in that, The energy consumption decision model includes a state representation network, a feature transformation network, a decision generation network, and a control decision network set in a sequential manner. The state representation network is used to generate light strip state features by taking the operation perception information as input. The feature transformation network is used to convert the light strip state features into transformation result features in the decision feature representation domain. The decision generation network is constructed based on a reinforcement learning decision network trained on an initial policy and is used to output energy consumption decision features and energy consumption adjustment analysis results based on the transformation result features and the energy consumption control instructions. The control decision network is used to make decisions based on the energy consumption decision features to obtain the power control decision result. The energy consumption decision model is obtained by combining the light strip operating status and energy consumption feedback data to perform phased training on the state representation network, feature transformation network and decision generation network of the initial decision model to perform light strip state feature and feature association constraints, and by combining the energy consumption control task dataset to perform phased training on the initial decision model to perform light strip state extraction and energy consumption control decision generation.

3. The method according to claim 2, characterized in that, The step of loading the operational sensing information and the energy consumption regulation command description into the energy consumption decision model to extract the light strip status and generate energy consumption control decisions, and obtaining the energy consumption regulation analysis results and power control decision results of the operational sensing information, includes: The operational perception information is loaded into the state representation network for state feature parsing to obtain the state features of the light strip; The light strip state features are loaded into the feature transformation network for decision transformation, so as to transform the light strip state into the decision feature representation domain of the decision generation network and obtain the transformation result features; The transformation result features and the energy consumption regulation instruction description are loaded into the decision generation network for decision inference to obtain the energy consumption decision features and the energy consumption regulation analysis results. The energy consumption regulation analysis results are obtained by feature transformation of the energy consumption decision features based on the hidden layer of the decision generation network. The energy consumption decision features are loaded into the control decision network for decision-making and execution to obtain the power control decision result.

4. The method according to claim 3, characterized in that, The feature transformation network includes a first sequence modeling subnetwork and a second sequence modeling subnetwork. The step of loading the light strip state features into the feature transformation network for decision transformation, thereby transforming the light strip state to the decision feature representation domain of the decision generation network, yields the following transformation result features: The light strip state features are loaded into the first sequence modeling subnetwork for unit-level feature representation, so as to transform the light strip state into a continuous state representation domain and obtain a unit state representation vector. The unit state representation vector is loaded into the second sequence modeling subnetwork to model the state association relationship, thereby obtaining the transformation result feature.

5. The method according to claim 2, characterized in that, The state representation network is obtained by performing phased training on a preset energy consumption control instruction feature parsing network and an operation state feature parsing network based on a state dependency modeling mechanism, based on the correspondence between the operating state of the light strip and the energy consumption control target.

6. The method according to claim 2, characterized in that, The method further includes: Acquire the operating status and energy consumption feedback data of the light strips across operating scenarios, the operating status and energy consumption feedback data of the light strips in the target operating scenario, and the control feedback sample set; Based on the light strip operation status and energy consumption feedback data across the operation scenarios and the light strip operation status and energy consumption feedback data of the target operation scenario, the state representation network, feature transformation network and decision generation network of the initial decision model are trained in the first stage of light strip state features and feature association constraints. During the training process, the network parameters of the decision generation network are fixed and the network parameters of the state representation network and the feature transformation network are optimized and updated until the first optimization convergence state is reached. Based on the energy consumption control task dataset formed by the light strip operation status and energy consumption feedback data of the target operation scenario and the control feedback sample set, the initial decision model that meets the first optimization convergence state is trained in the second stage of light strip state extraction and energy consumption control decision generation. During the training process, the network parameters of the state representation network and the feature transformation network are fixed and the network parameters of the decision generation network and the control judgment network are optimized and updated until the second optimization convergence state is reached. The initial decision model that conforms to the second optimization convergence state is determined as the energy consumption decision model.

7. The method according to claim 6, characterized in that, The light strip operating status and energy consumption feedback data corresponding to the cross-operation scenario includes a first light strip operating status sample, a first energy consumption control instruction information, and a first control target reference description corresponding to the first light strip operating status sample. The light strip operating status and energy consumption feedback data corresponding to the target operating scenario includes a second light strip operating status sample, a second energy consumption control instruction information, and a second control target reference description corresponding to the second light strip operating status sample. The first energy consumption control instruction information and the second energy consumption control instruction information are used to form the decision guidance information required by the decision generation network when making decision inferences. The first control target reference description is the control result reference value corresponding to the control state generation based on the first energy consumption control instruction information. The second control target reference description is the control result reference value corresponding to the control state generation based on the second energy consumption control instruction information. The first phase of training is implemented through the following process: Using the first light strip operating state sample or the second light strip operating state sample as input to the state representation network, state feature parsing is performed to obtain the first sample light strip state feature corresponding to the first light strip operating state sample or the second sample light strip state feature corresponding to the second light strip operating state sample. The first sample light strip state feature or the second sample light strip state feature is used as the input of the feature transformation network to perform decision transformation, so as to transform the first sample light strip state feature and the second sample light strip state to the decision feature representation domain of the decision generation network respectively, to obtain the first transformation result feature and the second transformation result feature; The first sample adjustment analysis result and the second sample adjustment analysis result are obtained by using the first conversion result feature and the first energy consumption control instruction information to form an association constraint unit, or the second conversion result feature and the second energy consumption control instruction information to form an association constraint unit as the input of the decision generation network. The first optimization deviation amount is determined based on the deviation between the first sample adjustment analysis result and the first control target reference description, and the deviation between the second sample adjustment analysis result and the second control target reference description; The network parameters of the decision generation network are fixed, and the state representation network and the feature transformation network are trained according to the first optimization bias until the first optimization convergence state is reached.

8. The method according to claim 7, characterized in that, The target operating scenario's light strip operating status and energy consumption feedback data also includes the status type identifier of the second light strip operating status sample. The control feedback sample set includes sample control commands and the control feedback truth values ​​corresponding to the sample control commands. The second stage of training is implemented through the following process, including: Using the second light strip operating state sample as the input of the state representation network that conforms to the first optimized convergence state, state feature analysis is performed to obtain the third sample light strip state feature; The third sample light strip state features are used as input to the feature transformation network that conforms to the first optimized convergence state to perform decision transformation, and the third transformation result features are obtained. Using the associated constraint unit formed by the second conversion result features and the second energy consumption control instruction information, or the sample control instruction as the input of the decision generation network, decision deduction is performed to obtain the third sample energy consumption decision features, the third sample adjustment analysis results, and the control feedback conclusion. The energy consumption decision features of the third sample are loaded into the control and decision network for decision execution to obtain the sample decision result; The second optimization deviation is determined based on the third sample adjustment analysis result, the second adjustment target reference description, the adjustment feedback conclusion, the adjustment feedback truth value, the sample judgment result, and the state type identifier; The network parameters of the state representation network and the feature transformation network are fixed, and the decision generation network and the control decision network are trained according to the second optimization deviation until the second optimization convergence state is reached.

9. The method according to claim 7, characterized in that, Both the first energy consumption control instruction information and the second energy consumption control instruction information include multi-level control requests, which include parameter description requests, state evaluation requests, and strategy deduction requests. The parameter description request is used to trigger the decision generation network to extract and quantify parameters. The first or second control target reference description corresponding to the parameter description request has a deterministic quantitative reference. The state assessment request is used to trigger the decision generation network to perform a comprehensive state assessment. The first or second control target reference description corresponding to the state assessment request has a non-quantitative assessment reference. The strategy deduction request is used to trigger the decision generation network to derive and generate a strategy for the light strip state.

10. A server system, characterized in that, Includes a server, the server being used to perform the method according to any one of claims 1-9.