Intelligent lamp strip dynamic regulation and control method and system based on reinforcement learning
By adopting knowledge distillation and hierarchical models in the smart light strip system, combined with federated learning and differential privacy protection, the problems of high computational complexity and insufficient privacy protection on edge devices are solved, efficient and secure personalized lighting control is achieved, and user experience and system transparency are improved.
Patent Information
- Application Number
- CN202510782664.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing reinforcement learning-based smart light strip control systems are difficult to run in real time on resource-constrained edge devices, and suffer from problems such as insufficient user privacy data protection, poor balance of multi-user needs, and insufficient data usage transparency.
Knowledge distillation technology is used to build a lightweight reinforcement learning model, which is divided into a shared layer and a personalized layer. Federated learning and differential privacy protection are combined to achieve multi-device collaborative learning through a secure aggregation protocol, and verifiable computational proofs are generated to ensure that data usage complies with the authorized scope.
Efficiently running reinforcement learning models on edge devices provides flexible privacy protection and verifiable data usage processes, improves system response speed and user satisfaction, reduces the risk of privacy leakage, balances the lighting preferences of multiple users, and improves system transparency and user trust.
Smart Images

Figure CN120671187A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent lighting control technology, and more specifically, to a method and system for dynamically controlling intelligent light strips based on reinforcement learning. Background Art
[0002] With the rapid development of IoT technology and the growing smart home market, smart light strips, as important lighting control devices, have gained widespread adoption in homes, offices, and commercial spaces. By adjusting parameters such as brightness, color temperature, and color, these light strips provide users with a personalized lighting experience, meeting lighting needs in diverse scenarios.
[0003] Currently, smart light strip control methods primarily include preset mode control, sensor-based automatic control, and machine learning-based intelligent control. Among these, reinforcement learning-based intelligent control methods have become a research hotspot due to their ability to continuously optimize control strategies through interaction with the environment.
[0004] However, the prior art has the following deficiencies:
[0005] First, existing smart light strip control systems based on reinforcement learning usually use complex deep neural network models, which have large parameter scales and computational complexity. For example, typical deep reinforcement learning models often contain millions of parameters, requiring hundreds of MB of storage space and a large amount of floating-point computing resources. This high computing demand makes it difficult for the model to run in real time on resource-constrained edge devices (such as smart home gateways, light controllers, etc.), resulting in high system response delays and affecting user experience. In order to provide personalized lighting services, the system needs to collect and analyze user behavior data, preference information and other personal privacy data. Traditional solutions usually upload this data to the cloud for centralized processing and model training. This approach has obvious risks of privacy leakage. User lighting preference data often contains a wealth of sensitive information such as living habits, work and rest patterns, activity patterns, etc. Once leaked, it may pose a serious threat to user privacy. Existing privacy protection solutions usually adopt a unified protection strategy. Applying the same level of privacy protection measures to all types of data is a "one-size-fits-all" approach with obvious flaws. For less sensitive data (such as basic brightness preferences), excessive privacy protection will reduce data availability and affect the quality of personalized services. For highly sensitive data (such as detailed behavioral patterns), a fixed level of protection may not be sufficient to provide adequate privacy protection. In environments shared by multiple users (such as home public areas and open office spaces), different users often have different or even conflicting preferences for lighting environments. Existing systems lack an effective multi-user demand balancing mechanism, making it difficult to coordinate the lighting preferences of multiple users while protecting the privacy of each user, resulting in unsatisfactory lighting control effects in shared spaces. Existing systems also lack transparency and verifiability in data usage. Users often cannot understand how their personal data is used, nor can they verify whether the system is using the data within the authorized scope, which further exacerbates users' concerns about privacy protection.
[0006] Therefore, a new technical solution is needed that can efficiently run reinforcement learning models on resource-constrained edge devices while providing flexible privacy protection mechanisms and verifiable data usage processes to solve the above technical problems. Summary of the Invention
[0007] The present invention provides a method and system for dynamic control of smart light strips based on reinforcement learning, which solves technical problems in related technologies such as high computational complexity, difficulty in running on edge devices, and insufficient protection of user privacy data.
[0008] The present invention provides a method for dynamic control of an intelligent light strip based on reinforcement learning, comprising:
[0009] Build a lightweight reinforcement learning model using knowledge distillation technology, compressing the complex teacher model into a student model that can run on edge devices;
[0010] The lightweight reinforcement learning model is divided into a shared layer and a personalized layer. The shared layer extracts general lighting features, and the personalized layer adapts to specific user preferences.
[0011] Based on a layered lightweight reinforcement learning model, the federated learning principle is used to process local data, calculate the model gradient of the personalization layer and add differential privacy noise, and realize multi-device collaborative learning through a secure aggregation protocol;
[0012] Based on the local data processing process, a data value assessment algorithm is constructed to dynamically allocate privacy budgets and adjust the privacy protection level;
[0013] Based on the dynamic allocation of privacy process, cryptographic proof is generated for the calculation process using user data, enabling users to verify that data usage is in compliance with the authorized scope, and based on the output of the sharing layer and the personalization layer combined with user status and environmental information, the control parameters of the smart light strip are generated.
[0014] Furthermore, the steps of constructing a lightweight reinforcement learning model using knowledge distillation technology include:
[0015] Training a knowledge distillation network consisting of a teacher model and a student model. The teacher model is a complex reinforcement learning model, and the student model is a lightweight network structure.
[0016] The output distribution generated by the teacher model is used to guide the training of the student model, and the student model is optimized through a combination of task loss and knowledge distillation loss;
[0017] Perform computational graph optimization on the student model, including weight quantization, computational graph fusion, and redundant operation elimination.
[0018] Furthermore, the organization of the shared layer and the personalized layer includes:
[0019] The shared layer stores general lighting knowledge, is responsible for extracting general lighting features, sharing them among multiple users, and updating them regularly;
[0020] The personalization layer stores user-specific preference parameters, which are stored and trained only on the local device and optimized for specific users;
[0021] When a specific user is detected, the system automatically switches to the corresponding personalized layer and generates personalized control parameters in combination with the shared layer output.
[0022] Furthermore, the step of processing local data using the federated learning principle includes:
[0023] Use user data on the local device to train the model and calculate the model parameter gradients;
[0024] Add differential privacy noise to the calculated gradient to protect user privacy;
[0025] A secure aggregation protocol is used to calculate the global gradient mean without leaking the gradients of individual devices;
[0026] The shared layer model is updated using the global gradient, while the personalized layer is updated only with local data.
[0027] Furthermore, the steps of dynamically allocating the privacy budget and adjusting the privacy protection level include:
[0028] Build data value assessment algorithms to quantify the sensitivity and utility of different types of data;
[0029] Dynamically allocate privacy budgets for different types of data based on data value and usage scenarios;
[0030] Provide a user interface that allows users to set the global privacy level and the protection strength for specific data types;
[0031] Automatically adapting the noise scale in differential privacy according to the allocated privacy budget.
[0032] Furthermore, the step of generating a cryptographic proof includes:
[0033] Define the constraint system of the calculation, which represents the calculation process and its constraints;
[0034] Generate proof key and verification key;
[0035] Generate a proof using the proof key and input data, proving that the computation was performed correctly and complies with predefined constraints.
[0036] Provides a user control interface that allows users to set the scope of authorization for data usage, including the allowed computing types, data access frequency, and persistence.
[0037] Furthermore, the lightweight reinforcement learning model is a deep Q network, the input of which includes ambient light, time, user historical preferences and activity type status information, and the output includes the brightness, color temperature and color control parameters of the light strip.
[0038] Furthermore, the sharing among multiple users balances the needs of different users through the following steps:
[0039] The number of users in the testing environment;
[0040] Get the personalized layer model for each user;
[0041] Calculate the weight of each user's preference based on the number of users, location, and activity type;
[0042] Based on weighted combination, control parameters are generated that take into account the needs of each user.
[0043] Furthermore, the personalization layer training adopts a meta-learning method to quickly adapt to new user preferences, and adopts a progressive learning strategy to gradually refine the personalization model as user data accumulates.
[0044] The present invention provides a reinforcement learning-based intelligent light strip dynamic control system, which is used to execute the above-mentioned reinforcement learning-based intelligent light strip dynamic control method, including:
[0045] Model compression module, used to achieve knowledge distillation and model lightweight processing;
[0046] Layered architecture module, used to manage the structural organization of shared and personalized layers;
[0047] Federated learning module for performing local training, differential privacy protection, and secure aggregation;
[0048] Privacy management module, used for data value assessment and dynamic allocation of privacy budget;
[0049] Verification control module, used to generate calculation proof and output control of light strip parameters.
[0050] The beneficial effects of the present invention are: through knowledge distillation and model optimization technology, the size of the reinforcement learning model is compressed, the memory usage is reduced, and the reasoning speed is improved;
[0051] By leveraging a hierarchical model structure and federated learning technology, we achieve the goal of providing high-quality personalized services while protecting user data privacy. User preference data is always retained on the local device, and unprocessed raw data is never uploaded or shared, fundamentally eliminating the risk of data leakage. Furthermore, the system service quality improves rather than decreases, increasing user satisfaction and reducing the need for manual intervention.
[0052] Through data value assessment and a multi-level privacy budget allocation mechanism, this method can dynamically adjust the privacy protection level according to data sensitivity and usage scenarios, achieving differentiated protection for different types of data. Sensitive data (such as user behavior patterns) receives stronger protection, while insensitive data (such as ambient light) uses a lower protection level to improve performance, reducing the risk of privacy leakage while maintaining personalized effects.
[0053] In a multi-user shared environment, this method can balance the lighting preferences of different users and create a coordinated and consistent lighting environment while protecting the privacy of each user, thus reducing user conflicts and improving the satisfaction of the shared space experience.
[0054] This method implements complex reinforcement learning algorithms locally on edge devices, reducing dependence on network connections, improving system response speed and stability, increasing system offline availability, and reducing user-perceived latency to a level imperceptible to the human eye. It can also operate normally even when the network is unstable or disconnected.
[0055] Through verifiable computing technology, users can verify that the use of their data is within the scope of authorization, which increases system transparency and accountability. This increases users' trust in the system and makes them willing to share more data to obtain better services, forming a virtuous cycle.
[0056] The optimized lightweight model is more energy-efficient when running on edge devices. Compared with cloud computing solutions, computing energy consumption is reduced, network transmission energy consumption is reduced, and lighting energy consumption is also reduced by intelligently adjusting light strip parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 This is a flow chart of a method for dynamic control of intelligent light strips based on reinforcement learning in the present invention;
[0058] Figure 2 This is a bar chart comparing the lightweight model in the method of the present invention with the traditional model in terms of model size, memory usage, and inference latency;
[0059] Figure 3 It is a line graph showing the changing trend of the number of manual interventions by different user types over the usage time after the deployment of the smart light strip system;
[0060] Figure 4 This is a radar chart comparing the solution of the present invention with traditional cloud computing solutions and simple local solutions in terms of five key performance indicators;
[0061] Figure 5 It is an area chart showing the changing trends of personalized accuracy, response time, and energy efficiency of the solution of the present invention under different privacy protection level settings;
[0062] Figure 6 It is a pie chart showing the energy consumption distribution during the operation of the method of the present invention. DETAILED DESCRIPTION
[0063] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.
[0064] At least one embodiment of the present invention discloses a method for dynamic control of intelligent light strips based on reinforcement learning, such as Figure 1 Shown, including:
[0065] Step 1: Build a lightweight reinforcement learning model using knowledge distillation technology to compress the complex teacher model into a student model that can run on edge devices;
[0066] According to an embodiment of the present application, this step uses knowledge distillation technology to compress a complex reinforcement learning model into a lightweight model that can run on edge devices, including the following sub-steps:
[0067] Step 1.1, teacher model training;
[0068] A fully functional reinforcement learning teacher model is trained as the knowledge source. This teacher model uses a Deep Q-Network (DQN) architecture. Its input includes state information such as ambient light, user preferences, and time. Its output is the light strip control action (brightness, color temperature, color, and other parameters). The reward signal is obtained through interaction with the environment to optimize the strategy.
[0069] The teacher model is trained by minimizing the temporal difference error. It should be understood that this training process utilizes standard DQN techniques, such as experience replay and target networks, to improve training stability and efficiency. It should be noted that teacher model training can be performed on high-performance servers without resource constraints.
[0070] Step 1.2, student model construction;
[0071] Build a student model with reduced parameters, using a lightweight network architecture, such as reducing the number of layers, reducing the number of neurons in each layer, and using separable convolutions. The input and output of the student model remain consistent with the teacher model, but the internal structure is more streamlined.
[0072] The student model provided in this application adopts a simplified network structure, including a 3-layer fully connected network, the number of input layer neurons is the state vector dimension (usually including features such as ambient light, time, user historical preferences, etc., with a dimension of about 20 to 30), the number of hidden layer neurons is 64 and 32, and the number of output layer neurons is the action space dimension (usually including control parameters such as brightness, color temperature, and color, with a dimension of about 5 to 10). The activation function uses ReLU to improve computational efficiency. In contrast, the teacher model adopts a more complex structure, usually containing 5 to 6 layers of neural networks, with more neurons in each layer (such as 128, 256, etc.).
[0073] Step 1.3, knowledge distillation implementation;
[0074] Knowledge distillation is used to transfer the teacher model's knowledge to the student model. The distillation process optimizes two objectives: the student model's direct learning loss from the environment (task loss), and the difference between the output distributions of the student and teacher models (KL divergence). These two components are weighted together using a balancing factor, enabling the student model to both learn from the environment and mimic the behavior of the teacher model.
[0075] L distill =α distill ·L task +(1-α distill )·D KL (π T ||π S );
[0076] Among them L distill is the knowledge distillation loss, α distill is the knowledge distillation balance factor, ranging from [0, 1], L task is the mission loss, D KL is the KL divergence (Kullback-Leibler divergence), π T and π S are the strategies (probability distributions) of the teacher model and the student model respectively.
[0077] In some embodiments, the student model may optionally employ a convolutional neural network structure, particularly when the input state contains spatially correlated features (e.g., room layout, multi-region lighting distribution, etc.). In this case, the student model may include one or two lightweight convolution layers, such as depthwise separable convolution instead of standard convolution, to reduce the number of parameters and computational complexity.
[0078] Optionally, the knowledge distillation process can employ temperature scaling techniques, introducing a temperature parameter to adjust the "softness" of soft labels. Furthermore, in other implementations, a progressive distillation strategy can be employed, dividing the knowledge distillation process into multiple stages, each focusing on a different aspect of the model. This approach is particularly effective on edge devices, where computing resources are particularly constrained.
[0079] In the smart light strip scenario, knowledge distillation is specifically applied as follows: A teacher model is first trained on a high-performance server using a large amount of simulated data and limited real-world user data. It learns how to adjust light strip parameters based on time, ambient light, user activity, and other state information. For example, it can automatically dim the lights when the user is watching a movie, provide an appropriate color temperature for reading, or gradually increase the brightness in the morning to wake the user. After training, the teacher model can accurately predict user lighting preferences in different scenarios.
[0080] The teacher model then generates a large number of input-output pairs, and the student model learns by imitating these outputs. As a result, the student model can acquire the knowledge refined by the teacher model with a smaller parameter size and achieve similar performance levels.
[0081] Step 1.4, progressive training and optimization;
[0082] A progressive training strategy is adopted, starting with simple tasks to train the student model and gradually increasing the complexity of the tasks. At the same time, computational graph optimization is performed, including weight quantization, computational graph fusion, and redundant operation elimination, to further reduce the model's computing resource requirements.
[0083] It's important to note that progressive training effectively avoids optimization difficulties faced by the student model when faced with complex tasks. By increasing the difficulty of the task incrementally, the student model can learn more smoothly. Furthermore, computational graph optimization is a hardware-specific optimization technique that fully leverages the computing characteristics of the target device to further improve operational efficiency.
[0084] Step 2: Divide the lightweight reinforcement learning model into a shared layer and a personalized layer. The shared layer extracts general lighting features, and the personalized layer adapts to specific user preferences.
[0085] The method provided in this application divides the reinforcement learning model into a shared layer and a personalized layer to achieve efficient distribution and privacy protection of the model. It includes the following sub-steps:
[0086] Step 2.1, model hierarchical division;
[0087] The lightweight reinforcement learning model is divided into two parts:
[0088] Shared layer: The underlying network extracts common lighting features and learns universal lighting patterns.
[0089] Personalization layer: The top-level network that adapts to specific user preferences and stores user personalized parameters.
[0090] According to the embodiment of the present application, the overall structure of the model adopts a serial approach, that is, the input state is first processed by the shared layer to obtain an intermediate representation, and then the intermediate representation is converted into the final control action by the personalized layer. Here, a simplified formula is retained:
[0091] f(s)=f personal (f shared (s));
[0092] Where f(s) represents the complete model output function, that is, the mapping from the input state to the final control action; f shared is a shared layer function that processes the model part of the general lighting features; f personalis the personalized layer function, which processes the model part of the user's specific preferences; s is the input state vector, which contains features such as ambient lighting, time, and user preferences.
[0093] In practice, the shared layer typically consists of a two-layer fully connected neural network, with raw state features as input and an intermediate representation vector as output. For example, the input layer receives user-independent environmental state information such as ambient light, time, and weather. This information is processed by two hidden layers, each containing 48 and 32 neurons, to output a 16-dimensional intermediate representation vector that encodes general lighting environment characteristics.
[0094] The personalization layer consists of a one- or two-layer fully connected network. Its input is the output of the shared layer (intermediate representation vector) and user-specific preference features (such as historical user interaction data and explicitly set preferences). Its output is the light strip control action. For example, the personalization layer receives the 16-dimensional vector output by the shared layer, combines it with the user's specific 10-dimensional preference features, and processes it through a hidden layer containing 24 neurons to ultimately output the action parameters that control the light strip.
[0095] It should be understood that this hierarchical structure enables the model to learn general rules while adapting to individual differences, achieving a balance between efficient reuse of model structure and personalized customization.
[0096] Step 2.2, shared layer training and distribution;
[0097] The shared layer is trained using anonymous data from multiple users to learn general lighting knowledge. Once trained, it is distributed to all end devices via a secure channel, serving as the foundation for the personalized layer. The shared layer is updated regularly, but much less frequently than the personalized layer.
[0098] It should be noted that the data used for shared layer training has been anonymized and does not contain identifiable personal information. In addition, shared layer updates use incremental learning to minimize the impact on end devices.
[0099] Step 2.3: local training of the personalization layer;
[0100] On the user's local device, the personalized layer is trained using the user's own data, based on a fixed shared layer. This personalized layer training utilizes a small-sample learning approach to quickly adapt to user preferences. The personalized layer training process is entirely local, and the original user data never leaves the device, ensuring data privacy.
[0101] Furthermore, the personalization layer training can optionally employ meta-learning methods, enabling the model to quickly adapt to new user preferences using a small number of samples (5 to 10). Furthermore, the local training process can employ a progressive learning strategy, gradually refining the personalization model as user data accumulates.
[0102] In the smart home scenario, an example of a hierarchical model architecture is as follows: When the system is deployed in a new home, the shared layer has already been pre-trained using anonymized data from a large number of homes, incorporating common lighting patterns, such as adjusting color temperature with changing sunlight and dynamically adjusting brightness based on ambient light intensity. New users only need to train the personalized layer to adapt to their specific preferences, such as warmer or cooler color temperatures or brighter or dimmer lighting. This significantly reduces the onboarding time for new users; typically, only three to five days of usage data is needed to form an accurate personalized model.
[0103] In a multi-user home, the system first detects the number of users in the environment and maintains a separate personalization layer for each user, while sharing the same shared layer. A personalization layer model is then generated for each user. Based on the number of users, location, and activity type, each user's preferences are weighted. Based on this weighted combination, control parameters are generated that take into account the needs of each user. When different users are detected (e.g., through phone location or wearable devices), the system automatically switches to the corresponding personalization layer, providing a customized lighting experience while maintaining consistency in the underlying lighting patterns.
[0104] Step 3: Based on a layered lightweight reinforcement learning model, the federated learning principle is used to process local data, calculate the model gradient of the personalization layer, add differential privacy noise, and implement multi-device collaborative learning through a secure aggregation protocol.
[0105] According to an embodiment of the present application, this step implements multi-device collaborative learning based on the principle of federated learning while protecting user privacy, and includes the following sub-steps:
[0106] Step 3.1, local data processing and training;
[0107] Each terminal device uses local data for model training and calculation of model gradients. During the training process, the original user data is always retained on the local device and is not uploaded or shared.
[0108] In some implementations, local training can employ an asynchronous training strategy, allowing devices to determine the training frequency based on their computing power and data volume. For example, devices with greater computing power or frequently updated data can perform local training more frequently, while devices with limited computing power can train less frequently, thereby balancing system load.
[0109] Optionally, local training can employ importance sampling techniques to select training samples based on the importance of the data (e.g., novelty, representativeness) rather than simply using all available data. This approach can improve training efficiency, especially when data distribution is uneven.
[0110] Step 3.2, differential privacy gradient processing;
[0111] Add noise to the locally calculated gradient to achieve differential privacy protection. The core formula retained here is:
[0112]
[0113] in represents the gradient after adding noise; g orig is the original gradient, The mean is 0 and the variance is Gaussian noise distribution; C is the gradient clipping threshold, used to limit the gradient norm; σ noise is the noise scale parameter, and the privacy budget ε budget Related.
[0114] Differential privacy is a mathematically rigorous privacy protection method that adds carefully calibrated noise to the data, ensuring that any analysis results do not reveal individual information. In this application, differential privacy protection makes it impossible to reverse-infer the user's original data from the gradient information, even during multi-device collaborative learning.
[0115] In some embodiments, Laplace noise can be used in addition to Gaussian noise. Furthermore, the system can optionally employ gradient compression and quantization techniques, first compressing the gradient (e.g., sparsifying, quantizing) and then adding noise, thereby reducing communication overhead and improving privacy protection efficiency.
[0116] Step 3.3, security aggregation implementation;
[0117] A secure aggregation protocol is used to calculate the global mean gradient without leaking individual device gradients. The aggregation process involves taking a weighted average of the noisy gradients from multiple devices to determine the global update direction.
[0118] It should be noted that encrypted communication is used during the secure aggregation process to ensure the security of the transmission process. In addition, the aggregation process only focuses on the statistical characteristics of the gradient rather than specific individual data, further enhancing privacy protection.
[0119] In some implementations, secure aggregation can employ homomorphic encryption techniques, allowing operations to be performed directly on encrypted gradients, further enhancing privacy protection. For example, gradients can be encrypted using a semi-homomorphic encryption scheme (such as Paillier encryption), then aggregated in an encrypted state, with only the aggregated result decrypted.
[0120] Optionally, the system can implement a hierarchical aggregation strategy, first performing local aggregation between geographically close devices, and then globally aggregating the local aggregation results. This approach can reduce communication latency and improve system scalability, making it particularly suitable for large-scale deployment scenarios.
[0121] Therefore, through local training, differential privacy protection, and secure aggregation, this application achieves the goal of improving model performance by utilizing multi-device data while protecting user data privacy.
[0122] Step 4: Based on the local data processing process, a data value assessment algorithm is constructed to dynamically allocate privacy budget and adjust the privacy protection level;
[0123] According to an embodiment of the present application, this step implements dynamic adjustment of the privacy protection level based on data sensitivity and usage scenarios, and includes the following sub-steps:
[0124] Step 4.1, data value assessment;
[0125] Build a data value assessment algorithm to quantify the sensitivity and utility of different types of data. This algorithm considers two key dimensions of data: sensitivity (the privacy risk that data leakage may cause) and utility (the contribution of data to model performance).
[0126] In practice, the data value assessment algorithm uses a weighted sum approach, balancing sensitivity and utility using weight coefficients. For example, user activity patterns and habit data have a higher sensitivity score (e.g., 0.8 to 0.9), while ambient lighting data has a lower sensitivity score (e.g., 0.2 to 0.3). Utility is assessed by calculating the data's impact on model performance. This can be measured using an influence function or model perturbation method, mapping the data's contribution to improving model accuracy to the [0, 1] range.
[0127] In some implementations, data value assessment can utilize a multi-factor model that considers multiple data dimensions, such as sensitivity, utility, timeliness, and scarcity, and uses a weighted combination to derive a final data value score. This multi-factor model enables a more comprehensive assessment of data value and adapts to the needs of diverse application scenarios.
[0128] Optionally, the system can adopt an adaptive weight adjustment mechanism to dynamically adjust the weight of each factor based on user feedback and system performance, so that the evaluation results are more in line with actual needs.
[0129] It should be understood that data value assessment is the basis of dynamic privacy protection. By accurately assessing the value of different data, the system can allocate privacy resources in a targeted manner and achieve refined privacy protection.
[0130] Step 4.2, multi-level privacy budget allocation;
[0131] Dynamically allocate a privacy budget based on data value and usage scenarios. The privacy budget is a core parameter in differential privacy theory that controls the amount of noise added to the data. Smaller values result in higher privacy protection but lower data utility; larger values result in lower privacy protection but higher data utility.
[0132] In this application, privacy budget allocation adopts a multi-tiered strategy, assigning different privacy budgets to data of different values. Specifically, the system first sets a base privacy budget and then adjusts it based on the data value score and scenario factors. For high-value (high sensitivity, low utility) data, a smaller privacy budget is allocated, and more noise is added to protect the data; for low-value (low sensitivity, high utility) data, a larger privacy budget is allocated, and less noise is added to maintain data validity.
[0133] In some implementations, privacy budget allocation can employ an adaptive allocation strategy based on user historical behavior, taking into account historical user behavior statistics related to the data item, such as past data usage frequency and user interest in that data type. This approach can dynamically adjust the privacy protection level based on the user's actual usage patterns, making the system more personalized.
[0134] Optionally, the system can implement multi-level privacy budget allocation, allocating privacy budgets for data of different granularities (such as single records, data sets, user profiles), and ensuring that the overall privacy protection level meets the requirements through a combination mechanism.
[0135] In the smart light strip application scenario, the specific application example of multi-level privacy budget allocation is as follows: when the system operates in a private home scene (such as a bedroom), for the user's sleep pattern data, the system allocates a smaller privacy budget (such as ε i = 0.5), adding a larger noise protection data; for medium-value user lighting adjustment history, the system allocates a medium privacy budget (such as ε i =1.0); for ambient lighting data, the system allocates a larger privacy budget (such as ε i =2.0), adding small noise to maintain data validity.
[0136] Step 4.3, user privacy preference configuration;
[0137] Provides a user interface that allows users to set privacy protection preferences, including the global privacy level and the protection strength for specific data types. The system adjusts the privacy budget allocation strategy based on user configuration to ensure that user privacy needs are met.
[0138] It's important to note that user privacy preferences are presented in an easy-to-understand format, such as "high / medium / low" privacy protection levels, rather than directly exposing technical parameters. Furthermore, the system provides visual feedback on privacy-functionality tradeoffs to help users understand the impact of different privacy settings on system functionality.
[0139] Step 4.4, adaptive noise scale adjustment;
[0140] Automatically adjust the noise scale in differential privacy based on the allocated privacy budget. The noise scale is a direct parameter that controls the amount of noise added to the data and is inversely proportional to the privacy budget.
[0141] This application uses an adaptive noise adjustment algorithm to calculate the optimal noise scale based on parameters such as privacy budget and failure probability. In addition, the system also considers the distribution characteristics of the data and adopts different noise patterns for different types of data to improve the effectiveness of privacy protection.
[0142] Therefore, through data value assessment, multi-level privacy budget allocation, user privacy preference configuration and adaptive noise scale adjustment, this application achieves dynamic adjustment of privacy protection strength and strikes a balance between protecting user privacy and providing high-quality services.
[0143] Step 5: Based on the dynamic privacy allocation process, a cryptographic proof is generated for the computational process using user data, enabling the user to verify that the data usage complies with the authorized scope. The control parameters of the smart light strip are generated based on the output of the sharing layer and the personalization layer, combined with the user status and environmental information.
[0144] This application introduces verifiable computing technology to ensure that users can verify that their data usage complies with the authorized scope, including the following sub-steps:
[0145] Step 5.1, calculation proof generation;
[0146] Generate a cryptographic proof for each computation using user data. This proof proves that the computation was performed correctly and that predefined constraints (such as data usage limits) were met without revealing the specific content of the original data.
[0147] In its implementation, computational proof generation is based on zero-knowledge proof (ZKP) technology, specifically zero-knowledge succinct non-interactive argument of knowledge (zk-SNARK). The system first represents the computational process as an arithmetic circuit or constraint system, then uses a zk-SNARK generator to generate a proving key and a verification key. When the computation is performed, the system generates a proof using the proving key and input data.
[0148] The main steps of proof generation include:
[0149] Define the constraint system of the calculation, which represents the calculation process and its constraints;
[0150] Generate a proof key and a verification key using system parameters;
[0151] Generate a proof using the proving key, public input, and private input.
[0152] It should be understood that zero-knowledge proof technology allows one party (the prover) to prove to another party (the verifier) that a statement is true, without revealing any information other than that the statement is true. In this application, this means that the system can prove that its use of user data complies with predefined rules without revealing the specific content of the data.
[0153] Step 5.2, user verification is performed;
[0154] Users can use the verification function to check the validity of the computational proof, confirming that the data is being used as expected. This process is highly efficient, as it only requires verifying the key, the digest of the computational result, and the proof. It doesn't require access to the original data or re-execution of the computation.
[0155] It’s important to note that the verification process can be performed locally on the user’s device, without sending data to an external server. Furthermore, the verification result is deterministic: for a given input and calculation process, the verification result is either valid or invalid, with no ambiguity.
[0156] Step 5.3, authorization scope management;
[0157] Provides a user control interface that allows users to set the scope of authorization for data usage, including the allowed calculation types, data access frequency and persistence, etc. The system constrains data usage based on user authorization settings.
[0158] Furthermore, this application enables fine-grained control over the scope of authorization, allowing users to set different authorization rules for different types of data. For example, a user could allow the system to use lighting preference data for model training, but prohibit the use of that data for user profiling or behavior analysis.
[0159] In the smart light strip application scenario, verifiable computing is used as follows: When the system needs to train a model using a user's lighting preference data, it first checks the authorization scope set by the user to ensure that the operation complies with the user's permission. The system then generates a computation proof, confirming that the computation performed was indeed within the authorization scope and that only the necessary data fields were used.
[0160] For example, a user might authorize the system to use their lighting preference data for model training, but prohibit the use of that data for user profiling or behavioral analysis. When the system trains the model, it generates a certificate indicating that the computations performed were limited to model training and no profiling-related computations were performed. This certificate can be verified by the user's device, ensuring that the system complies with the data usage rules.
[0161] Through this step, the application improves system transparency and user trust, allowing users to be confident that their data is only used for authorized purposes without affecting the quality of model optimization.
[0162] A reinforcement learning-based intelligent light strip dynamic control system, used to execute the above-mentioned reinforcement learning-based intelligent light strip dynamic control method, comprising:
[0163] Model compression module, used to achieve knowledge distillation and model lightweight processing;
[0164] Layered architecture module, used to manage the structural organization of shared and personalized layers;
[0165] Federated learning module for performing local training, differential privacy protection, and secure aggregation;
[0166] Privacy management module, used for data value assessment and dynamic allocation of privacy budget;
[0167] Verification control module, used to generate calculation proof and output control of light strip parameters.
[0168] Here, the present invention provides an implementation example:
[0169] The method of this embodiment has been tested in a smart home environment. The test environment is a residence of approximately 120 square meters, consisting of six areas: living room, dining room, master bedroom, second bedroom, study, and kitchen. Each area is equipped with smart light strips. The family members include a couple and a teenager, each with different lighting preferences and schedules.
[0170] The hardware devices used in this house include: a smart home central controller (equipped with a quad-core 1.5GHz processor and 2GB of RAM); environmental sensors in each room (light, temperature, and motion detection); smart light strips that support adjustable brightness, color temperature, and RGB color; and the user's smartphone and wearable device (for user identification and preference collection).
[0171] There are differences in lighting preferences among family members: adult males prefer brighter, cool-toned light (about 5500K color temperature) for reading and working; adult females prefer warm and soft light (about 3800K color temperature); and teenagers have varying lighting preferences for different activities, including colorful dynamic lighting effects for gaming and moderately bright, neutral light for studying.
[0172] In addition, family members are also concerned about privacy protection, especially for data that may reveal personal routines and behavioral patterns. They hope to protect their personal privacy while enjoying the convenience of smart lighting.
[0173] The specific implementation process of this embodiment in the above-mentioned smart home environment is as follows:
[0174] First, a complex teacher model was trained on a high-performance server. This model uses a five-layer deep Q-network structure, with each layer containing 128 to 256 neurons. The input consists of a 32-dimensional state vector (including information such as ambient light, time, user activity, and historical preferences), and the output is a 10-dimensional action vector (which controls parameters such as the brightness, color temperature, and RGB values of the light strip). The teacher model is trained on a dataset containing approximately 50,000 state-action-reward examples, including simulated data and limited real user data.
[0175] Subsequently, a lightweight student model was constructed, consisting of a three-layer neural network with layer sizes of 32, 64, and 32, and containing only 4.2% of the parameters of the teacher model. Through a knowledge distillation process using soft labels with a temperature of 2.5 and a balance factor of α = 0.7, the student model successfully acquired the core knowledge of the teacher model after 100 rounds of training.
[0176] Finally, the computational graph of the student model was optimized, including 8-bit weight quantization, computational graph fusion, and redundant operation elimination, further compressing the model size to 1.8% of the original teacher model and reducing memory usage to 4.3% of the original.
[0177] In actual deployment, the student model is divided into a shared layer and a personalized layer. The shared layer contains the first two layers of the network (32-64), which extract general lighting features; the personalized layer contains the last layer of the network (32-output layer), which adapts to specific user preferences.
[0178] The system maintains a separate personalized layer model for each of the three family members. When a specific user is detected (via smartphone location or wearable device identification), the system automatically switches to the corresponding personalized layer. For example, if a teenager is detected in the room, the system loads a personalized layer that matches their preferences, providing a lighting environment suitable for learning or entertainment.
[0179] The shared layer is trained on anonymized multi-family data and updated monthly, while the personalized layer is incrementally trained daily on the local device to ensure rapid adaptation to changes in user preferences. In this family, the personalized layer only needed five days of interaction data to accurately fit the preferences of a specific user.
[0180] The smart home controller in the home performs local data processing and training, and the user's raw behavior data and lighting preference data are always retained locally. The system participates in a federated learning process once a week, collaborating with hundreds of other homes to optimize the shared layer model.
[0181] During federated learning, locally computed gradients are first processed using differential privacy. For highly sensitive data (such as nighttime activity patterns), a smaller privacy budget (ε = 0.6) is used, with greater noise added. For less sensitive data (such as ambient lighting responses), a larger privacy budget (ε = 2.0) is used, with less noise added.
[0182] Furthermore, the system implements a hierarchical aggregation strategy, first performing local aggregation among 10 to 15 geographically close households, and then submitting the local aggregation results to the central server for global aggregation. This approach not only reduces communication latency but also provides an additional layer of privacy protection.
[0183] The system dynamically adjusts the privacy protection level based on the data type and usage scenario. For example, for data collected during the day in public areas (such as the living room), the system sets a lower privacy protection level (allowing more data sharing to improve performance); while for data collected at night in private areas (such as the bedroom), the system sets a higher privacy protection level.
[0184] Family members set their privacy preferences through a smartphone app, using simple "high / medium / low" privacy level options, which the system automatically converts into corresponding technical parameters. During the test, adult family members chose "medium" privacy protection, while teenagers chose "high" privacy protection.
[0185] The system maintains a separate privacy budget account for each family member and tracks its consumption in real time to ensure it stays within user-defined limits. When the privacy budget for a particular type of data is about to be exhausted, the system will suspend collection of that data or prompt the user to decide whether to adjust their privacy settings.
[0186] The system generates a cryptographic proof for each computation that uses user data. For example, when training the personalization layer using the user’s lighting preference data, the system generates a proof showing that:
[0187] Only authorized data fields were accessed;
[0188] The calculation is only used for model training purposes;
[0189] No copies of the original data are kept.
[0190] Users can verify the proof of computation at any time through a smartphone app, which displays the purpose, time, and scope of data usage, as well as the verification results. During the testing period, all computation proofs generated by the system were successfully verified, reinforcing family members' trust in the system.
[0191] The system also implements fine-grained control over the scope of authorization. For example, a young user can set the system to use only their lighting preference data to optimize their personal experience, prohibiting their data from being used for global model improvement or other purposes. The system strictly adheres to these authorization rules and proves its compliance through verifiable computation.
[0192] A three-month practical application test in the aforementioned home environment verified the technical effectiveness of this implementation. The following is detailed verification data for two of the most important technical effects:
[0193] The actual operation of the lightweight model of this embodiment on a home smart controller (quad-core 1.5GHz processor, 2GB RAM) is as follows:
[0194] The original teacher model is 243MB in size and occupies 485MB of memory during inference. A single inference takes an average of 126 milliseconds, making it impossible to run in real time on the target hardware.
[0195] After knowledge distillation and optimization, the model size is only 4.4MB (compression rate 98.2%), the runtime memory usage is 22.5MB (reduction 95.4%), and the average time for a single inference is only 3.7 milliseconds, which can run smoothly on the smart home controller.
[0196] In actual testing, when the system detects a user entering a room or a change in ambient light, it analyzes and makes decisions within 5 milliseconds and adjusts the light strip parameters within 10 milliseconds. The total user-perceived latency is below the threshold of human perception. Even when the network is disconnected, the system maintains full functionality, achieving an offline availability rate of 99.8%.
[0197] In a 30-day comparative test, the smart light strip system using this method reduced average response time by 93.6%, improved system reliability by 17.8%, and reduced energy consumption by 68.4% compared to a cloud computing-based control group. The advantages of the localized system were particularly evident during periods of network instability, with the user experience score exceeding the control group by 41.2%.
[0198] Analysis of usage data from test households over a three-month period validated this implementation's ability to provide high-quality personalized services while protecting user privacy:
[0199] The number of manual user interventions recorded by the system gradually decreased from an average of 8.7 times per person per day in the first week to 2.4 times in the fourth week, and only 1.8 times in the twelfth week, a reduction of 79.3%, indicating that the system successfully learned user preferences and provided a lighting environment that met expectations.
[0200] The system's prediction accuracy for user activity improved from 62.5% in the first week to 87.3% in the fourth week, and reached 93.6% in the twelfth week. Privacy-preserving measures also ensured the security of user data. The noise added by differential privacy techniques reduced the success rate of reconstructing original user data from shared gradients to less than 5.8%, well below the 15% threshold required by privacy protection standards.
[0201] In the user satisfaction survey, family members rated the system's personalized performance at 4.6 out of 5, and their trust in privacy protection measures at 4.4 out of 5, both higher than the control group's 3.2 and 2.8. Notably, the use of verifiable computing increased user trust in the system from 3.5 at the beginning of the test to 4.4 by the end, demonstrating that transparent data usage policies effectively enhance user trust.
[0202] In tests of multi-user shared scenarios (e.g., family members in the living room), the system was able to find a lighting solution that balanced the needs of multiple users 91.4% of the time, with a conflict rate of only 8.6%, lower than the 34.7% of traditional systems. This demonstrates that this implementation can effectively address the issue of balancing personalized needs in multi-user environments.
[0203] To sum up, the actual application test results show that this implementation method successfully achieves the technical goal of efficiently running reinforcement learning models on edge devices and providing high-quality personalized services while protecting user privacy, verifying the technical effect and practical value of this application.
[0204] like Figures 2 to 6 As shown, there are bar charts comparing the lightweight model in the method of the present invention with the traditional model in terms of model size, memory usage and inference delay; a line chart showing the changing trend of the number of manual interventions by different user types over usage time after the deployment of the smart light strip system; a radar chart comparing the solution of the present invention with the traditional cloud computing solution and the simple local solution in five key performance indicators; an area chart showing the changing trend of personalized accuracy, response time and energy efficiency of the solution of the present invention under different privacy protection level settings; and a pie chart showing the energy consumption distribution of the method of the present invention during operation.
[0205] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.
Claims
1. A dynamic control method for intelligent light strips based on reinforcement learning, characterized in that: include: Build a lightweight reinforcement learning model using knowledge distillation technology, compressing the complex teacher model into a student model that can run on edge devices; The lightweight reinforcement learning model is divided into a shared layer and a personalized layer. The shared layer extracts general lighting features, and the personalized layer adapts to specific user preferences. Based on a layered lightweight reinforcement learning model, the federated learning principle is used to process local data, calculate the model gradient of the personalization layer and add differential privacy noise, and realize multi-device collaborative learning through a secure aggregation protocol; Based on the local data processing process, a data value assessment algorithm is constructed to dynamically allocate privacy budgets and adjust the privacy protection level; Based on the dynamic allocation of privacy process, cryptographic proof is generated for the calculation process using user data, enabling users to verify that data usage is in compliance with the authorized scope, and based on the output of the sharing layer and the personalization layer combined with user status and environmental information, the control parameters of the smart light strip are generated.
2. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The steps of constructing a lightweight reinforcement learning model using knowledge distillation technology include: Training a knowledge distillation network consisting of a teacher model and a student model. The teacher model is a complex reinforcement learning model, and the student model is a lightweight network structure. The output distribution generated by the teacher model is used to guide the training of the student model, and the student model is optimized through a combination of task loss and knowledge distillation loss; Perform computational graph optimization on the student model, including weight quantization, computational graph fusion, and redundant operation elimination.
3. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The organization of the shared layer and the personalized layer includes: The shared layer stores general lighting knowledge, is responsible for extracting general lighting features, sharing them among multiple users, and updating them regularly; The personalization layer stores user-specific preference parameters, which are stored and trained only on the local device and optimized for specific users; When a specific user is detected, the system automatically switches to the corresponding personalized layer and generates personalized control parameters in combination with the shared layer output.
4. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The steps of processing local data using the federated learning principle include: Use user data on the local device to train the model and calculate the model parameter gradients; Add differential privacy noise to the calculated gradient to protect user privacy; A secure aggregation protocol is used to calculate the global gradient mean without leaking the gradients of individual devices; The shared layer model is updated using the global gradient, while the personalized layer is updated only with local data.
5. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The steps of dynamically allocating the privacy budget and adjusting the privacy protection level include: Build data value assessment algorithms to quantify the sensitivity and utility of different types of data; Dynamically allocate privacy budgets for different types of data based on data value and usage scenarios; Provide a user interface that allows users to set the global privacy level and the protection strength for specific data types; Automatically adapting the noise scale in differential privacy according to the allocated privacy budget.
6. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The steps of generating a cryptographic proof include: Define the constraint system of the calculation, which represents the calculation process and its constraints; Generate proof key and verification key; Generate a proof using the proof key and input data, proving that the computation was performed correctly and complies with predefined constraints. Provides a user control interface that allows users to set the scope of authorization for data usage, including the allowed computing types, data access frequency, and persistence.
7. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The lightweight reinforcement learning model is a deep Q network, whose input includes ambient light, time, user historical preferences and activity type status information, and whose output includes the brightness, color temperature and color control parameters of the light strip.
8. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 3, characterized in that: Sharing among multiple users balances the needs of different users through the following steps: The number of users in the testing environment; Get the personalized layer model for each user; Calculate the weight of each user's preference based on the number of users, location, and activity type; Based on weighted combination, control parameters are generated that take into account the needs of each user.
9. The method for dynamic control of intelligent light strips based on reinforcement learning according to claim 1, characterized in that: The personalization layer training adopts a meta-learning method to quickly adapt to new user preferences, and adopts a progressive learning strategy to gradually refine the personalization model as user data accumulates.
10. A dynamic control system for intelligent light strips based on reinforcement learning, characterized in that: A method for dynamically controlling a smart light strip based on reinforcement learning, for executing any one of claims 1 to 9, comprising: Model compression module, used to achieve knowledge distillation and model lightweight processing; Layered architecture module, used to manage the structural organization of shared and personalized layers; Federated learning module for performing local training, differential privacy protection, and secure aggregation; Privacy management module, used for data value assessment and dynamic allocation of privacy budget; Verification control module, used to generate calculation proof and output control of light strip parameters.
Citation Information
Cited By
Complex scene-oriented AI large model lightweight deployment method
CN120930709A