Power consumer demand response portrait optimization method, system, equipment and medium
By generating reinforcement learning models and using closed-loop feedback mechanisms to optimize power user profiles, the problems of insufficient real-time performance and dynamic adaptability of user profiles in existing technologies are solved, and real-time adjustment and multi-objective optimization of user profiles are realized.
Patent Information
- Application Number
- CN202510949247.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-14
AI Technical Summary
Existing methods for generating power user profiles lack real-time update capabilities, making it difficult to accurately capture users' real-time response characteristics. They also cannot dynamically adjust according to changes in the real-time environment, and their feedback mechanisms are simplistic, resulting in poor accuracy and timeliness in response prediction.
By acquiring data, performing preprocessing and classification, an initial user profile is generated. A reinforcement learning environment is defined using a Markov decision process, a reinforcement learning model is generated based on a deep Q-network, and the user profile feature vector is adjusted using a closed-loop feedback optimization mechanism to obtain the final user profile.
It enables real-time adjustment of user profiles, adapting to changes in user behavior and the external environment, meeting various demand response objectives, and improving the accuracy and timeliness of response prediction.
Smart Images

Figure CN120952072A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of smart grid and demand response management technology, specifically to methods, systems, equipment, and media for optimizing demand response profiles of power users. Background Technology
[0002] In smart grids and demand response management, power companies need to use user profiles to predict electricity demand and optimize resource scheduling to achieve load balancing and power system stability. Traditional user profile generation methods typically rely on static user attributes and historical electricity consumption data. These methods include user profiles generated based on historical data analysis and profile optimization techniques based on simple feedback adjustments. However, as the real-time nature and complexity of grid demand response continue to increase, methods relying solely on static profiles are insufficient to reflect users' immediate demand response capabilities in a timely manner.
[0003] Current demand response user profile generation primarily relies on machine learning and statistical analysis methods. Typical methods include static analysis based on historical electricity consumption data and user profile generation based on feedback optimization. These methods suffer from the following shortcomings: Strong static profiles and insufficient dynamic response capabilities: Traditional user profiles are mainly generated from historical data and lack real-time update capabilities. Therefore, they struggle to accurately capture users' real-time response characteristics in demand response scenarios, reducing the accuracy of response prediction. Poor adaptability and difficulty in keeping up with external environmental changes: Most existing profile optimization methods lack adaptive mechanisms and cannot dynamically adjust according to changes in the real-time environment, resulting in profiles that fail to reflect users' electricity demand and response capabilities at different times. Limited feedback mechanisms and optimization objectives: Although some methods introduce feedback mechanisms, they often focus only on a single indicator, such as load response volume, failing to comprehensively consider multi-dimensional optimization objectives such as response time and user satisfaction, making it difficult to achieve comprehensive optimization of demand response profiles. Delayed feedback updates and poor real-time performance: Existing methods typically update profiles based on offline feedback, unable to adjust strategies in real-time during demand response, resulting in poor timeliness of user profiles in demand response. Summary of the Invention
[0004] In view of the above-mentioned problems, the present invention is proposed.
[0005] Therefore, the technical problem solved by this invention is: how to acquire external data for demand response profiles through a power user demand response profile optimization method, generate profile features and feature vectors, generate a reinforcement learning model by defining a reinforcement learning environment and reward function, adjust the profile using the reinforcement learning model, and adjust the feature vectors based on the feedback from the reinforcement learning model, ultimately obtaining and outputting the final user profile. This effectively solves the problem of profiles needing to adjust accordingly to changes in user behavior and the external environment. Simultaneously, the profile optimization process can satisfy various demand responses and dynamically optimize the profile based on the actual effects of the demand responses, thereby ensuring that the user profile can adapt to constantly changing electricity prices and environmental conditions.
[0006] To address the aforementioned technical problems, this invention provides the following technical solution: a method for optimizing power user demand response profiles, comprising the following steps:
[0007] Data is acquired, preprocessed, and classified to generate an initial user profile. The initial user profile is then structured, and user profile feature vectors are obtained. The demand response problem is modeled as a Markov decision process. A reinforcement learning environment is defined based on the Markov decision process. A reinforcement learning model is generated based on a deep Q-network. Using the reinforcement learning model and the reinforcement learning environment, the user profile feature vectors are adjusted. The parameters of the initial user profile are updated using a closed-loop feedback optimization mechanism. The profile feature weights are adjusted using a profile feature adjustment mechanism. The final user profile is then obtained and output.
[0008] As a preferred embodiment of the power user demand response profile optimization method of the present invention, the following steps are included: acquiring data, performing preprocessing and classification, generating an initial user profile, structuring the initial user profile, and obtaining user profile feature vectors, including: using a clustering algorithm to divide users into different demand response groups, generating the initial user profile for the different demand response groups, and processing the initial user profile.
[0009] As a preferred embodiment of the power user demand response profiling optimization method described in this invention, the demand response problem is modeled as a Markov decision process, and a reinforcement learning environment is defined based on the Markov decision process. This includes: modeling the demand response problem as the Markov decision process, and defining the state, action, and reward function of the reinforcement learning environment. The beneficial effects of this preferred embodiment are that it provides a mathematical framework for the Markov decision process, ensuring the theoretical reliability of the reinforcement learning model, and making the decision process easy to understand and use through the explicit definition of state transition probabilities and reward functions.
[0010] As a preferred embodiment of the power user demand response profile optimization method of the present invention, the method includes: adjusting the profile features by utilizing the reinforcement learning model, including: defining the state, actions, and reward function of the reinforcement learning environment, including: calculating the current user profile features, demand response state, and external environment features using a state vector function to obtain the state of the reinforcement learning; defining the actions of the reinforcement learning environment by adjusting the response intensity, changing the incentive strategy, and controlling the response delay; and defining the reward function of the reinforcement learning environment by combining the actions of the learning environment with customer satisfaction.
[0011] As a preferred embodiment of the power user demand response profile optimization method described in this invention, the generation of a reinforcement learning model based on a deep Q-network includes: training the discrete space using the deep Q-network, iteratively training using the Q-value update formula, and generating the reinforcement learning model. The beneficial effects of this preferred embodiment are that the discrete action space quantifies continuous adjustment parameters into finite discrete options, reducing the search space, avoiding the gradient calculation and optimization difficulties in the continuous action space, reducing the randomness of action selection, and improving the stability of the learning algorithm.
[0012] As a preferred embodiment of the power user demand response profile optimization method of the present invention, the method includes: using the reinforcement learning model, adjusting the user profile feature vector based on the reinforcement learning environment, and updating the parameters of the initial user profile using a closed-loop feedback optimization mechanism. This includes: adjusting the initial user profile according to the action definition in the reinforcement learning environment using a profile feature vector update formula; collecting demand response feedback through the closed-loop feedback optimization mechanism; calculating the feedback loss to determine the convergence state of the reinforcement learning model and the initial profile; and then adjusting the reinforcement learning model parameters using the feedback formula. The beneficial effects of this preferred embodiment are that the closed-loop feedback mechanism enables the system to continuously learn based on actual demand response effects, avoids the limitations of static profile models, achieves dynamic updates of user profiles, and, through actual effect feedback, allows the system to automatically identify and correct deviations in profile features, improving its autonomous operation capability.
[0013] As a preferred embodiment of the power user demand response profile optimization method of the present invention, the following is included: adjusting the profile feature weights using a profile feature adjustment mechanism to obtain and output the final user profile, comprising: adjusting the feature weights in the initial user profile based on the performance of the reinforcement learning model in the demand response using a feature weighting matrix.
[0014] This invention provides an optimization system for profiling and responding to electricity user demands.
[0015] To address the aforementioned technical problems, this invention further provides the following technical solution: a power user demand response profile optimization system, comprising: a data collection and profile construction module, which acquires data, performs preprocessing and classification, generates an initial user profile, structures the initial user profile, and obtains user profile feature vectors; a reinforcement learning environment definition module, which models the demand response problem as a Markov decision process and defines a reinforcement learning environment based on the Markov decision process; a reinforcement learning modeling module, which generates a reinforcement learning model based on a deep Q-network; a profile optimization module, which, in reinforcement learning, adjusts the initial user profile based on the user profile features according to the reinforcement learning environment and updates the parameters of the initial user profile using a closed-loop feedback optimization mechanism; and an information feedback and profile output module, which adjusts the profile feature weights using a profile feature adjustment mechanism, obtains the final user profile, and outputs it.
[0016] The present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the power user demand response profile optimization method.
[0017] The present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the power user demand response profile optimization method.
[0018] The beneficial effects of this invention are as follows: Through the online learning mechanism of the reinforcement learning model, user profiles are adjusted in real time, ensuring that the profiles can be adjusted accordingly to changes in user behavior and the external environment; by combining a multi-objective reward function that considers load response, response time, and user satisfaction, the profile optimization process can simultaneously meet multiple demand response objectives; through the closed-loop feedback mechanism of reinforcement learning, the profiles are dynamically optimized based on the actual effects of demand response, enabling profile adjustments to reflect user behavior in real time; and by utilizing the strong adaptability of the reinforcement learning model in different response environments, the user profiles can adapt to constantly changing electricity prices and environmental conditions. Attached Figure Description
[0019] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 The above is a flowchart of the power user demand response profile optimization method provided in one embodiment of the present invention. Detailed Implementation
[0021] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0022] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a method for optimizing power user demand response profiles, including:
[0023] S100: Acquire data, perform preprocessing and classification, generate initial user profiles, structure the initial user profiles, and obtain user profile feature vectors.
[0024] S200: Model the demand response problem as a Markov decision process, and define a reinforcement learning environment based on the Markov decision process.
[0025] S300: Generates reinforcement learning models based on deep Q-networks.
[0026] S400: Utilizes a reinforcement learning model and, based on the reinforcement learning environment, adjusts the user profile feature vector and updates the parameters of the initial user profile using a closed-loop feedback optimization mechanism.
[0027] S500: Use the profile feature adjustment mechanism to adjust the profile feature weights, obtain the final user profile, and output it.
[0028] It should be noted that traditional user profiles are mainly generated based on historical data and lack the ability to be updated in real time. Therefore, they are difficult to accurately capture the real-time response characteristics of users in demand response scenarios, which reduces the accuracy of response prediction. Most existing profile optimization methods lack adaptive mechanisms and cannot be dynamically adjusted according to changes in the real-time environment. At the same time, although traditional methods introduce feedback mechanisms, they often only focus on a single indicator, such as load response volume, and fail to comprehensively consider the multi-dimensional optimization goals of response time and user satisfaction, making it difficult to achieve comprehensive optimization of demand response profiles. In addition, existing methods usually update profiles based on offline feedback, which cannot adjust strategies in real time during the demand response process, resulting in poor timeliness of user profiles in demand response.
[0029] Therefore, to address the aforementioned issues, the response profile is optimized through steps S100-S500. First, electricity consumption data, user attribute data, and external environment data are acquired and preprocessed to obtain the basic input data for the demand response profile. Second, the basic input data is classified to generate profile features and feature vectors. Next, a reinforcement learning model is generated by defining a reinforcement learning environment and combining it with a reinforcement learning task. Then, the profile features are adjusted using the reinforcement learning model, and the feature vectors are adjusted based on the feedback from the reinforcement learning model to obtain the final user profile. Finally, the final user profile is output, providing data support for the adjustment strategies in the demand response profile. This effectively solves the problem of the profile being able to adjust accordingly to changes in user behavior and the external environment. Furthermore, the profile optimization process can satisfy various demand responses and dynamically optimize the profile based on the actual effects of the demand responses, thereby ensuring that the user profile can adapt to constantly changing electricity prices and environmental conditions.
[0030] Example 2, refer to Figure 1 This is the second embodiment of the present invention, which provides a method for optimizing the profile of power user demand response.
[0031] In this embodiment of the invention, in step S100, data is acquired, preprocessed and classified to generate an initial user profile, the initial user profile is structured, and user profile feature vectors are obtained. This includes acquiring electricity consumption data, user attribute data and external environment data, cleaning, standardizing and time-aligning the data, using the K-means clustering algorithm to divide users into different groups, generating an initial user profile, and then obtaining user profile feature vectors by structuring the initial user profile.
[0032] Specifically, the forms of data standardization are as follows:
[0033]
[0034] Where X represents the original data, μ is the mean, and σ is the standard deviation. norm This is the standardized data.
[0035] It should be noted that time alignment involves aligning data collected from different data sources according to a fixed time window and filling in missing values using interpolation methods to ensure the consistency of data time sequence.
[0036] Specifically, the structured initial portrait feature vector is represented as follows:
[0037] X = [x1, x2, ..., x n ];
[0038] Where X is the user's profile feature vector, containing information related to load characteristics, response potential, and device attributes, x1, x2, x3, x4, x5, x6, x7, x8, x9, x1, x1, x2, x9, x1, x1, x1, x2, x3 ...n Specific parameters in the preset load characteristics, user attribute characteristics, and response potential.
[0039] For example, the predefined load characteristics can be average daily load, peak-to-valley difference, and seasonal load fluctuations; user attribute characteristics can be user type, geographical location, and equipment type; and response potential can be user's historical demand response participation and reduction capabilities.
[0040] In one implementation, a clustering algorithm is used in S200 to divide users into different demand response groups. Alternatively, a decision tree classification algorithm can be used for classification. A decision tree model is constructed based on user attribute data, and the classification is performed step-by-step according to set age and electricity consumption through the decision nodes of the tree structure, making the classification results clear and interpretable.
[0041] In another possible implementation, clustering algorithms are used in S200 to classify users into different demand response groups. Alternatively, support vector machine (SVM) classification can be employed. In generating electricity demand response user profiles, the SVM classification method constructs multi-dimensional feature vectors of age and electricity consumption, and uses historical demand response data to train the model to find the optimal classification boundary. This method can classify users into high, medium, and low response groups, accurately scheduling power grid allocation.
[0042] In this embodiment of the invention, step S200 models the demand response problem as a Markov decision process and defines a reinforcement learning environment based on the Markov decision process, including modeling the demand response problem as a Markov decision process and defining the state, action, and reward function of the reinforcement learning environment.
[0043] It should be noted that Markov decision processes define states, actions, and reward functions, providing a mathematical framework.
[0044] Specifically, the specific manifestation of a state, as defined by a Markov decision process, is as follows:
[0045] S t =[X,DR t E t ];
[0046] Wherein, state S t Describe the current user profile feature X and the demand response status DR. t and external environmental characteristics E t .
[0047] It should be noted that the Demand Response Status (DR) t This can be based on load reduction, response time, and external environmental characteristics E. t This formula can be applied to factors such as electricity prices and weather. It ensures that the current user profile features do not depend on historical features, providing a mathematical basis for subsequent decision optimization.
[0048] It should be noted that defining actions through Markov decision processes includes adjusting response intensity, increasing or decreasing load reduction, and modifying user response intensity; adjusting incentive strategies, modifying user response incentives, and stimulating user participation; and controlling response latency, selecting appropriate response times, and ensuring timely demand response.
[0049] Specifically, the reward function defined through a Markov decision process takes the following form:
[0050] R = w1·L + w2·T + w3·U;
[0051] Where R is the total reward value, L is the load response amount, T is the response time reward, U is the user satisfaction, and w1, w2, and w3 are weighting coefficients.
[0052] It should be noted that the reward function needs to comprehensively consider load response, response time, and customer satisfaction, and provide corresponding power distribution strategies for different users by adjusting the weighting coefficients.
[0053] In one implementation, the Markov decision process can also be implemented using particle swarm optimization (PSO). PSO transforms the portrait optimization problem into an optimization process in a multi-dimensional space by simulating the foraging behavior of bird flocks. Each particle represents a portrait configuration scheme, and the search is guided by individual and group experience. For example, each particle in the particle swarm represents a set of portrait feature parameter configurations, and the feature parameters are adjusted based on individual and global optimal experiences to achieve iterative optimization and convergence of portrait features.
[0054] In another possible implementation, the reinforcement learning model can also be computed using a genetic algorithm. The genetic algorithm optimizes user profile features by simulating the natural selection process, encoding user profile parameters as chromosomes, and searching for the optimal solution through selection, crossover, and mutation operations. For example, the weights of each dimension of the user profile features can be encoded as real-number chromosomes. A multi-objective function integrating load response, response time, and user satisfaction can be used, combining excellent features from different user groups through crossover operations, and exploring new feature combinations through mutation operations.
[0055] In this embodiment of the invention, step S300, which generates a reinforcement learning model based on a deep Q-network, includes: training a discrete space using a deep Q-network, iteratively training using a Q-value update formula, and generating a reinforcement learning model.
[0056] Specifically, the Q-value update formula is expressed as follows:
[0057] Q(S t A t )=Q(S t A t)+α·[R+γ·maxQ(S t+1 ,A′)-Q(S t A t )];
[0058] Among them, Q(S) t A t This indicates that at time step t, the system is in state S. t Take action A t Then, the expected cumulative reward, where α is the learning rate, γ is the discount factor, R is the immediate reward, and maxQ(S) is the maximum cumulative reward. t+1 A') indicates that at the next time step (t+1), the system is in a new state S. t+1 At that time, the maximum Q value that can be obtained among all possible actions A′ is the maximum expected reward brought by the optimal future action.
[0059] It should be noted that Deep Q-Network takes user profile features, demand response status, and external environment information as input states. It approximates the Q-value function through a deep neural network and outputs expected cumulative rewards such as adjusting load reduction and changing incentive strategies. Through continuous interaction with the environment and parameter updates, DQN learns strategies that can maximize the multi-objective reward function.
[0060] In this embodiment of the invention, step S400 utilizes a reinforcement learning model to adjust the user profile feature vector based on the reinforcement learning environment, and uses a closed-loop feedback optimization mechanism to update the parameters of the initial user profile. This includes: adjusting the initial user profile according to the action definition in the reinforcement learning environment using the profile feature vector update formula; collecting feedback on demand response through the closed-loop feedback optimization mechanism; calculating the feedback loss to determine the convergence state of the reinforcement learning model and the initial profile; and then adjusting the reinforcement learning model parameters through the feedback formula.
[0061] Specifically, in reinforcement learning, user profile features are dynamically adjusted based on the selected action, which manifests as follows:
[0062] X new =X old +β·ΔX;
[0063] Among them, X new X represents the updated user profile feature vector. old Let β represent the user profile feature vector before the update, β be the update rate, and ΔX be the feature adjustment amount based on policy optimization.
[0064] Instantiated, reinforcement learning action definitions are optimization measures taken for the current state, such as adjusting load response strategies and optimizing response time.
[0065] In step S400, feedback on the demand response is collected through the closed-loop feedback optimization mechanism, the feedback loss is calculated to determine the convergence state of the reinforcement learning model and the initial profile, and then the parameters of the reinforcement learning model are adjusted through the feedback formula.
[0066] Specifically, the specific form of feedback loss calculation is as follows:
[0067]
[0068] Where L is the loss function, which measures the error between the Q-value output by the current reinforcement learning model and the target Q-value. The smaller the loss function, the closer the model's prediction is to the ideal target; This represents the summation of all time steps in the training sequence from time step t to T. 1 / 2 is a constant coefficient before the loss function, often used for standardizing the squared difference loss to simplify subsequent gradient calculations; Q(S t A t ) indicates that the current model at time step t, for state S t And Action A t The predicted Q value, i.e., the expected cumulative reward under the current strategy; R t This represents the actual immediate reward at the current time step, reflected in state S. t Take action A t The overall feedback obtained; γ is a discount factor that determines the degree to which future rewards affect the current target Q value. The larger the value, the more the model focuses on long-term returns; maxQ(S t+1 [A′)] represents the maximum Q value among all possible actions A′ at the next time step (t+1), which is the expected reward under the optimal decision in the future.
[0069] Specifically, the feedback adjustment formula is expressed as follows:
[0070]
[0071] Where, θ new This represents the updated model parameters, i.e., the latest parameter values of the reinforcement learning model after this feedback adjustment; θ old This represents the model parameters before the update, i.e., the parameters of the reinforcement learning model before the optimization iteration; ∈ represents the feedback adjustment learning rate, used to adjust the magnitude of each parameter update. The numerical range is usually 0-1. The larger the learning rate, the faster the parameter adjustment; the smaller the learning rate, the smoother the parameter convergence. This represents the gradient of the loss function L with respect to the parameter θ, which is the rate at which the loss function changes in the direction of the current model parameters. It is used to indicate how the parameters should be adjusted to reduce the value of the loss function, thereby optimizing the model performance.
[0072] It should be noted that the policy parameters θ in reinforcement learning include the Q-value function and the parameters of the policy network. During optimization, the policy parameters are iteratively updated using an adaptive learning rate to ensure that the policy converges to the optimal solution. Specifically, this is manifested as follows:
[0073]
[0074] Among them, W π b is the weight of the policy network. π Let α be the bias, α be the learning rate, L be the loss function, and W be the weight. π,new For the updated network weight parameters, W π,old b represents the network weight parameters before the update. π,new For the updated policy network bias parameters, b π,old These are the policy network bias parameters before the update.
[0075] In one implementation, the closed-loop feedback mechanism can also adjust user profile features through online learning. The online learning algorithm can process streaming data in real time and gradually update model parameters. In this invention, online learning can immediately update user profile features whenever new demand response data arrives, without waiting for a complete feedback cycle, ensuring the system's real-time responsiveness.
[0076] In another possible implementation, the reinforcement learning model feedback data can also be obtained through an expert system. The expert system builds an expert knowledge base and inference rules, updates the profile features according to predefined business rules and expert experience, and balances the power grid dispatching needs and user experience by allocating weights and incentive sensitivity when a user fails to meet the set electricity consumption standard within a month.
[0077] In this embodiment of the invention, step S500 uses a profile feature adjustment mechanism to adjust the profile feature weights, obtain the final user profile and output it, including adjusting the feature weights in the initial user profile based on the performance of the reinforcement learning model in demand response and using a feature weighting matrix.
[0078] In optimizing user profile features, to ensure that user profiles reflect users' true needs and response behaviors, the reinforcement learning model re-optimizes and adjusts the profile features after each cycle. By aggregating newly collected data features and combining them with real-time feedback response results, the user profile features are updated to ensure that the user profile can dynamically adapt to the constantly changing external environment. By utilizing the profile feature adjustment mechanism, based on the reinforcement learning model's performance in demand response, the weights of various features in the profile are adjusted. For example, for users with a more proactive load response, the weight of their response potential feature is increased to enhance prediction accuracy. The specific form of updating the feature weights in each iteration is as follows:
[0079] Wf,new =W f,old +β·ΔW;
[0080] Where β is the learning rate, ΔW represents the incremental update of the feature weights in each round of training by the reinforcement learning model, and W f W is the feature weighting matrix defined in this invention. f,new For the updated feature weighting matrix, W f,old This is the feature weighting matrix before the update.
[0081] It should be noted that the optimization of the profile features in step S500 can be fed back to the policy parameters of reinforcement learning in step S400 for coordinated optimization. Changes in the profile will affect the user's demand response behavior. Therefore, the policy parameters and profile features need to be optimized in conjunction. After each update of the reinforcement learning model, the profile is re-optimized according to the new policy. During the optimization process, the policy parameters are continuously updated and fed back to the reinforcement learning model. The optimization iteration and convergence judgment are judged by the iterative convergence condition. When the reinforcement learning model has reached the convergence state in multiple iterations, the iterative convergence condition is met.
[0082] In summary, this invention utilizes the online learning mechanism of a reinforcement learning model to adjust user profiles in real time, ensuring that the profiles adapt to changes in user behavior and the external environment. By combining a multi-objective reward function that considers load response, response time, and user satisfaction, the profile optimization process can simultaneously satisfy multiple demand response objectives. Through a closed-loop feedback mechanism of reinforcement learning, the profiles are dynamically optimized based on the actual effects of demand response, allowing profile adjustments to reflect user behavior in real time. Furthermore, the strong adaptability of reinforcement learning models in different response environments enables user profiles to adapt to constantly changing electricity prices and environmental conditions.
[0083] Example 3, referring to Figure 1 This is the third embodiment of the present invention. This embodiment provides a power user demand response profile optimization system, including a data collection and profile construction module, which acquires data, performs preprocessing and classification, generates an initial user profile, structures the initial user profile, and obtains user profile feature vectors; a reinforcement learning environment definition module, which models the demand response problem as a Markov decision process and defines a reinforcement learning environment based on the Markov decision process; a reinforcement learning modeling module, which generates a reinforcement learning model based on a deep Q-network; a profile optimization module, which adjusts the user profile feature vectors based on the reinforcement learning model and the reinforcement learning environment, and updates the parameters of the initial user profile using a closed-loop feedback optimization mechanism; and an information feedback and profile output module, which adjusts the profile feature weights using a profile feature adjustment mechanism, obtains the final user profile, and outputs it.
[0084] Example 4, the fourth embodiment of the present invention, differs from the previous three embodiments in that: if the function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0085] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0086] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.
[0087] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0088] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for optimizing the demand response profile of electricity users, characterized by: include, Data is acquired, preprocessed, and classified to generate an initial user profile. The initial user profile is then structured to obtain the user profile feature vector. The demand response problem is modeled as a Markov decision process, and a reinforcement learning environment is defined based on the Markov decision process. Reinforcement learning models are generated based on deep Q-networks. Using the reinforcement learning model and based on the reinforcement learning environment, the user profile feature vector is adjusted, and the parameters of the initial user profile are updated using a closed-loop feedback optimization mechanism. The user profile feature weights are adjusted using a profile feature adjustment mechanism to obtain and output the final user profile.
2. The power user demand response profile optimization method as described in claim 1, characterized in that, Data is acquired, preprocessed, and then classified to generate an initial user profile. Specifically, a clustering algorithm is used to divide users into different demand response groups, and the initial user profile is generated for each of the different demand response groups.
3. The power user demand response profile optimization method as described in claim 2, characterized in that, The demand response problem is modeled as a Markov decision process, and a reinforcement learning environment is defined based on the Markov decision process. Specifically, the demand response problem is modeled as the Markov decision process, and the state, action, and reward function of the reinforcement learning environment are defined.
4. The power user demand response profile optimization method as described in claim 3, characterized in that, A reinforcement learning model is generated based on a deep Q-network. Specifically, the model is trained on a discrete space using a deep Q-network and iteratively trained using the Q-value update formula to generate the reinforcement learning model.
5. The power user demand response profile optimization method as described in claim 4, characterized in that, Using the reinforcement learning model and based on the reinforcement learning environment, the user profile feature vector is adjusted, and the parameters of the initial user profile are updated using a closed-loop feedback optimization mechanism, including the following steps: Based on the action definition in the reinforcement learning environment, the initial user profile is adjusted using the profile feature vector update formula. Feedback on demand response is collected through a closed-loop feedback optimization mechanism, and the feedback loss is calculated to determine the convergence state of the reinforcement learning model and the initial profile. Finally, the parameters of the reinforcement learning model are adjusted using the feedback formula.
6. The power user demand response profile optimization method as described in claim 5, characterized in that, The feature weights of the user profile are adjusted using a profile feature adjustment mechanism to obtain and output the final user profile. Specifically, based on the performance of the reinforcement learning model in the demand response, the feature weights in the initial user profile are adjusted using a feature weighting matrix.
7. The power user demand response profile optimization method as described in claim 6, characterized in that, The portrait feature weights are adjusted using a portrait feature adjustment mechanism. Specifically, the portrait feature weights are fed back to the closed-loop feedback mechanism in the strong chemical environment to provide support for the closed-loop feedback mechanism.
8. A power user demand response profile optimization system, employing the power user demand response profile optimization method as described in any one of claims 1 to 7, characterized in that, include: The data collection and profile building module acquires data, performs preprocessing and classification, generates an initial user profile, structures the initial user profile, and obtains the user profile feature vector. The reinforcement learning environment definition module models the demand response problem as a Markov decision process and defines the reinforcement learning environment based on the Markov decision process. The reinforcement learning modeling module generates reinforcement learning models based on deep Q-networks. The user profile optimization module uses the reinforcement learning model and the reinforcement learning environment to adjust the user profile feature vector and uses a closed-loop feedback optimization mechanism to update the parameters of the initial user profile. The information feedback and profile output module uses a profile feature adjustment mechanism to adjust the weights of profile features, obtain the final user profile, and output it.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the power user demand response profile optimization method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the power user demand response profile optimization method as described in any one of claims 1 to 7.