Virtual power plant personalized incentive method based on psychological account representation learning

Through the method of psychological account characterization learning, the problems of limited incentive effect and high cost in the demand response of virtual power plants are solved, and the user participation and long-term response stability are improved. Through self-calibration Transformer mechanism and multi-objective optimization, personalized incentives are generated, which improves user participation and system reliability.

CN120525286AInactive Publication Date: 2025-08-22GUIZHOU XIANGBIN NEW ENERGY TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510701687.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing virtual power plant demand response incentive methods fail to make full use of psychological account theory, resulting in limited incentive effects and high costs, inability to achieve precise and personalized incentives, and lack a mechanism to capture the dynamic changes of user psychological accounts, and cannot adapt to the dynamic changes in power grid status and user preferences, resulting in a decrease in user participation and poor long-term response stability.

Method used

Using a method based on psychological account representation learning, the high-level Transformer psychological account representation learning mechanism is self-calibrated, user characteristic data is obtained and processed, the dynamic weights of the user's four psychological accounts are calculated, personalized incentives are generated, and the user response probability is predicted. Combined with multi-objective optimization and psychological account transfer learning mechanism, incentive configuration is optimized.

Benefits of technology

It has increased user participation by 85.0%, reduced incentive costs by 73.0%, and improved long-term participation stability by 99.6%, and achieved accurate portrayal of users' multi-dimensional psychological characteristics and personalized incentive configuration, which has significantly improved the economicality and system reliability of virtual power plants' demand response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120525286A_ABST
    Figure CN120525286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power systems and artificial intelligence, in particular to a virtual power plant personalized incentive method based on psychological account representation learning. Comprising the following steps: S1, acquiring and processing user feature data which comprises a demographic feature, a historical behavior feature and a context feature; s2, processing the user characteristic data by using a self-calibration high-order transform psychological account representation learning mechanism to obtain user representation including four psychological accounts of economy, society, environment and convenience; s3, capturing mutual influence among accounts through a psychological account cross enhancement mechanism; s4, calculating dynamic weights of the four psychological accounts of the user; s5, generating and distributing personalized incentives for the four psychological accounts; and S6, predicting a user response probability and optimizing excitation configuration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power systems and artificial intelligence technologies, and in particular to a personalized incentive method for virtual power plants based on mental account representation learning. Background Art

[0002] As the global energy transition deepens, the power system is undergoing a transformation from traditional centralized power generation to a new distributed, renewable energy-based power system. In this process, demand-side management (DSM) has garnered widespread attention as a key means of balancing supply and demand and improving system resilience. Virtual power plants (VPs), an innovative power resource aggregation and management model, are rapidly gaining traction worldwide. By integrating distributed energy resources, controllable loads, and energy storage equipment into a centrally dispatched virtual power generation unit, they provide power services similar to those of traditional power plants.

[0003] Demand response is a core function in virtual power plant operations. By guiding users to change their electricity consumption behavior to respond to grid demand, it helps balance power supply and demand, alleviate grid pressure, and improve system reliability. However, the effective implementation of demand response relies heavily on active user participation, and motivating users to actively participate in demand response is a key challenge in virtual power plant operations.

[0004] Traditional demand response incentives are primarily based on economic incentives, such as time-of-day pricing, capacity pricing, and direct load control, encouraging user participation by offering financial subsidies or discounts. Existing technologies include designing price differences for different time periods to encourage users to shift their electricity consumption, and providing corresponding financial subsidies based on the user's adjustable load capacity. While these approaches have promoted user participation to a certain extent, their effectiveness is often limited, and they are subject to high costs and poor sustainability.

[0005] In recent years, researchers have gradually recognized that user decision-making in demand response is influenced by a variety of factors, including not only economic factors but also multidimensional psychological factors such as social norms, environmental awareness, and convenience. While some approaches have achieved some success by simply categorizing users and providing differentiated incentives for different types of users, these approaches remain relatively crude in their depiction of user psychology and fail to fully account for the complex psychological mechanisms underlying user decision-making.

[0006] Research in behavioral economics shows that human economic decision-making often defies the completely rational assumptions of traditional economics, but is instead influenced by various psychological factors and cognitive biases. Mental accounting theory, a key theory in behavioral economics, reveals that people often allocate resources across different mental accounts during emergency decision-making and make decisions based on the characteristics of these accounts. However, existing virtual power plant demand response incentive methods have not fully leveraged mental accounting theory to optimize incentive effectiveness.

[0007] On the technical level, with the development of artificial intelligence technology, deep learning and representation learning have shown great potential in user behavior modeling. However, existing modeling mainly focuses on the extraction of electricity consumption behavior characteristics and fails to effectively model the user's psychological decision-making mechanism, especially failing to capture the user's psychological account allocation strategy when facing different types of incentives.

[0008] The main problems with the current virtual power plant demand response incentive methods are as follows. First, the traditional single economic incentive method ignores the psychological characteristics of user decision-making behavior, concentrates all incentives on economic accounts, and fails to trigger users' social, environmental and convenience psychological accounts, resulting in limited incentive effects and high costs; second, although the existing group incentive method takes into account user differences, the portrayal of user psychological characteristics is still relatively rough, and fails to achieve refined personalized incentives. Finally, there is a lack of a mechanism to capture the dynamic changes of users' psychological accounts, and it cannot adapt to the time-varying characteristics of users' psychological preferences.

[0009] With the expansion of the scale of virtual power plants and the diversification of user types, these problems have become increasingly prominent, seriously restricting the effectiveness and economy of virtual power plant demand response. Therefore, there is an urgent need for a new method that can accurately characterize the characteristics of users as mostly luggage and realize personalized incentive configuration, so as to achieve efficient demand response incentives at low cost.

[0010] Furthermore, traditional incentive methods are typically designed in a static manner, making them difficult to adapt to the dynamic changes in grid conditions and user preferences. Over the long term, users may experience incentive fatigue, leading to decreased participation and a lack of effective mechanisms to maintain long-term incentives.

[0011] To address these challenges, designing a new method based on the mental accounting theory of behavioral economics and advanced deep learning technology that can accurately understand the psychological characteristics of customer users and achieve personalized incentive configuration has become the key to improving the demand response effect of virtual power plants. Summary of the Invention

[0012] The technical problem solved by the present invention is to provide a personalized incentive method for virtual power plants based on mental account representation learning, accurate representation and dynamic optimization configuration of user mental accounts, improve user participation, reduce incentive costs, and enhance long-term response stability.

[0013] The basic solution provided by the present invention is a personalized incentive method for a virtual power plant based on mental account representation learning, which is characterized by comprising the following steps: S1. Acquire and process user feature data, wherein the user feature data includes demographic features, historical behavior features, and contextual features; S2. Use the self-calibrated high-order transformer mental account representation learning mechanism to process user feature data and obtain user representations that include four mental accounts: economic, social, environmental, and convenience; S3. Capture the mutual influence between accounts through the cross-reinforcement mechanism of mental accounts; S4. Calculate the dynamic weights of the user's four mental accounts; S5. Generate and distribute personalized incentives for four mental accounts; S6. Predict user response probability and optimize incentive configuration.

[0014] Furthermore, the S1 comprises the following steps: S1.1. Obtaining user demographic data For each user i in the virtual power plant, demographic characteristics are extracted from the user database. The demographic characteristics include user age, gender, education level, family structure, residential type, income level, and occupation category:

[0015] Dimensions representing demographic characteristics, represents the jth demographic characteristic of user i; S1.2. Obtaining user historical behavior feature data , collect user behavior data on historical participation in demand response activities. The historical behavior characteristics include past participation rate, response time, historical power consumption pattern, device usage frequency, peak power adjustment capability, and historical incentive response sensitivity:

[0016] in Dimensions that characterize historical behavior, represents the j-th historical behavior feature value of user i at time t; S1.3. Obtaining system context feature data , the context features include time features, weather conditions, grid status, and social practices:

[0017] in The dimension representing the contextual features, represents the j-th context feature value at time t; S1.4. Use feature concatenation to construct user comprehensive input feature vector , the demographic features, user historical behavior features, and context feature vectors are concatenated to obtain a complete user input feature vector:

[0018] in, , is the total feature dimension after splicing; S1.5, normalize the concatenated feature vectors. Apply normalization to the numerical features in :

[0019] in, is the j-th feature value of user i, and are the mean and standard deviation of the j-th feature in the training data.

[0020] Furthermore, the step S2 includes the following steps: S2.1. Perform multi-head self-attention speed three, apply layer normalization to the input feature vector, and then generate query, key and value matrices through linear transformation:

[0021]

[0022]

[0023] in, is a learnable query, key, and value transformation matrix, is the dimension of the attention head, and LayerNorm is the layer normalization function; S2.2, based on FlashAttention-2 optimized multi-head self-attention calculation, the computational complexity is reduced to :

[0024] Where h is the number of attention heads, is the output projection matrix; S2.3. Applying SwiGLU feed-forward network to enhance representation capabilities:

[0025]

[0026]

[0027] in, , , is the weight matrix of the feedforward network, is the corresponding bias vector, is the hidden dimension of the feedforward network, represents the Hadamard product; S2.4, through adaptive mixing weight parameters Integrate the multi-head self-attention output and the feedforward network output to form a comprehensive user representation vector:

[0028] in, is an adaptive hybrid weight parameter, which is dynamically adjusted according to the current context to synthesize the representation vector Divide into four sub-vectors of equal length, corresponding to the representation of the four mental accounts of economy, society, environment and convenience:

[0029] in, are the representation vectors of the economic account, social account, environmental account and convenience account of user i at time t respectively.

[0030] Furthermore, the step S3 includes the following steps: S3.1. Calculate the cross-reinforcement strength between accounts. For each mental account , calculate the cross-enhancement strength of account j to account k:

[0031] in, is the cross-influence weight vector of account k, is the corresponding bias term, Is the sigmoid function, ensuring that the cross enhancement strength is Within the range, represents the j-th mental account representation of user i at time t; S3.2. Apply the cross-enhancement account update criteria and update the representation of each mental account based on the calculated cross-enhancement strength:

[0032] in, is the cross transformation matrix from account j to account k, which maps the representation of account j to a representation space compatible with account k; S3.3. After applying the cross-increment of mental accounts, the final mental account representation of user i at time t is updated to: .

[0033] Further, the S4 includes the following steps: S4.1. Combine user comprehensive criteria and system context features to form the input for weight calculation:

[0034] in, is the user comprehensive representation vector, is the system context feature vector; S4.2. Enhance the nonlinear expression capability of weight calculation through multi-layer perceptron:

[0035] in, is the weight matrix of the multilayer perceptron, is the corresponding bias vector, is the rectified linear unit activation function; S4.3. Apply Softmax normalization to the raw output of the multilayer perceptron:

[0036] in, are the normalized weights of the four mental accounts, corresponding to the economic account, social account, environmental account, and convenience account respectively; S4.4. Combine the representations of the four mental accounts according to their weights to obtain a weighted comprehensive representation of the user:

[0037] in, It is a comprehensive representation after weighted combination.

[0038] Furthermore, the step S5 includes the following steps: S5.1 builds an incentive generation network, which accepts the user's weighted comprehensive representation and the system budget constraint as input and outputs a four-dimensional incentive vector:

[0039] Multi-layer fully connected network with non-linear activation functions and normalization layers:

[0040] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, LayerNorm is the normalization function; S5.2. Based on the user weighted comprehensive representation and system budget constraints, the four-dimensional incentive vector is calculated through the incentive generation network:

[0041] in, Represent the incentive values ​​for the economic account, social account, environmental account, and convenience account respectively. For the economic account, the incentive value is the amount of monetary reward; for the social account, the incentive value indicates the intensity of social recognition and honor incentives; for the environmental account, the incentive value indicates the degree of emphasis on environmental impact and contribution; for the convenience account, the incentive value indicates the intensity of convenience compensation and service improvement; S5.3. Adjust the generated excitation vector:

[0042] Among them, N is the number of users who need to be motivated at the current moment, is the incentive budget constraint of the system at time t.

[0043] Furthermore, the step S6 includes the following steps: S6.1. Build a response prediction network that receives the user's four-dimensional incentive vector, mental account weights, and original feature vector as input and outputs the predicted probability of the user's response:

[0044] Specifically, it is a multi-layer fully connected network with a nonlinear activation function and a Sigmoid output layer:

[0045] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, Is the Sigmoid function, ensure that the output is in the range Inside, represents the probability of user response; S6.2. Based on the user's four-dimensional incentive vector, mental account weights, and original feature vector, the response prediction network calculates the probability of the user's response:

[0046] in, represents the probability prediction of user i responding to demand at time t; S6.3. Construct a multi-objective optimization framework, including three loss functions: prediction loss function, budget constraint loss function, and weight diversity loss function; Prediction loss function: Use the two-compartment cross loss to measure the accuracy of the response prediction:

[0047] in, It is a binary label of the user's actual response, 1 indicates response and 0 indicates no response; Budget constraint loss: Ensure that the total incentive does not exceed the system budget:

[0048] in, is the weight coefficient of the budget constraint loss; Weight diversity loss, which encourages diversity in weight distribution:

[0049] Among them, KL is KL divergence, which is a four-dimensional uniform distribution , is the weight of diversity loss S6.4. Calculate the total loss and optimize the model parameters, combining the three loss functions into a total loss function:

[0050] Update the model parameters using gradient descent:

[0051]

[0052] in, is the learning rate, and The total loss function is respectively and gradient.

[0053] Furthermore, the method further comprises the following steps: S7: Applying a user decision optimization mechanism, wherein S7 includes the following steps: S7.1. Construct a decision complexity evaluation network to assess the decision burden imposed by incentive schemes on users:

[0054]

[0055] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function; S7.2. Based on the user's four-dimensional incentive vector, mental account weight, and original feature vector, the decision burden index is calculated using the decision complexity evaluation network:

[0056] in, It represents the decision complexity caused by the incentive scheme to user i at time t. The larger the value, the higher the decision burden. S7.3. Optimize the incentive presentation method based on decision complexity. When the decision complexity exceeds the preset threshold, simplify the incentive presentation method:

[0057] Among them, Simplify is the incentive simplification function, is the complexity threshold parameter, and the simplification strategies include: when When you are working on something, keep the original thinking motivation unchanged; when , retain the incentives of the two mental accounts with the highest weights, and set the others to zero; when When , the value retains the weight to engage the incentive of the mental account, and the others are set to zero.

[0058] Furthermore, the method further comprises the following steps: S8. Implementing a mental account transfer learning mechanism, wherein S8 specifically includes the following steps: S8.1. Establish a meta-learning framework, using a model-independent learning method, to find initialization parameters through meta-training for a series of tasks that users have already thought about:

[0059] in, is the task distribution, which means adopting users from the existing user group. is the loss function for a single user task, is the initialization parameter obtained by meta-learning; S8.2. Based on the initialization parameters obtained by meta-learning and a small amount of new user interaction data, the model parameters for the new user are quickly adapted through gradient descent:

[0060] in, is the transfer learning rate, which controls the degree of deviation of the new parameters from the meta-parameters, is the gradient of the meta-loss function with respect to the meta-parameters; S8.3 uses different learning rates for different mental accounts and dynamically adjusts the learning rate based on historical response data:

[0061]

[0062] in, is the learning rate of the jth mental account at time t, is the learning rate adjustment coefficient, is the gradient of the damage function with respect to the j-th mental account parameter.

[0063] Furthermore, the method further comprises the following steps: S9. Deploy and operate the personalized incentive system for the virtual power plant, which specifically includes the following steps: S9.1. System initialization, loading pre-trained model parameters, including self-calibration high-order Transformer model parameters, mental accounting weight calculation model parameters, incentive generation and response prediction model parameters, and decision complexity assessment model parameters; S9.2. The system continuously receives user feature data and environmental context data, executes S1-S7, generates a personalized incentive plan for each user, and adjusts the incentive presentation method based on the decision complexity assessment results; S9.3. Collect real-time user response data to the incentive scheme as labels for model training and perform model update steps regularly; S9.4. Regularly evaluate system performance, including indicators such as user participation rate, incentive cost, and long-term participation stability. Adjust model hyperparameters and training strategies based on the evaluation results to continuously optimize system performance.

[0064] The principles and advantages of the present invention are: The present invention uses a self-calibrated high-order Transformer mental accounting representation learning mechanism to process user feature data. This mechanism includes three key components: multi-head self-attention calculation, SwiGLU feedforward network, and hybrid residual connection and mental accounting representation generation. The multi-head self-attention mechanism optimized based on FlashAttention-2 reduces the computational complexity to O(L√N log N), significantly improving the efficiency of representation learning; the SwiGLU feedforward network enhances representation capabilities; and the hybrid residual connection and mental accounting representation generation achieve accurate capture of user mental accounts. This combination of technologies enables the present invention to extract high-quality mental accounting representations from the complex and ever-changing user feature space.

[0065] The dynamic weight calculation mechanism for mental accounts in this paper is based on comprehensive user representations and system contextual features. It uses the Softmax function to ensure weight normalization and employs a multi-layer perceptron to enhance the nonlinearity of weight calculation. The dynamic weights of users and the weighted combination of these weights are calculated using w_i^(t) = Softmax(MLP([H_i^(t), C^(t)])) and S_i^(t) = ∑_j=1^4 w_i^(t)[j]·A_i^(t)[j]. This mechanism accurately reflects the dynamic changes in a user's mental account weights in different contexts, providing a scientific basis for incentivizing personalized configuration.

[0066] Regarding personalized incentive generation and allocation, this invention generates a four-dimensional incentive vector based on comprehensive user representation and budget constraints. This vector targets four mental accounts: economic, social, environmental, and convenience. The user's response probability is predicted based on the incentive vector and the weights of these mental accounts. Accurate incentive configuration and response prediction are achieved through the use of I_i^(t) = [I_i^econ, I_i^soc, I_i^env, I_i^conv] = f_θ(S_i^(t), B^(t)) and R_i^(t) = g_φ(I_i^(t), w_i^(t), X_i^(t)). This mechanism ensures optimal allocation of incentive resources, achieving the "micro-incentive, large-response" effect.

[0067] The present invention also adopts a multi-objective optimization training mechanism, combining binary cross-entropy loss, budget constraint loss, and KL divergence loss to balance prediction accuracy, budget constraints, and weight diversity. By balancing the importance of each loss term through weight coefficients, the model training is fully optimized.

[0068] As the preferred technical solution of the present invention, a psychological account cross-enhancement mechanism is introduced. The mutual influence between accounts is captured through C_{i,j,k}^(t)= σ(W_{jk}·A_i^(t)[j] + b_{jk}) and A_i^(t)[k]= A_i^(t)[k] + ∑_{j≠k} C_{i,j,k}^(t)·(W_{jk}^trans·A_i^(t)[j]), and the gated update mechanism is applied to dynamically adjust the cross-influence intensity to achieve the "micro-stimulus and big response" effect.

[0069] The present invention also includes an adaptive learning rate adjustment mechanism, which uses different learning rates for different mental accounts and dynamically adjusts the learning rate based on historical response data to accelerate the integration of new users into the system. Accurate parameter updates are achieved through η_j^(t+1) = η_j^(t)·exp(β·∇_θj L_total^(t)) and θ_j^(t+1) = θ_j^(t) - η_j^(t+1)·∇_θj L_total^(t).

[0070] In order to reduce user decision fatigue, the present invention introduces a user decision burden optimization mechanism, which evaluates the decision complexity and dynamically adjusts the incentive presentation method through D_i^(t) = h_ψ(I_i^(t), w_i^(t), X_i^(t)) and I_i^(t),opt = Simplify(I_i^(t), D_i^(t), τ).

[0071] To address the cold start problem of new users, this paper designs a mental account transfer learning mechanism. Through θ_new = θ_meta + γ·∇_θmeta L_meta and L_meta = E_{task~p(task)} [L_task(θ_meta)], the mental account model of new users is used to initialize the new user, and the meta-learning method is used to accelerate the modeling of new user mental accounts.

[0072] Through the above-mentioned technical solution, the present invention achieves precise characterization of user mental accounts and personalized configuration of incentives in virtual power plant demand response. Compared with the most advanced existing clustered incentive schemes, user participation rates increased by 85.0%, incentive costs decreased by 73.0%, and long-term participation stability improved by 99.6%. This invention effectively solves the technical problems of low participation, high costs, and poor response stability caused by the single economic incentive approach in traditional virtual power plant demand response, which ignores the multidimensional psychological characteristics of users. It provides a new technical path for virtual power plant demand response.

[0073] The beneficial effects of the present invention are as follows: First, by constructing a dynamic representation learning model that includes four mental accounts: economic, social, environmental, and convenience, it achieves accurate modeling of users' multi-dimensional psychological preferences, which is more comprehensive and accurate than traditional methods in characterizing users' psychological characteristics; second, the mental account representation learning mechanism based on the self-calibrated high-order Transformer architecture greatly improves the efficiency and quality of feature extraction and representation learning; third, through the mental account cross-enhancement mechanism, the mutual influence between accounts is captured, the "micro-incentive-large-response" effect is achieved, and the incentive cost is significantly reduced; fourth, the introduction of the user decision burden optimization mechanism reduces user decision fatigue and improves the user experience; finally, through the mental account transfer learning mechanism, the cold start problem of new users is effectively solved, and the integration process of new users is accelerated. These technological innovations have enabled the present invention to achieve significant results in improving user participation, reducing incentive costs, and enhancing long-term response stability, greatly improving the economic benefits and system reliability of virtual power plant demand response. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 This is a flow chart of an embodiment of a personalized incentive method for a virtual power plant based on mental account representation learning according to the present invention; Figure 2 This is a personalized incentive architecture diagram of an embodiment of a personalized incentive method for a virtual power plant based on mental account representation learning according to the present invention; Figure 3 This is a flowchart of a self-calibrated Transformer mental account representation learning process in an embodiment of a virtual power plant personalized incentive method based on mental account representation learning according to the present invention; Figure 4This is a flow chart of a mental account cross-enhancement mechanism in an embodiment of a virtual power plant personalized incentive method based on mental account representation learning according to the present invention; Figure 5 A flowchart of a user decision burden optimization mechanism of an embodiment of a virtual power plant personalized incentive method based on mental account representation learning according to the present invention; Figure 6 A flowchart of a mental account transfer learning mechanism according to an embodiment of a virtual power plant personalized incentive method based on mental account representation learning according to the present invention; DETAILED DESCRIPTION The following is further described in detail through specific implementation methods: The embodiment is basically as shown in the attached Figure 1 As shown: A personalized incentive method for a virtual power plant based on mental account representation learning is characterized by comprising the following steps: S1. Acquire and process user feature data, wherein the user feature data includes demographic features, historical behavior features, and contextual features; S2. Use the self-calibrated high-order transformer mental account representation learning mechanism to process user feature data and obtain user representations that include four mental accounts: economic, social, environmental, and convenience; S3. Capture the mutual influence between accounts through the cross-reinforcement mechanism of mental accounts; S4. Calculate the dynamic weights of the user's four mental accounts; S5. Generate and distribute personalized incentives for four mental accounts; S6. Predict user response probability and optimize incentive configuration.

[0075] Said S1 comprises the following steps: S1.1. Obtaining user demographic data For each user i in the virtual power plant, demographic characteristics are extracted from the user database. The demographic characteristics include user age, gender, education level, family structure, residential type, income level, and occupation category:

[0076] Dimensions representing demographic characteristics, represents the jth demographic characteristic of user i; S1.2. Obtaining user historical behavior feature data , collect user behavior data on historical participation in demand response activities. The historical behavior characteristics include past participation rate, response time, historical power consumption pattern, device usage frequency, peak power adjustment capability, and historical incentive response sensitivity:

[0077] in Dimensions that characterize historical behavior, represents the j-th historical behavior feature value of user i at time t; S1.3. Obtaining system context feature data , the context features include time features, weather conditions, grid status, and social practices:

[0078] in The dimension representing the contextual features, represents the j-th context feature value at time t; S1.4. Use feature concatenation to construct user comprehensive input feature vector , the demographic features, user historical behavior features, and context feature vectors are concatenated to obtain a complete user input feature vector:

[0079] in, , is the total feature dimension after splicing; the feature splicing operation forms a higher-dimensional vector by connecting the feature vectors from different sources end to end, effectively retaining the complete information of various features. Figure 1 As shown, the feature processing module receives multi-source data and outputs a comprehensive feature vector.

[0080] S1.5, normalize the concatenated feature vectors. Apply normalization to the numerical features in :

[0081] in, is the j-th eigenvalue of user i, and are the mean and standard deviation of the j-th feature in the training data.

[0082] The S2 is as Figure 3 The following steps are shown: S2.1. Perform multi-head self-attention speed three, apply layer normalization to the input feature vector, and then generate query, key and value matrices through linear transformation:

[0083]

[0084]

[0085] in, is a learnable query, key, and value transformation matrix, is the dimension of the attention head, and LayerNorm is the layer normalization function; S2.2, based on FlashAttention-2 optimized multi-head self-attention calculation, the computational complexity is reduced to :

[0086] Where h is the number of attention heads, is the output projection matrix; S2.3. Applying SwiGLU feed-forward network to enhance representation capabilities:

[0087]

[0088]

[0089] in, , , is the weight matrix of the feedforward network, is the corresponding bias vector, is the hidden dimension of the feedforward network, represents the Hadamard product; S2.4, through adaptive mixing weight parameters Integrate the multi-head self-attention output and the feedforward network output to form a comprehensive user representation vector:

[0090] in, is an adaptive hybrid weight parameter, which is dynamically adjusted according to the current context to synthesize the representation vector Divide into four sub-vectors of equal length, corresponding to the representation of the four mental accounts of economy, society, environment and convenience:

[0091] in, are the representation vectors of the economic account, social account, environmental account and convenience account of user i at time t respectively.

[0092] The S3 Figure 4 The following steps are shown: S3.1. Calculate the cross-reinforcement strength between accounts. For each mental account , calculate the cross-enhancement strength of account j to account k:

[0093] in, is the cross-influence weight vector of account k, is the corresponding bias term, Is the sigmoid function, ensuring that the cross enhancement strength is Within the range, represents the j-th mental account representation of user i at time t; S3.2. Apply the cross-enhancement account update criteria and update the representation of each mental account based on the calculated cross-enhancement strength:

[0094] in, is the cross transformation matrix from account j to account k, which maps the representation of account j to a representation space compatible with account k; S3.3. After applying the cross-increment of mental accounts, the final mental account representation of user i at time t is updated to: .

[0095] The S4 comprises the following steps: S4.1. Combine user comprehensive criteria and system context features to form the input for weight calculation:

[0096] in, is the user comprehensive representation vector, is the system context feature vector; S4.2. Enhance the nonlinear expression capability of weight calculation through multi-layer perceptron:

[0097] in, is the weight matrix of the multilayer perceptron, is the corresponding bias vector, is the rectified linear unit activation function; S4.3. Apply Softmax normalization to the raw output of the multilayer perceptron:

[0098] in, are the normalized weights of the four mental accounts, corresponding to the economic account, social account, environmental account, and convenience account respectively; S4.4. Combine the representations of the four mental accounts according to their weights to obtain a weighted comprehensive representation of the user:

[0099] in, It is a comprehensive representation after weighted combination.

[0100] The S5 comprises the following steps: S5.1 builds an incentive generation network, which accepts the user's weighted comprehensive representation and the system budget constraint as input and outputs a four-dimensional incentive vector:

[0101] Multi-layer fully connected network with non-linear activation functions and normalization layers:

[0102] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, LayerNorm is the normalization function; S5.2. Based on the user weighted comprehensive representation and system budget constraints, the four-dimensional incentive vector is calculated through the incentive generation network:

[0103] in, Represent the incentive values ​​for the economic account, social account, environmental account, and convenience account respectively. For the economic account, the incentive value is the amount of monetary reward; for the social account, the incentive value indicates the intensity of social recognition and honor incentives; for the environmental account, the incentive value indicates the degree of emphasis on environmental impact and contribution; for the convenience account, the incentive value indicates the intensity of convenience compensation and service improvement; S5.3. Adjust the generated excitation vector:

[0104] Among them, N is the number of users who need to be motivated at the current moment, is the incentive budget constraint of the system at time t.

[0105] The S6 comprises the following steps: S6.1. Build a response prediction network that receives the user's four-dimensional incentive vector, mental account weights, and original feature vector as input and outputs the predicted probability of the user's response:

[0106] Specifically, it is a multi-layer fully connected network with a nonlinear activation function and a Sigmoid output layer:

[0107] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, Is the Sigmoid function, ensure that the output is in the range Inside, represents the probability of user response; S6.2. Based on the user's four-dimensional incentive vector, mental account weights, and original feature vector, the response prediction network calculates the probability of the user's response:

[0108] in, represents the probability prediction of user i responding to demand at time t; S6.3. Construct a multi-objective optimization framework, including three loss functions: prediction loss function, budget constraint loss function, and weight diversity loss function; Prediction loss function: Use the two-compartment cross loss to measure the accuracy of the response prediction:

[0109] in, It is a binary label of the user's actual response, 1 indicates response and 0 indicates no response; Budget constraint loss: Ensure that the total incentive does not exceed the system budget:

[0110] in, is the weight coefficient of the budget constraint loss; Weight diversity loss, which encourages diversity in weight distribution:

[0111] Among them, KL is KL divergence, which is a four-dimensional uniform distribution , is the weight of diversity loss S6.4. Calculate the total loss and optimize the model parameters, combining the three loss functions into a total loss function:

[0112] Update the model parameters using gradient descent:

[0113]

[0114] in, is the learning rate, and The total loss function is respectively and gradient.

[0115] The following steps are also included: S7, applying the user decision optimization mechanism, said S7 is as follows Figure 5 The following steps are involved: S7.1. Construct a decision complexity evaluation network to assess the decision burden imposed by incentive schemes on users:

[0116]

[0117] in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function; S7.2. Based on the user's four-dimensional incentive vector, mental account weight, and original feature vector, the decision burden index is calculated using the decision complexity evaluation network:

[0118] in, It represents the decision complexity caused by the incentive scheme to user i at time t. The larger the value, the higher the decision burden. S7.3. Optimize the incentive presentation method based on decision complexity. When the decision complexity exceeds a preset threshold, simplify the incentive presentation method:

[0119] Among them, Simplify is the incentive simplification function, is the complexity threshold parameter, and the simplification strategies include: when When you are working on something, keep the original thinking motivation unchanged; when , retain the incentives of the two mental accounts with the highest weights, and set the others to zero; when When , the value retains the weight to engage the incentive of the mental account, and the others are set to zero.

[0120] The following steps are also included: S8. Implement the mental account transfer learning mechanism. Figure 6 The specific steps include: S8.1. Establish a meta-learning framework, using a model-independent learning method, to find initialization parameters through meta-training for a series of tasks that users have already thought about:

[0121] in, is the task distribution, which means adopting users from the existing user group. is the loss function for a single user task, is the initialization parameter obtained by meta-learning; S8.2. Based on the initialization parameters obtained by meta-learning and a small amount of new user interaction data, the model parameters for the new user are quickly adapted through gradient descent:

[0122] in, is the transfer learning rate, which controls the degree of deviation of the new parameters from the meta-parameters, is the gradient of the meta-loss function with respect to the meta-parameters; S8.3 uses different learning rates for different mental accounts and dynamically adjusts the learning rate based on historical response data:

[0123]

[0124] in, is the learning rate of the jth mental account at time t, is the learning rate adjustment coefficient, is the gradient of the damage function with respect to the j-th mental account parameter.

[0125] The following steps are also included: S9. Deploy and operate the personalized incentive system for the virtual power plant, which specifically includes the following steps: S9.1. System initialization, loading pre-trained model parameters, including self-calibration high-order Transformer model parameters, mental accounting weight calculation model parameters, incentive generation and response prediction model parameters, and decision complexity assessment model parameters; S9.2. The system continuously receives user feature data and environmental context data, executes S1-S7, generates a personalized incentive plan for each user, and adjusts the incentive presentation method based on the decision complexity assessment results; S9.3. Collect real-time user response data to the incentive scheme as labels for model training and perform model update steps regularly; S9.4. Regularly evaluate system performance, including indicators such as user participation rate, incentive cost, and long-term participation stability. Adjust model hyperparameters and training strategies based on the evaluation results to continuously optimize system performance.

[0126] The following is an explanation of the implementation process of this solution through specific actual cases: This example was tested in a northern winter heating region, where power grid fluctuations are significant, to verify the adaptability and stability of the proposed method under extreme load fluctuations. The test platform included 2,800 residential, 85 commercial, and 23 industrial users, all equipped with smart meters and home / enterprise energy management systems. The hardware environment consisted of an AMD EPYC7763 CPU (64 cores, 2.45GHz), 512GB of RAM, and two NVIDIA H100 GPUs (80GB). The software environment used Python 3.10, PyTorch 2.0.2, and the proprietary distributed computing framework VirtualPlantCore 3.5. The experiment ran from November 15, 2024, to February 15, 2025, covering the entire peak heating season. During this period, 78 demand response events were triggered, an average of six per week, with individual durations ranging from 1.5 to 3 hours, depending on the actual grid load. The ambient temperature fluctuates from -25°C to 5°C, with daily temperature differences of up to 30°C, providing ideal conditions for testing the system's performance in extreme environments.

[0127] Comparison plan 1: Dynamic Economic Incentive Scheme: This scheme utilizes a dynamic pricing mechanism based on real-time grid load. The incentive amount increases exponentially with load intensity, starting at 0.5 yuan / kWh and rising to 2.5 yuan / kWh during periods of load intensity. This scheme uses historical data to train an LSTM prediction model before implementation, enabling 24-hour forecasts of load curves and determining incentive amounts. This scheme represents a relatively advanced economic incentive mechanism in the current electricity market and has been implemented in multiple regional power grids. While it emphasizes the transmission of economic signals, it lacks consideration of users' non-economic psychological factors and lacks customization capabilities, resulting in the same incentive level for all users.

[0128] Comparison plan 2: A personalized hybrid incentive scheme combines machine learning and behavioral economics principles to build a personalized incentive model based on historical user response data. This scheme uses the XGBoost algorithm to update the user model weekly, providing each user with a hybrid incentive package consisting of economic incentives, environmental points, and community rankings. While this scheme achieves a certain degree of personalization, its low update frequency fails to capture short-term changes in user preferences. Furthermore, its mental account segmentation is crude, considering only the economic, environmental, and social dimensions while ignoring convenience. Furthermore, this scheme fails to consider the cross-influence between mental accounts, calculating each incentive dimension independently and failing to generate synergistic effects.

[0129] Experimental steps: First, data collection and feature engineering are carried out to collect basic user information, device parameters, historical electricity consumption data, and historical response records. For residential users, 14 demographic characteristics are collected, including age, occupation, family structure, income level, and education level; device information includes 9 items such as heating equipment type, number of smart appliances, and total installed capacity; historical behavioral data includes 12 time series features such as hourly electricity consumption, peak-to-valley ratio, and electricity consumption volatility over the past three months. For commercial and industrial users, additional specific features are collected, such as industry type, production plan flexibility, and backup power supply status. Contextual features include outdoor temperature, humidity, wind speed, weather conditions, weekday / holiday marking, and time period type. All features are screened through a feature selection algorithm, and ultimately 42 valid features are retained.

[0130] Then the model is deployed and optimized. Comparative solution 1 uses the price elasticity model in economics to determine the incentive amount under different load levels, and trains the load forecasting model based on historical data. Comparative solution 2 uses the XGBoost algorithm to train the user response prediction model, and determines the personalized incentive parameters based on historical response data. The solution of the present invention first makes specific adjustments to the self-calibrated high-order Transformer, increases the number of attention heads to 12, to adapt to the complexity of user behavior in extreme winter environments; at the same time, optimizes the SwiGLU feedforward network structure, and expands the number of neurons in the middle layer from 1024 to 2048 to enhance the model's expression ability. With regard to the cross-enhancement mechanism of psychological accounts, the cross-influence between economic accounts and convenience accounts is particularly strengthened in this embodiment, and the weight coefficient is increased from 0.35 of the standard configuration to 0.65 to meet the special needs of users for convenience in extreme weather conditions.

[0131] Then, when the demand response event is triggered, the test is implemented. To ensure fairness, the same number of users are randomly selected to apply the three incentive schemes. For the scheme of the present invention, a streaming computing framework is used to process user data in real time, the mental account weight is updated every 15 minutes, and the incentive combination is dynamically adjusted. During extreme load events (grid frequency is lower than 49.85Hz or the load exceeds 92% of the design capacity), the scheme of the present invention starts the "deep characterization extraction" mode, and the sampling frequency is increased to 5 minutes / time to achieve more accurate user status capture. At the same time, in order to verify the robustness of the scheme, 30% of the events were randomly selected during the experiment to simulate communication delays or data loss to test the fault tolerance of the system.

[0132] Finally, data collection, analysis, and evaluation were conducted. Smart meter terminals recorded user response data in real time, including load reduction, response delay, duration of response, and post-event rebound. User experience data was collected through instant feedback from mobile apps, using a 5-minute sampling method to record user decision-making processes and subjective experiences. System performance data was collected via a distributed monitoring platform, including metrics such as computational latency, memory usage, communication overhead, and model accuracy. After outlier detection and data cleaning, Bayesian A / B testing was used to evaluate the effectiveness of different solutions, with a confidence level of 99.5% (p < 0.005) to ensure the reliability of the conclusions.

[0133] Test methods and standards: To comprehensively evaluate the performance of each solution, this embodiment designs a multi-dimensional testing method. The first is the response effect evaluation, using the weighted response index (WRI) to measure the quality of user responses. The calculation formula is WRI = 0.5×(DR / DRmax) + 0.3×(1-DT / DTmax) + 0.2×(DR_dur / T), where DR is the actual load reduction, DRmax is the target load reduction, DT is the response delay time, DTmax is the maximum acceptable delay time (set to 10 minutes), DR_dur is the continuous response time, and T is the total duration of the event. This indicator comprehensively considers the response volume, response speed, and response continuity to more comprehensively reflect the response quality.

[0134] Next, economic benefits are assessed. The integrated rate of return (CIRR) is used to measure the input-output ratio. The calculation formula is CIRR = (VDR - IC - OC) / (IC + OC), where VDR is the market value of load reduction, IC is the incentive cost, and OC is the operating cost. The marginal diminishing returns (MDER) is also used to assess the sensitivity of increased incentives to increased response volume. The formula is MDER = (△DR / DR) / (△IC / IC). A value closer to 1 indicates higher incentive efficiency.

[0135] The third area is user experience assessment. This involves using a modified Technology Acceptance Model (TAM) questionnaire, comprised of 18 questions across four dimensions: perceived usefulness, perceived ease of use, perceived enjoyment, and continued use intention. Furthermore, the mobile app records user decision-making process data, including information viewing time, decision completion time, number of steps, and number of interface interactions. This data is then used to construct a cognitive load index (CLI) to assess decision burden.

[0136] The fourth step is system performance evaluation, which involves testing throughput, latency, scalability, and stability. Throughput testing assesses system processing capabilities by simulating concurrent requests from users of varying scales (up to 100,000 queries per second). Latency testing records the end-to-end time from request issuance to response completion. Scalability testing analyzes the relationship between system resource consumption and user scale. Stability testing assesses system recovery capabilities by injecting faults (network outages, server downtime, data loss, etc.).

[0137] Finally, the environmental adaptability assessment specifically tests system performance under extreme weather conditions. Temperature Sensitivity Analysis (TSA) assesses the impact of varying temperature conditions on user responses; Temperature-Response Elasticity Coefficient (TREC) quantifies the impact of temperature changes on response volume; and Extreme Event Response Rate (EERR) evaluates system performance under extreme deviations in grid frequency or load.

[0138] The specific experimental results are shown in Table 1: Table 1

[0139] Under extreme load conditions, the proposed solution demonstrated remarkable performance advantages. First, the average response rate under extreme load reached 92.4%, 41.9% higher than Comparative Solution 2. This significant improvement stems from the self-calibrated high-order Transformer in our invention, which accurately captures the dramatic changes in users' psychological states under extreme conditions, particularly behavioral adjustments to sudden temperature fluctuations. Furthermore, through a cross-reinforcement mechanism for mental accounts, it strengthens the influence of economic accounts on convenience accounts, effectively alleviating users' excessive focus on convenience in extremely cold weather. This mechanism enables the system to leverage the weight of users' convenience accounts with relatively low economic incentives, achieving a "small win, big win" effect.

[0140] The temperature sensitivity of the proposed solution is only 0.37% / °C, 78.0% lower than that of Comparative Solution 2, demonstrating the system's strong adaptability to ambient temperature fluctuations. This is because mental account representation learning can capture the impact of temperature changes on users' mental account weights in real time and offset this impact by dynamically adjusting the incentive mix. In particular, when the temperature drops sharply, the system automatically increases the weight of the social account and convenience account, weakening the role of the economic account, thus avoiding the vicious cycle of "lower temperature, higher incentives required."

[0141] Response latency is a key indicator of demand response speed. The average response latency for the solution presented in this invention is only 1.7 minutes, a 67.3% reduction compared to Comparative Solution 2. This efficient response stems from a decision-burden optimization mechanism that assesses the user's current cognitive load and, accordingly, simplifies the incentive presentation, significantly reducing decision complexity. For example, during the evening rush hour during extreme weather, when user anxiety levels are high, the system automatically simplifies the decision interface, transforming complex, multi-dimensional incentive combinations into a straightforward, one-click response option while preserving the internal structure of the incentive combination.

[0142] In terms of system performance, despite employing a complex deep learning model, the optimized computing architecture and FlashAttention-2 technology reduce CPU consumption per inference to just 0.18 seconds, a 72.3% reduction compared to Comparative Solution 2. This advantage enables the system to run on resource-constrained edge devices, significantly reducing deployment costs. Especially when communication bandwidth is limited, the local inference capability ensures continuous system operation and improves overall reliability.

[0143] User experience testing showed that under extreme weather conditions, the proposed solution achieved a high satisfaction rating of 4.7 out of 5, a 46.9% improvement over Comparative Solution 2. In-depth interviews revealed that users particularly valued the system's "intelligent understanding" capability, accurately grasping the changing psychological needs of users in different environments. 88.3% of users stated that the system's incentive package perfectly met their diverse needs in extreme weather, taking into account both financial compensation and comfort and convenience, making them feel understood and respected.

[0144] Response continuity and load rebound are key indicators of demand response quality. The solution of the present invention achieved a high response continuity of 95.3% and a load rebound of only 8.4%, representing improvements of 31.6% and a reduction of 70.6%, respectively, compared to Comparative Solution 2. This demonstrates that the present invention effectively alleviates user "response fatigue," particularly during periods of consecutive days of extreme weather. By dynamically adjusting mental account weights and incentive combinations, users' willingness to respond is maintained. The low load rebound is attributed to the cross-reinforcement mechanism of mental accounts. This mechanism, by strengthening the role of environmental and social accounts, fosters more stable user response habits and reduces "retaliatory" electricity consumption.

[0145] In terms of new user adaptation, the solution of the present invention requires only 1.8 events for a new user to reach a stable response state, a 75.0% reduction compared to Comparative Solution 2. This is due to the mental account transfer learning mechanism, which can quickly build a mental account model for a new user based on a small amount of interaction data. In particular, for users from the south who have moved to the north, the system can identify their electricity usage habits and quickly adjust model parameters to help them adapt to the demand response mechanism in this new environment.

[0146] Finally, in terms of system stability, the proposed solution performed exceptionally well despite network outages, server downtime, and other failures, achieving an average recovery time of only 3.4 seconds, an 87.7% improvement over Comparative Solution 2. This is due to its layered fault recovery mechanism and local inference backup strategy. Even in the event of a complete communication outage, the local model can still provide near-optimal incentives based on cached user data, ensuring continuous system operation. Across 24 random failures simulated during the experiment, the proposed solution maintained 89.7% service availability, significantly higher than the 62.3% achieved by the comparative solution.

[0147] Example 2: This example proposes an implementation plan for ensuring long-term user response stability, primarily addressing the issues of diminishing incentives and user fatigue that can occur during long-term demand response for virtual power plants. This solution combines a self-calibrating high-order Transformer mental accounting representation learning mechanism, a mental accounting cross-reinforcement mechanism, and a user decision burden optimization mechanism to build a complete system for ensuring long-term user stability.

[0148] First, this embodiment uses an expanded feature set to obtain more comprehensive long-term user behavior data. Compared to the basic embodiment, it adds long-term stability-related features such as the user's historical response decay rate, seasonal participation pattern, and social network influence factor. The feature vector is represented as X_i^(t) = Concat[F_demo(i), F_behavior(i,t), F_context(t), F_decay(i,t), F_seasonal(i,t), F_social(i,t)], where F_decay(i,t) represents the response decay rate feature of user i before time t, F_seasonal(i,t) represents the seasonal participation pattern feature of user i, and F_social(i,t) represents the social network influence factor of user i. The response decay rate is calculated using exponential smoothing: F_decay(i,t) = α·R_i^(t-1) + (1-α)·F_decay(i,t-1), where α is the smoothing coefficient, set to 0.3, and R_i^(t-1) is the user's response behavior at the previous moment. Seasonal participation pattern characteristics are derived by extracting the periodic components of historical user participation data through Fourier time series decomposition. Social network influence factors are constructed by processing user social relationship data using graph neural networks.

[0149] During the mental accounting representation learning phase, this embodiment optimizes the self-calibrated high-order Transformer architecture by adding a time-aware layer and an adaptive history memory unit. The time-aware layer enhances the temporal information representation of the feature vector through position encoding: X_i^(t)_time = X_i^(t) + PE(t), where PE(t) represents the temporal position encoding, calculated using the alternating sine-cosine form: PE(t)_2j = sin(t / 10000^(2j / d)), PE(t)_2j+1 = cos(t / 10000^(2j / d)), where j is the position index and d is the feature dimension. The adaptive historical memory unit stores the user's mental account representations for the past k time windows and calculates the correlation between the current representation and the historical representation through the attention mechanism: A_i^(t)_hist =Attention(A_i^(t), [A_i^(t-1), A_i^(t-2), ..., A_i^(tk)]), where k is set to 12, corresponding to one year's monthly data, which can effectively capture the user's annual behavioral pattern changes.

[0150] To address the dynamic nature of long-term users' mental accounts, this embodiment introduces an enhanced cross-reinforcement mechanism for mental accounts. This mechanism not only considers the immediate cross-influence between accounts but also incorporates a time decay factor and a periodic reinforcement function. The time decay factor naturally decays the cross-reinforcement effect over time: C_{i,j,k}^(t) = σ(W_{jk}·A_i^(t)[j] + b_{jk})·exp(-λ·Δt), where λ is the decay coefficient, set to 0.05, and Δt is the time since the last stimulus. The periodic reinforcement function boosts the weight of specific mental accounts at specific time points (such as holidays and seasonal changes): w_i^(t)[j] = w_i^(t)[j] + γ·sin(2π·t / T + φ_j), where γ is the periodic reinforcement strength, set to 0.15, T is the period length (e.g., 365 days), and φ_j is the phase offset of each mental account. Through this dynamic time adjustment mechanism, the system can adapt to the natural changes in users' long-term behavior patterns and avoid the attenuation of incentive effects over time.

[0151] To address decision fatigue among long-term users, this embodiment further improves the user decision burden optimization mechanism. This introduces a dynamic decision complexity assessment model that comprehensively considers a user's historical decision-making behavior, current cognitive load, and the complexity of the incentive scheme: D_i^(t) = h_ψ(I_i^(t), w_i^(t), X_i^(t), H_i^(t)), where H_i^(t) represents the characteristics of the user's historical decision-making behavior. When the user's decision complexity exceeds a threshold (D_i^(t) > τ), the system initiates a three-level simplification strategy: in the first level, the incentive components of low-weighted mental accounts are merged; in the second level, the incentive components of the two highest-weighted mental accounts are retained; and in the third level, the incentive scheme based on the user's best historical response is automatically adopted. The simplification process is expressed as: I_i^(t),opt = Simplify_level(I_i^(t), D_i^(t), τ), where the simplification level is determined by the difference between D_i^(t) and τ. Experiments show that this optimization mechanism can significantly reduce the decision-making burden of users in the long-term participation process and improve user experience satisfaction by 79.8%.

[0152] This embodiment also incorporates an incentive variation and innovation mechanism for long-term users. By regularly introducing novel incentive elements, it prevents users from becoming accustomed to fixed incentive models. The variation incentive generation function is: I_i^(t)_variant = I_i^(t) + ε·Novel(t, w_i^(t)), where ε is the variation intensity parameter, ranging from [0.1 to 0.3]. Novel(t, w_i^(t)) is a novel incentive generation function based on time and the weight of the user's mental account. This function automatically adjusts the incentive combination at regular intervals (e.g., 30 days), changing the presentation and structure of the incentives while maintaining the overall incentive intensity. For example, for environmental mental accounts, the system might change from "displaying reduced carbon emissions" to "visualizing protected forest area," stimulating user interest and motivation for continued participation.

[0153] To evaluate the stability of users' responses to long-term incentives, this embodiment designs a multi-dimensional response stability evaluation index, including the response rate fluctuation coefficient CV_i^(t), the incentive sensitivity change rate S_i^(t), and the long-term satisfaction trend T_i^(t). The response rate fluctuation coefficient is calculated as follows: CV_i^(t) = std(R_i^(tk:t)) / mean(R_i^(tk:t)), where k is the evaluation window length. The incentive sensitivity change rate is expressed as: S_i^(t) = (R_i^(t) / I_i^(t)) / (R_i^(tk) / I_i^(tk)). The long-term satisfaction trend is obtained by the linear regression slope of user feedback data. The system dynamically adjusts the user's incentive strategy based on these indicators: when a downward trend in the user stability indicator is detected, the automatic adjustment mechanism is triggered to rebalance the weight distribution of the four mental accounts and appropriately increase the cross-reinforcement strength.

[0154] During the model training phase, this embodiment adopts a multi-timescale optimization strategy, taking into account both short-term response accuracy and long-term stability. The loss function is expanded to: L_total = L_pred + L_budget + L_div + λ_3·L_stability, where L_stability = mean(CV_i^(t)) + β·mean(|1-S_i^(t)|), λ_3 is a trade-off parameter set to 0.4, and β is a sensitivity balance coefficient set to 0.6. By introducing the stability loss term, the model can optimize the short-term response effect while ensuring long-term response stability.

[0155] To address the natural decline in user interest, this embodiment introduces an adaptive incentive reinforcement mechanism. This mechanism monitors the changing trend of user response rates and adjusts the incentive strategy in advance when potential decline signals are detected. The decline signal detection function is: D_signal(i,t) = 1 if mean(R_i^(t-3:t))<0.85·mean(R_i^(t-6:t-3)) else 0. When a decline signal is detected, the system temporarily increases the incentive strength of the user's most heavily weighted mental account: I_i^(t)[j_max] = I_i^(t)[j_max]·(1 + boost_factor), where j_max is the index of the most heavily weighted mental account and boost_factor is the reinforcement coefficient. The initial value is set to 0.25 and then dynamically adjusted based on user response. Experiments have shown that this mechanism can effectively prevent user interest decline and reduce long-term participation rate fluctuations by 67.3%.

[0156] The above are only embodiments of the present invention. Common knowledge such as the known specific structures and characteristics in the scheme are not described in detail here. Ordinary technicians in the field are aware of all common technical knowledge in the technical field of the invention before the application date or priority date, can obtain all existing technologies in the field, and have the ability to apply conventional experimental means before that date. Ordinary technicians in the field can improve and implement this scheme in combination with their own abilities under the inspiration given by this application. Some typical known structures or known methods should not become obstacles for ordinary technicians in the field to implement this application. It should be pointed out that for those skilled in the art, without departing from the structure of the present invention, several variations and improvements can be made, which should also be regarded as the scope of protection of the present invention. These will not affect the effect of the implementation of the present invention and the practicality of the patent. The scope of protection required by this application shall be based on the content of its claims, and the specific implementation methods and other records in the specification can be used to interpret the content of the claims.

Claims

1. A personalized incentive method for virtual power plants based on mental account representation learning, characterized by: The following steps are involved: S1. Acquire and process user feature data, wherein the user feature data includes demographic features, historical behavior features, and contextual features; S2. Use the self-calibrated high-order transformer mental account representation learning mechanism to process user feature data and obtain user representations that include four mental accounts: economic, social, environmental, and convenience; S3. Capture the mutual influence between accounts through the cross-reinforcement mechanism of mental accounts; S4. Calculate the dynamic weights of the user's four mental accounts; S5. Generate and distribute personalized incentives for four mental accounts; S6. Predict user response probability and optimize incentive configuration.

2. A virtual power plant personalized incentive method based on mental account representation learning according to claim 1, characterized in that: Said S1 comprises the following steps: S1.

1. Obtaining user demographic data For each user i in the virtual power plant, demographic characteristics are extracted from the user database. The demographic characteristics include user age, gender, education level, family structure, residential type, income level, and occupation category: Dimensions representing demographic characteristics, represents the jth demographic characteristic of user i; S1.

2. Obtaining user historical behavior feature data , collect user behavior data on historical participation in demand response activities. The historical behavior characteristics include past participation rate, response time, historical power consumption pattern, device usage frequency, peak power adjustment capability, and historical incentive response sensitivity: in Dimensions that characterize historical behavior, represents the j-th historical behavior feature value of user i at time t; S1.

3. Obtaining system context feature data , the context features include time features, weather conditions, grid status, and social practices: in The dimension representing the contextual features, represents the j-th context feature value at time t; S1.

4. Use feature concatenation to construct user comprehensive input feature vector , the demographic features, user historical behavior features, and context feature vectors are concatenated to obtain a complete user input feature vector: in, , is the total feature dimension after splicing; S1.5, normalize the concatenated feature vectors. Apply normalization to the numerical features in : in, is the j-th feature value of user i, and are the mean and standard deviation of the j-th feature in the training data.

3. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 2 is characterized by: The S2 comprises the following steps: S2.

1. Perform multi-head self-attention speed three, apply layer normalization to the input feature vector, and then generate query, key and value matrices through linear transformation: in, is a learnable query, key, and value transformation matrix, is the dimension of the attention head, and LayerNorm is the layer normalization function; S2.2, based on FlashAttention-2 optimized multi-head self-attention calculation, the computational complexity is reduced to : Where h is the number of attention heads, is the output projection matrix; S2.

3. Applying SwiGLU feed-forward network to enhance representation capabilities: in, , , is the weight matrix of the feedforward network, is the corresponding bias vector, is the hidden dimension of the feedforward network, represents the Hadamard product; S2.4, through adaptive mixing weight parameters Integrate the multi-head self-attention output and the feedforward network output to form a comprehensive user representation vector: in, is an adaptive hybrid weight parameter, which is dynamically adjusted according to the current context to synthesize the representation vector Divide into four sub-vectors of equal length, corresponding to the representation of the four mental accounts of economy, society, environment and convenience: in, are the representation vectors of the economic account, social account, environmental account and convenience account of user i at time t respectively.

4. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 3 is characterized by: The S3 includes the following steps: S3.

1. Calculate the cross-reinforcement strength between accounts. For each mental account , calculate the cross-enhancement strength of account j to account k: in, is the cross-influence weight vector of account k, is the corresponding bias term, Is the sigmoid function, ensuring that the cross enhancement strength is Within the range, represents the j-th mental account representation of user i at time t; S3.

2. Apply the cross-enhancement account update criteria and update the representation of each mental account based on the calculated cross-enhancement strength: in, is the cross transformation matrix from account j to account k, which maps the representation of account j to a representation space compatible with account k; S3.

3. After applying the cross-increment of mental accounts, the final mental account representation of user i at time t is updated to: 。 5. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 4 is characterized by: The S4 comprises the following steps: S4.

1. Combine user comprehensive criteria and system context features to form the input for weight calculation: in, is the user comprehensive representation vector, is the system context feature vector; S4.

2. Enhance the nonlinear expression capability of weight calculation through multi-layer perceptron: in, is the weight matrix of the multilayer perceptron, is the corresponding bias vector, is the rectified linear unit activation function; S4.

3. Apply Softmax normalization to the raw output of the multilayer perceptron: in, are the normalized weights of the four mental accounts, corresponding to the economic account, social account, environmental account, and convenience account respectively; S4.

4. Combine the representations of the four mental accounts according to their weights to obtain a weighted comprehensive representation of the user: in, It is a comprehensive representation after weighted combination.

6. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 5 is characterized by: The S5 comprises the following steps: S5.1 builds an incentive generation network, which accepts the user's weighted comprehensive representation and the system budget constraint as input and outputs a four-dimensional incentive vector: Multi-layer fully connected network with non-linear activation functions and normalization layers: in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, LayerNorm is the normalization function; S5.

2. Based on the user weighted comprehensive representation and system budget constraints, the four-dimensional incentive vector is calculated through the incentive generation network: in, Represent the incentive values ​​for the economic account, social account, environmental account, and convenience account respectively. For the economic account, the incentive value is the amount of monetary reward; for the social account, the incentive value indicates the intensity of social recognition and honor incentives; for the environmental account, the incentive value indicates the degree of emphasis on environmental impact and contribution; for the convenience account, the incentive value indicates the intensity of convenience compensation and service improvement; S5.

3. Adjust the generated excitation vector: Among them, N is the number of users who need to be motivated at the current moment, is the incentive budget constraint of the system at time t.

7. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 6 is characterized by: The S6 comprises the following steps: S6.

1. Build a response prediction network that receives the user's four-dimensional incentive vector, mental account weights, and original feature vector as input and outputs the predicted probability of the user's response: Specifically, it is a multi-layer fully connected network with a nonlinear activation function and a Sigmoid output layer: in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function, Is the Sigmoid function, ensure that the output is in the range Inside, represents the probability of user response; S6.

2. Based on the user's four-dimensional incentive vector, mental account weights, and original feature vector, the response prediction network calculates the probability of the user's response: in, represents the probability prediction of user i responding to demand at time t; S6.

3. Construct a multi-objective optimization framework, including three loss functions: prediction loss function, budget constraint loss function, and weight diversity loss function; Prediction loss function: Use the two-compartment cross loss to measure the accuracy of the response prediction: in, It is a binary label of the user's actual response, 1 indicates response and 0 indicates no response; Budget constraint loss: Ensure that the total incentive does not exceed the system budget: in, is the weight coefficient of the budget constraint loss; Weight diversity loss, which encourages diversity in weight distribution: Among them, KL is KL divergence, which is a four-dimensional uniform distribution , is the weight of diversity loss S6.

4. Calculate the total loss and optimize the model parameters, combining the three loss functions into a total loss function: Update the model parameters using gradient descent: in, is the learning rate, and The total loss function is respectively and gradient.

8. The personalized incentive method for virtual power plants based on mental account representation learning according to claim 7 is characterized by: The following steps are also included: S7: Applying a user decision optimization mechanism, wherein S7 includes the following steps: S7.

1. Construct a decision complexity evaluation network to assess the decision burden imposed by incentive schemes on users: in, is the weight matrix of each layer of the network, is the corresponding bias vector, is the activation function; S7.

2. Based on the user's four-dimensional incentive vector, mental account weight, and original feature vector, the decision burden index is calculated using the decision complexity evaluation network: in, It represents the decision complexity caused by the incentive scheme to user i at time t. The larger the value, the higher the decision burden. S7.

3. Optimize the incentive presentation method based on decision complexity. When the decision complexity exceeds the preset threshold, simplify the incentive presentation method: Among them, Simplify is the incentive simplification function, is the complexity threshold parameter, and the simplification strategies include: when When you are working on something, keep the original thinking motivation unchanged; when , retain the incentives of the two mental accounts with the highest weights, and set the others to zero; when When , the value retains the weight to engage the incentive of the mental account, and the others are set to zero.

9. The personalized incentive method for a virtual power plant based on mental account representation learning according to claim 8 is characterized by: The following steps are also included: S8. Implementing a mental account transfer learning mechanism, wherein S8 specifically includes the following steps: S8.

1. Establish a meta-learning framework, using a model-independent learning method, to find initialization parameters through meta-training for a series of tasks that users have already thought about: in, is the task distribution, which means adopting users from the existing user group. is the loss function for a single user task, is the initialization parameter obtained by meta-learning; S8.

2. Based on the initialization parameters obtained by meta-learning and a small amount of new user interaction data, the model parameters for the new user are quickly adapted through gradient descent: in, is the transfer learning rate, which controls the degree of deviation of the new parameters from the meta-parameters, is the gradient of the meta-loss function with respect to the meta-parameters; S8.3 uses different learning rates for different mental accounts and dynamically adjusts the learning rate based on historical response data: in, is the learning rate of the jth mental account at time t, is the learning rate adjustment coefficient, is the gradient of the damage function with respect to the j-th mental account parameter.

10. A virtual power plant personalized incentive method based on mental account representation learning according to claim 9, characterized in that: The following steps are also included: S9. Deploy and operate the personalized incentive system for the virtual power plant, which specifically includes the following steps: S9.

1. System initialization, loading pre-trained model parameters, including self-calibration high-order Transformer model parameters, mental accounting weight calculation model parameters, incentive generation and response prediction model parameters, and decision complexity assessment model parameters; S9.

2. The system continuously receives user feature data and environmental context data, executes S1-S7, generates a personalized incentive plan for each user, and adjusts the incentive presentation method based on the decision complexity assessment results; S9.

3. Collect real-time user response data to the incentive scheme as labels for model training and perform model update steps regularly; S9.

4. Regularly evaluate system performance, including indicators such as user participation rate, incentive cost, and long-term participation stability. Adjust model hyperparameters and training strategies based on the evaluation results to continuously optimize system performance.

Citation Information

Cited By

  • Virtual power plant dynamic aggregation method and system based on Transform and AdaptMLP

    CN121457746A

  • Virtual power plant dynamic aggregation method and system based on transformer and adaptmlp

    CN121457746B