A smart scheduling method and system for power grid flexibility resources
Patent Information
- Application Number
- CN202610740129.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-27
- Publication Date
- 2026-09-22
AI Technical Summary
[0004]本申请提供了一种电网灵活性资源智能调度方法及系统,改善了现有技术中调度策略采用粗放式统一手段,难以在保护隐私前提下聚合多元个体形成精准互动,造成资源利用不充分与用户参与受限的技术问题
本申请技术方案通过提供的一种电网灵活性资源智能调度方法,首先,通过对目标用户的历史V2G行为数据进行时序事件序列构建、正弦位置编码注入和自注意力编码,获取行为表征向量,同时通过硬负例加权对比学习与多头解耦获取多维偏好向量,实现了用户行为时序规律与多维偏好的深度编码和结构化分解,为后续虚拟邻域构建及个性化决策提供了双重用户画像基础,解决了传统调度方法对用户行为建模粗糙、多种偏好因素难分离的问题。
Smart Images

Figure CN122797992A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of user recommendation technology, and in particular to a method and system for intelligent scheduling of power grid flexibility resources. Background Technology
[0002] V2G (Vehicle-to-Grid) user dispatch refers to a flexible resource management approach that guides electric vehicle users to participate in grid regulation through bidirectional charging and discharging facilities. By rationally arranging charging and discharging strategies, it enables a vast number of dispersed electric vehicle batteries to charge during off-peak hours and discharge during peak hours, thereby supporting peak shaving and valley filling and the absorption of renewable energy. Traditional dispatching methods typically involve the grid side or aggregators issuing standardized charging and discharging guidance instructions to all users based on a unified electricity price signal or fixed time period division. Users respond passively or not at all, and the dispatching strategy lacks awareness and adaptation to individual behavioral differences.
[0003] Traditional scheduling methods rely heavily on centralized decision-making, making it difficult to perceive and adapt to the significant differences in charging and discharging time preferences, site selection habits, price sensitivity, and range anxiety among a large number of dispersed users. The information transmission and mutual influence patterns existing in group behavior cannot be effectively captured and utilized. Historical behavioral data of individual users is stored on various local platforms and, due to privacy and security concerns, cannot be directly aggregated for centralized model training. This extensive scheduling strategy leads to a disconnect between recommended solutions and actual user behavior, limiting user participation and hindering the accurate aggregation and effective utilization of massive flexible resources by the power grid, thus restricting the efficient use of distributed resources by the power grid. Summary of the Invention
[0004] This application provides a method and system for intelligent scheduling of power grid flexibility resources, which improves the technical problems in the prior art where the scheduling strategy adopts a rough and uniform approach, making it difficult to aggregate multiple individuals to form precise interactions while protecting privacy, resulting in insufficient resource utilization and limited user participation.
[0005] This application discloses the following technical solution: Firstly, this application provides a method for intelligent scheduling of power grid flexibility resources, the method comprising: Acquire historical V2G behavior data of target users, and analyze to obtain behavioral representation vectors and multidimensional preference vectors of target users; Based on the behavior representation vector, a multidimensional virtual neighborhood of the target user is constructed, wherein each virtual neighborhood user in the multidimensional virtual neighborhood is labeled with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. Based on the homogeneous neighborhood in the multidimensional virtual neighborhood, a group decision model corresponding to the target user is established, and the group decision knowledge in the form of probability distribution generated by the group decision model on the preset query scenario set is obtained and sent to the target user's local area for distillation training to obtain a local behavior prediction model. The evidence input is constructed by combining the local behavior prediction model with the causal neighborhood, and the multidimensional preference vector is used as the condition input. The scheduling scheme decision is made by the scheduling decision-maker based on the conditional variational autoencoder to obtain a set of recommended scheduling schemes.
[0006] Secondly, this application provides a smart scheduling system for power grid flexibility resources, the system comprising: The behavior analysis module is used to acquire historical V2G behavior data of target users and analyze and obtain the behavioral representation vector and multidimensional preference vector of target users. The neighborhood construction module is used to construct a multidimensional virtual neighborhood of the target user based on the behavior representation vector. Each virtual neighborhood user in the multidimensional virtual neighborhood is marked with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. The group distillation module is used to establish a group decision model corresponding to the target user based on the homogeneous neighborhood in the multidimensional virtual neighborhood, and to obtain the group decision knowledge in the form of probability distribution generated by the group decision model on a preset query scenario set, and send it to the target user's local distillation training to obtain a local behavior prediction model. The scheme recommendation module is used to combine the local behavior prediction model with the causal neighborhood to construct evidence input, and use the multidimensional preference vector as condition input to make scheduling scheme decisions through a scheduling decision-maker based on conditional variational autoencoder to obtain a set of recommended scheduling schemes.
[0007] One or more technical solutions provided in this application have at least the following technical effects or advantages: The technical solution of this application provides an intelligent scheduling method for power grid flexibility resources. First, by constructing a time-series event sequence, injecting sinusoidal position encoding, and performing self-attention encoding on the historical V2G behavior data of the target user, a behavior representation vector is obtained. At the same time, a multi-dimensional preference vector is obtained through hard negative example weighted contrastive learning and multi-head decoupling. This achieves deep encoding and structured decomposition of the time-series patterns of user behavior and multi-dimensional preferences, providing a dual user profile foundation for subsequent virtual neighborhood construction and personalized decision-making. This solves the problems of coarse user behavior modeling and difficulty in separating multiple preference factors in traditional scheduling methods.
[0008] Furthermore, by screening homogeneous neighborhoods based on the cosine similarity of behavioral representation vectors, screening causal neighborhoods based on transfer entropy and permutation tests, and screening spatiotemporal neighborhoods based on spatiotemporal co-occurrence, a multidimensional virtual neighborhood with labeled neighborhood types is constructed. This enables comprehensive identification of user-related groups from three dimensions: behavioral similarity, causal transmission, and spatiotemporal collaboration. This solves the problems of isolated user modeling and the inability to discover and utilize the correlation patterns of group behavior in traditional scheduling.
[0009] Furthermore, a group decision-making model is established by weighting the federated aggregation according to the neighborhood type and the degree of association. After temperature scaling and differential privacy noise injection, secure group decision-making knowledge is generated. After being distributed, it is trained locally by distillation combined with preference alignment loss to obtain a local behavior prediction model. This realizes the secure differential transfer of group decision-making experience and the preservation of personalized preferences, and solves the problems of traditional federated learning and other methods ignoring the differences of participants, easily losing individual characteristics and having insufficient privacy protection.
[0010] Finally, by fusing the output of the local behavior prediction model with causal neighborhood signals to form evidence input, and using multidimensional preference vectors as conditions, a diverse set of candidate schemes is generated through a conditional variational autoencoder. After feasibility filtering, multidimensional scoring and difference constraint screening, a set of recommended scheduling schemes is output, realizing the generation and optimization of personalized scheduling schemes through multi-source evidence collaboration. This solves the problems of traditional recommendation methods lacking uncertainty modeling, ignoring the transmission of group behavior, and having insufficient feasibility and diversity of schemes.
[0011] In summary, the technical solution of this application achieves deep characterization and multidimensional preference decoupling of V2G user behavior data, construction and labeling of multiple types of virtual neighborhoods, privacy-preserving group experience federated distillation and local model training, multi-source evidence fusion and conditional variational autoencoder recommendation generation. Through temporal self-attention encoding and contrastive learning decoupling, transfer entropy causal inference and spatiotemporal co-occurrence analysis, weighted federated aggregation and differential privacy knowledge transfer, latent variable probability sampling and feasibility difference constraint screening, it effectively improves the technical problems in the prior art, such as the extensive nature of power grid flexibility resource scheduling strategies, difficulty in finely adapting to the personalized behavioral habits of dispersed users, and lack of privacy and security coordination means, resulting in low resource response efficiency and unstable user participation willingness. Attached Figure Description
[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1This is a schematic diagram of the overall process of a smart scheduling method for power grid flexibility resources provided in an embodiment of this application; Figure 2 A flowchart illustrating the construction of a multi-dimensional virtual neighborhood for an intelligent scheduling method of power grid flexibility resources provided in this application embodiment; Figure 3 This is a schematic diagram of the structure of a smart scheduling system for power grid flexibility resources provided in an embodiment of this application.
[0014] In the attached diagram, Figure 3 The components represented by each number are explained as follows: Behavior Analysis Module 11, Neighborhood Construction Module 12, Population Distillation Module 13, and Scheme Recommendation Module 14. Detailed Implementation
[0015] This application provides a method and system for intelligent scheduling of power grid flexibility resources, which addresses the technical problem that existing scheduling methods cannot fully adapt to and learn the personalized behavioral characteristics and group relationships of massive flexibility resources, resulting in recommended solutions deviating from actual needs and insufficient resource flexibility mining.
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that the numerical values in the embodiments are for illustrative purposes only and do not constitute a limitation on this application.
[0017] Example 1, as shown in the appendix Figure 1 As shown, this application provides a method for intelligent scheduling of power grid flexibility resources, the method comprising the following steps: S100: Obtain historical V2G behavior data of target users, and analyze and obtain the behavioral representation vector and multidimensional preference vector of target users; In this embodiment, V2G users refer to electric vehicle users participating in vehicle-to-grid interaction. They achieve bidirectional flow of electrical energy between the vehicle battery and the grid through bidirectional charging and discharging facilities, and are an important provider of grid flexibility resources. In grid flexibility resource scheduling, the behavior of dispersed V2G users exhibits significant individual differences. Different users have different preferences in terms of charging and discharging time, site selection, price sensitivity, and range anxiety. Traditional scheduling systems typically rely on simple statistics based on users' historical execution records, such as the charging frequency or site visits during different time periods over the past 30 days. This method can only capture shallow behavioral frequencies and lacks the ability to deeply model the temporal dependencies and contextual relationships in behavioral sequences. This results in a coarse representation of user behavior habits and a lack of structured decomposition, making it difficult to finely distinguish multi-dimensional factors such as time preferences, location preferences, price sensitivity, and convenience demands in behavioral habits, and thus failing to support subsequent accurate personalized scheduling decisions. Therefore, a user profiling method that can deeply encode behavioral sequence patterns and structurally decouple multi-dimensional preferences is needed. However, this method needs to solve the problem of modeling non-uniform time intervals between behavioral events and the problem of multiple preference factors being intertwined in behavioral data and difficult to separate independently.
[0018] Step S100 in the method provided in this application embodiment includes: Historical V2G behavior data of the target user is acquired and structured into behavior events containing timestamps, site numbers, behavior types, and state vectors, forming a behavior event sequence. For each behavior event, time feature embedding, site feature embedding, behavior type embedding, and state vector mapping are performed and concatenated to obtain an event embedding vector. Based on the real time interval between adjacent behavior events, a sinusoidal position code is generated and superimposed on the corresponding event embedding vector to obtain a location-aware event sequence. The location-aware event sequence is input into a pre-trained temporal encoder, and after multi-layer multi-head self-attention encoding, pooling is performed along the sequence dimension to obtain the behavior representation vector. Detailed explanation follows: In this embodiment, historical V2G behavior data refers to the complete behavior record generated by the target user's participation in vehicle-to-network interaction over a period of time, including operation logs from multiple dimensions such as charging behavior, discharging behavior, browsing behavior, and abandoned operations. The behavior event sequence refers to a structured event list formed by arranging historical V2G behavior data in ascending order of timestamps. Each event contains attribute fields such as time information, site information, behavior type, and environmental state. The event embedding vector refers to a fixed-length dense vector formed by concatenating the heterogeneous discrete and continuous attributes in the behavior event after embedding and mapping them separately, serving as a unified numerical representation that can be processed by the subsequent encoder. Sine wave position encoding refers to the sine function value calculated based on the actual time interval between adjacent behavior events, used to inject the time span information between events into the event embedding vector, enabling the encoder to perceive changes in the temporal density of the behavior. The location-aware event sequence refers to the event embedding vector sequence superimposed with sine wave position encoding, where each vector simultaneously carries behavioral semantic information and time interval information between events. The behavioral representation vector refers to a fixed-dimensional vector obtained by the temporal encoder after globally encoding the location-aware event sequence. It compresses the temporal patterns, habit preferences, and contextual dependency information in the user's complete behavioral sequence, serving as the user's overall behavioral profile.
[0019] In this step, firstly, in order to transform unstructured user logs into structured event sequences, it is necessary to obtain the target user's historical V2G behavior data, sort it in ascending order by timestamp, and uniformly represent each record as a behavior event containing timestamp, site number, behavior type, and state vector. This will yield a complete sequence of behavior events. For example, a user may have generated 327 behavior events in the past 90 days, with behavior types including six categories: start charging, end charging, start discharging, end discharging, browse sites, and cancel appointments.
[0020] Furthermore, in order to transform heterogeneous behavioral events into a unified numerical vector representation, it is necessary to perform embedding mapping on time features such as hours and days of the week, embedding mapping on site numbers, embedding mapping on behavior types, and linear mapping on state vectors such as initial SOC values and electricity price levels. Then, the various parts are concatenated to form an event embedding vector, which yields a 128-dimensional dense vector representation for each behavioral event. For example, the embedding vector of a charging behavior shows higher activation values in the dimensions of holiday periods and residential site.
[0021] Furthermore, in order to distinguish the time span information between events during the encoding process, a sinusoidal position code needs to be calculated based on the actual time interval between adjacent behavioral events. The longer the time interval, the lower the frequency component of the code. The sinusoidal position code is then superimposed element by element onto the event embedding vector to obtain a vector sequence carrying time density perception capability. For example, a sinusoidal code with a 4-hour interval between two charging behaviors makes the corresponding event respond more strongly in the low-frequency dimension, while an event with a 15-minute interval responds more strongly in the high-frequency dimension.
[0022] Finally, in order to extract a global representation of the user's overall behavioral habits, the location-aware event sequence needs to be input into a pre-trained temporal encoder. After six layers of multi-head self-attention encoding, average pooling is performed on the sequence dimension to obtain a 256-dimensional behavioral representation vector. This vector integrates the user's behavioral patterns at different times, at different stations, and under different states. For example, the representation vector of a weekday commuter shows typical patterns in the dimensions of weekday morning and evening peak hours and fixed stations.
[0023] For example, the pre-training process of the temporal encoder is as follows: Approximately 230 million historical behavioral event sequences from 500,000 V2G users were collected from multiple city public charging operation platforms (anonymized). 80% of the user data was randomly selected as the pre-training set, 10% as the validation set, and 10% as the test set. Masked behavior modeling was used as a self-supervised pre-training task. For each location-aware event sequence, 15% of the event embedding vector was randomly masked, and the masked positions were replaced with learnable mask vectors. The model architecture used a 6-layer Transformer encoder, each layer containing an 8-head self-attention mechanism, with a hidden dimension of 256 and a feedforward network dimension of 1024. The GELU activation function was used, and the maximum sequence length was set to 512. The encoder output was linearly projected onto a classification space of 6 behavior types and 3000 station numbers. The cross-entropy loss was calculated and summed to obtain the total pre-training loss. The optimizer used was AdamW, with an initial learning rate of 0.0001, weight decay of 0.01, a batch size of 64, and a maximum training epoch of 100 epochs. During training, a cosine annealing strategy was used for the learning rate. Training was terminated early if the loss on the validation set did not decrease for 10 consecutive rounds. Early termination was triggered after 78 training rounds. On the test set, the accuracy for behavior type prediction reached 0.892, and the accuracy for site number prediction reached 0.763, indicating that the encoder had fully learned the temporal patterns and contextual dependencies in the behavior sequences. After pre-training, the classification head was removed, but the encoder parameters were retained as a reusable temporal encoder in step S100.
[0024] Step S100 in the method provided in this application embodiment further includes: Based on the historical V2G behavior data, actual adoption behaviors are identified as positive examples, non-adoption behaviors as negative examples, and behaviors where the user's browsing time exceeds a preset time threshold but is not executed are identified as hard negative examples. Based on the positive examples, negative examples, and hard negative examples, contrastive sample pairs are constructed, and a preference encoder is trained using a contrastive learning method. The attention weight of the hard negative examples in the contrastive learning loss function is higher than that of the negative examples. Using multiple independent decoders, the preference vector output by the preference encoder is structurally decomposed into time preference subspaces, location preference subspaces, price sensitivity subspaces, convenience preference subspaces, and environmental awareness subspaces to obtain the multidimensional preference vector. The output of each subspace is a behavior probability. A detailed explanation follows: In this embodiment, hard negative examples refer to behavioral events where users spend a considerable amount of time browsing but ultimately abandon the action. For example, a user might stay on a website's details page for more than a preset time threshold of 180 seconds without performing a charging operation. These samples contain strong negative signals of user preference. The preference encoder refers to a neural network model trained through contrastive learning. Its input consists of the user's behavioral event sequence and contextual features, and its output is a comprehensive preference vector. The training objective is to make the representation of positive examples far removed from negative and hard negative examples in the feature space. The multidimensional preference vector refers to a set of preference subspace vectors obtained by decomposing the output of the preference encoder through multiple independent decoders. Each subspace corresponds to a behavioral preference dimension, including five dimensions: time preference, location preference, price sensitivity, convenience preference, and environmental awareness. Independent decoders refer to a set of parallel, lightweight, fully connected network heads connected to the output layer of the preference encoder. Each decoder maps a shared preference representation to the probability distribution space of its corresponding preference dimension, and the decoders are decoupled through orthogonal constraints.
[0025] In this step, firstly, in order to extract true preference signals from users' historical interaction behavior on the platform, it is necessary to identify three types of samples to construct contrastive learning data. Executed behaviors are positive examples, pushed but not executed behaviors are negative examples, and behaviors that are abandoned after browsing for more than 180 seconds are hard negative examples. Through contrastive learning, the preference encoder can clearly separate behaviors that users are interested in from those that they are indifferent to in the feature space. Hard negative examples are given higher attention weights in the loss function because they are rich in strong signals of users actively abandoning their actions.
[0026] Furthermore, in order to explicitly extract decision preferences of different dimensions from the preference vector, five independent decoding heads need to be set up to output time preferences such as the probability of choosing during different time periods during morning and evening peak hours, location preferences such as the probability of choosing charging stations in residential areas or commercial areas, price sensitivity such as the responsiveness to peak and off-peak electricity prices, convenience preferences such as the tolerance for charging and discharging time, and environmental awareness such as the green electricity preference coefficient. The output of each subspace is a probability value from 0 to 1, which can yield a multidimensional preference vector with separated dimensions. For example, a user's time preference subspace shows a charging probability of 0.78 during weekday evenings, a price sensitivity subspace shows a response probability of 0.85 during off-peak hours, and a convenience subspace shows an acceptable discharge depth of 20% to 50%.
[0027] For example, the preference encoder pre-training process is as follows: From the historical V2G behavior data of the target users and anonymized users on the same platform, 85,000 actual actions are extracted as positive examples, 123,000 actions that are pushed but not adopted are extracted as negative examples, and 12,000 actions that are viewed for more than 180 seconds but not executed are extracted as hard negative examples. For each positive example, one negative example and one hard negative example are randomly selected to construct a triplet comparison pair, for a total of 85,000 pairs. The preference encoder adopts a 6-layer fully connected network, with each layer having a dimension of 512, using the ReLU activation function and layer normalization. The input is a 256-dimensional vector of user behavior event sequence after embedding and average pooling, and the output is a 256-dimensional preference vector. The contrastive learning loss function adopts the weighted triplet loss, with the marginal value of the positive-negative example pair set to 0.5 and the marginal value of the positive-hard negative example pair set to 1.0. The loss weight of the hard negative example pair is 3 times that of the ordinary negative example pair. The optimizer used was Adam, with an initial learning rate of 0.0002, a batch size of 128, and a maximum training epoch of 50. During training, the learning rate decreased to 0.5 times its original value every 10 epochs. Early termination occurred if the loss on the validation set did not decrease for 5 consecutive epochs. Early termination was triggered after 42 epochs of training. On the validation set, the accuracy for distinguishing positive examples from negative examples reached 0.914, and the accuracy for distinguishing positive examples from hard negative examples reached 0.876, indicating that the preference encoder had sufficiently learned the differences between users' true preferences and abandonment behaviors. After training, the preference encoder parameters were frozen, and five independent decoders were connected for fine-tuning. Each decoder was a single-layer fully connected network, taking a 256-dimensional preference vector as input and outputting the probability distribution of the corresponding subspace.
[0028] For example, taking a target user as an example, historical V2G behavior data of the user over the past 90 days was collected, containing a total of 327 behavioral events, including 152 charging behaviors, 48 discharging behaviors, 89 browsing behaviors, 38 canceling appointment behaviors, and 23 hard negative examples of browsing duration exceeding 180 seconds. After time feature embedding, site feature embedding, behavior type embedding, and state vector mapping, each event obtains a 128-dimensional event embedding vector. This is then superimposed with sinusoidal position encoding based on the real time interval to form a location-aware event sequence. The sequence input is a pre-trained temporal encoder consisting of a 6-layer Transformer encoder and an average pooling layer, which outputs a 256-dimensional behavior representation vector. Dimensions 12 to 18 of this vector show high activation values during weekday periods, and dimension 67 shows significant activation at specific sites. At the preference encoder, a triplet comparison sample pair is constructed using execution behavior as a positive example, push notification not accepted as a negative example, and long browsing not accepted as a hard negative example. A 6-layer fully connected network preference encoder is trained to output a 256-dimensional preference vector, which is then decomposed into 32-dimensional subspace vectors by 5 independent decoders. The time preference subspace outputs the probability of charging during the evening period (0.78), the location preference subspace outputs the probability of choosing a residential site (0.64), the price sensitivity subspace outputs the probability of responding to off-peak electricity (0.85), the convenience preference subspace outputs the acceptable depth of discharge (20% to 50%), and the environmental awareness subspace outputs the green electricity preference coefficient (0.72).
[0029] In summary, this step constructs a behavior representation vector from historical V2G behavior data through event sequence construction, multi-feature embedding, sinusoidal positional encoding, and temporal encoder encoding. Simultaneously, it obtains a multi-dimensional preference vector through positive and negative example discrimination, hard negative example weighted contrastive learning, and multi-head decoupling. Compared with existing technologies, this step has the following advantages: First, by using a temporal encoder to perform self-attention modeling on behavior sequences with non-uniform time intervals, it preserves the long-short interval dependencies between behaviors, solving the problem that traditional statistical methods only focus on behavior frequency while ignoring temporal context. Second, through hard negative example weighted contrastive learning and orthogonal decoders, multi-dimensional preferences are explicitly decomposed into independent dimensions, solving the problem of multiple entangled and difficult-to-distinguish preference factors in traditional preference analysis, providing a dual user profile foundation for subsequent virtual neighborhood construction and scheduling recommendations.
[0030] S200: Based on the behavior representation vector, construct a multidimensional virtual neighborhood of the target user, wherein each virtual neighborhood user in the multidimensional virtual neighborhood is labeled with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. In this embodiment, the behavior of V2G users does not occur in isolation; charging habits, site selection, and responses to electricity prices among users exhibit behavioral propagation and mutual influence. Traditional methods model each user independently, such as relying solely on user history data for behavioral frequency statistics or time-series prediction, failing to identify other user groups with similar behavioral patterns, causal relationships, or spatiotemporal co-occurrence associations. This isolated individual modeling approach ignores the homogeneity of behavior and information propagation effects between groups, resulting in an understanding of user behavior limited to individual data while neglecting group interaction patterns. This makes it difficult to leverage group behavioral information to enhance the adaptability of subsequent scheduling decisions to individual user needs. Therefore, a neighborhood construction method is needed that can discover and label user-related groups from multiple dimensions. However, this method needs to address issues such as the quantitative standards for behavioral similarity, the directional determination of causal relationships, and the effective identification of spatiotemporal co-occurrence patterns.
[0031] The procedure for this step is as follows (see attached document). Figure 2 As shown.
[0032] Step S200 in the method provided in this application embodiment includes: Based on the cosine similarity between the behavioral representation vectors of multiple users, users exceeding a first preset threshold are selected as homogeneous virtual neighborhood users, and the homogeneous neighborhood is obtained. Based on the V2G participation rate time series of each user, transfer entropy is calculated to obtain the transfer entropy values between multiple user pairs. Significance is judged through a permutation test, and users whose transfer entropy values exceed a second preset threshold and whose significance level meets preset requirements are selected as causal virtual neighborhood users. A causal delay window is determined, and the causal neighborhood is obtained. Based on the spatiotemporal co-occurrence relationship between the behavioral events of each user, users whose co-occurrence frequency exceeds a third preset threshold within a preset time window and preset spatial distance are selected as spatiotemporal virtual neighborhood users, and the spatiotemporal neighborhood is obtained. The homogeneous neighborhood, the causal neighborhood, and the spatiotemporal neighborhood are combined to form the multidimensional virtual neighborhood. Detailed explanation follows: In this embodiment, a multidimensional virtual neighborhood refers to a set of virtual users constructed around the target user, consisting of various types of associated users. Each user is labeled with a different neighborhood type, providing a structured relationship network for subsequent group knowledge aggregation and causal signal extraction. Neighborhood type refers to a category label that distinguishes the nature of the association between users and the target user in the virtual neighborhood, including at least three types: homogeneous neighborhood, causal neighborhood, and spatiotemporal neighborhood. Homogeneous neighborhood refers to a set of virtual users whose V2G behavior patterns are highly similar to the target user, selected through cosine similarity between behavioral representation vectors, reflecting the homogeneous relationship characteristics of behavior. Causal neighborhood refers to a set of virtual users whose V2G participation behavior has a significant causal impact on the target user in the time series, selected through transfer entropy and permutation tests, reflecting the causal transmission relationship characteristics.
[0033] Spatiotemporal neighborhood refers to a set of virtual users whose V2G behavior frequently occurs within the same time period and similar spatial range. It is obtained through spatiotemporal co-occurrence relationships and reflects the spatiotemporal collaborative relationship characteristics. Transfer entropy is a causal measure based on information theory. It quantifies the predictive power of user A's behavior on user B's behavior by calculating the information gain of user A's V2G participation rate time series to the future values of user B's participation rate time series. A larger transfer entropy value indicates a stronger causal influence. Causal delay window refers to the time delay range of causal influence determined in the transfer entropy calculation, representing the typical time interval from a change in the source user's behavior to a response from the target user's behavior. Spatiotemporal co-occurrence relationship refers to the overlapping behavior patterns of two users simultaneously or sequentially appearing at the same or nearby charging stations within a preset time window and preset spatial distance.
[0034] In this step, firstly, in order to identify user groups with similar behavioral patterns, it is necessary to calculate the cosine similarity between the target user and other users' behavioral representation vectors. Users whose similarity exceeds the first preset threshold are selected to form a homogeneous neighborhood. The cosine similarity can reflect the directional consistency between two users in the high-dimensional behavioral feature space. For example, if the cosine similarity between the target user and a certain user is 0.92, which exceeds the first preset threshold of 0.85, then the user is included in the homogeneous neighborhood. A total of 127 homogeneous neighborhood users were selected.
[0035] Furthermore, to identify user groups that have a causal impact on the target user's behavior, it is necessary to calculate the transfer entropy based on the V2G participation rate time series of each user. The participation rate time series is counted hourly to determine whether users have engaged in V2G behavior within each hour. The transfer entropy order is set to 1. After calculating the transfer entropy values between user pairs, the significance is evaluated through 1000 random permutations tests. Users whose transfer entropy values exceed the second preset threshold of 0.15 and whose p-value is less than 0.01 are selected to form a causal neighborhood. At the same time, the time delay step corresponding to the peak value of the transfer entropy is extracted as the causal delay window. For example, if the transfer entropy value of user X to the target user is 0.23 and the p-value is 0.003, the causal delay window is 2 hours, indicating that user X's behavior will have an observable impact on the target user's behavior after about 2 hours. A total of 43 causal neighborhood users were selected.
[0036] Furthermore, in order to capture groups that exhibit spatiotemporal behavioral coordination with the target user, it is necessary to count the number of times each user co-occurs with the target user within a preset time window of 30 minutes and a preset spatial distance of 3km. Users whose co-occurrence count exceeds a third preset threshold of 5 times are selected to form a spatiotemporal neighborhood. Spatiotemporal co-occurrence reflects the synchronicity of behaviors between users in time and space. For example, if a user co-occurs with the target user charging at the same time and station 12 times in the past 30 days, exceeding the threshold of 5 times, they are included in the spatiotemporal neighborhood. A total of 68 spatiotemporal neighborhood users were selected.
[0037] Finally, in order to manage the three association dimensions in a unified manner, all virtual users selected from the homogeneous neighborhood, causal neighborhood, and spatiotemporal neighborhood need to be aggregated and labeled with their corresponding neighborhood types to obtain a multidimensional virtual neighborhood. It should be noted that the same user may belong to multiple neighborhood types at the same time. For example, if a user's behavioral representation vector is highly similar to that of the target user and also shows a significant causal influence in the transfer entropy test, then the user will be labeled as both a homogeneous neighborhood and a causal neighborhood. This overlapping relationship provides a flexible data foundation for weighting according to neighborhood type during subsequent group aggregation.
[0038] For example, taking a target user as an example, this user has already obtained a 256-dimensional behavioral representation vector in the aforementioned steps. Cosine similarity is calculated among 100,000 active V2G users in a similar user database, with a first preset threshold set to 0.85, resulting in 127 homogeneous neighboring users, with the highest similarity being 0.96. The V2G participation rate time series is statistically analyzed at hourly granularity, covering 1440 time points over the past 60 days. The transfer entropy order is set to 1. Transfer entropy is calculated for each of the 1000 candidate users, and 1000 permutation tests are performed. 43 causal neighboring users with transfer entropy values exceeding 0.15 and p-values less than 0.01 are selected, with a causal delay window distribution of 1 to 4 hours and a mode of 2 hours. Spatiotemporal co-occurrence analysis uses a preset time window of 30 minutes and a preset spatial distance of 3 km, counting the number of co-occurrences over the past 30 days, with a third preset threshold of 5, resulting in 68 spatiotemporal neighboring users. After merging and deduplicating the three types of neighborhoods, a total of 208 virtual neighborhood users were obtained, of which 26 users belonged to at least two types of neighborhoods, forming the target user's multidimensional virtual neighborhood.
[0039] In summary, this step constructs a multidimensional virtual neighborhood for the target user from three dimensions: behavioral similarity, causal influence, and spatiotemporal synergy. It utilizes cosine similarity to screen homogeneous neighborhoods, transfer entropy and permutation tests to screen causal neighborhoods, and spatiotemporal co-occurrence to screen spatiotemporal neighborhoods. Compared with existing technologies, this step has the following advantages: First, the multidimensional neighborhood discovery mechanism covers different levels of behavioral association. Homogeneous neighborhoods capture groups with similar behaviors, causal neighborhoods capture behavioral transmission relationships, and spatiotemporal neighborhoods capture synchronized behavioral patterns, overcoming the limitation of traditional methods that can only discover users associated with a single dimension. Second, the causal inference method combining transfer entropy and permutation tests provides a quantitative basis for judging the strength and significance of causal influence and gives a causal delay window, giving the introduction of subsequent causal signals a clear time range and solving the problem that traditional association analysis cannot distinguish between correlation and causal direction.
[0040] S300: Based on the homogeneous neighborhood in the multidimensional virtual neighborhood, establish a group decision model corresponding to the target user, and obtain the group decision knowledge in the form of probability distribution generated by the group decision model on a preset query scenario set, and send it to the target user's local area for distillation training to obtain a local behavior prediction model; In this embodiment, after constructing a multidimensional virtual neighborhood for the target user, a set of group users including homogeneous neighborhoods, causal neighborhoods, and spatiotemporal neighborhoods is obtained. However, how to effectively transfer the decision-making experience of homogeneous groups with similar behavioral patterns to the target user is a problem that urgently needs to be solved. Traditional federated learning methods typically aggregate the model parameters of all participants on a central server, using simple averaging or equal-weighted aggregation, ignoring the differences in the degree of behavioral similarity between different users and the target user. This results in the aggregated model being diluted by information from a large number of weakly correlated users, failing to accurately reflect the true decision-making patterns of the target user's group. Furthermore, directly distributing the aggregated model to the target user will result in the loss of the user's personalized behavioral characteristics, and there is a risk of privacy leakage during the transmission and use of group knowledge. Therefore, a method is needed that can aggregate group models differently according to neighborhood type and degree of correlation, and securely transfer group knowledge to individual models under strict privacy protection. This method needs to solve problems such as the reasonable quantification of aggregation weights, the privacy and security handling of group knowledge, and the preservation of individual differentiated preferences.
[0041] Step S300 in the method provided in this application embodiment includes: The process involves: obtaining local model parameters for the current status local decision-making models of each virtual neighboring user within the multidimensional virtual neighborhood; calculating aggregation weights based on the neighborhood type and correlation degree corresponding to each virtual neighboring user; weighting and aggregating the local model parameters with the target user's own model parameters to obtain a group decision-making model; performing inference and decision-making on a preset query scenario set based on the group decision-making model; temperature scaling and normalization of the inference output to generate group decision-making knowledge in the form of a probability distribution; injecting privacy-preserving noise from a Laplace distribution into the group decision-making knowledge, followed by truncation and normalization to obtain secure group decision-making knowledge; distributing the secure group decision-making knowledge to the target user's local environment; and using the target user's local behavior data to distill and train the target user's current status local decision-making model to obtain a local behavior prediction model. A detailed explanation follows: In this embodiment, the group decision-making model refers to the global model obtained by weighting and aggregating the local decision-making model parameters of each user within the same homogeneous neighborhood and the target user's own parameters, thus incorporating the shared knowledge of the behavioral group to which the target user belongs. The current local decision-making model refers to the behavior prediction model independently trained by each user based on their initial historical data. It has the same network structure but different parameters and is used to represent individual decision preferences. Local model parameters refer to the numerical set of all learnable weights and biases in the current local decision-making model, which is the basic unit of federated aggregation. Aggregation weights refer to the weighting coefficients assigned to each participating user based on the virtual neighborhood type and degree of association. To ensure normalization, the sum of the weights of all participants is usually 1; users with higher similarity or more direct association receive higher aggregation weights. Group decision-making knowledge refers to the probability distribution output by the group decision-making model after inference on a preset query scenario set, which is then sharpened and normalized by temperature scaling, reflecting the consensus-based behavioral choice tendency of the group in a specific scenario. Secure group decision-making knowledge refers to knowledge that has been modified by adding differential privacy noise with a Laplace distribution to the group decision-making knowledge and then truncating and normalizing it to maintain its probability distribution form, thus ensuring that individual user participation information is not leaked during transmission and use.
[0042] First, in order to aggregate collective intelligence, it is necessary to obtain the current local decision model parameters of each virtual neighboring user in the same neighborhood. These models have been independently trained locally by each user, and the parameters are transmitted to the aggregation node through the network. For example, among 127 users in the same neighborhood, each user uploads all the weights and biases of their current local decision model, which is about 12,000 floating-point parameters.
[0043] Furthermore, to avoid the dilution of group knowledge caused by equal-weighted aggregation, aggregation weights need to be calculated based on neighborhood type and degree of association. Different neighborhood weight coefficients are set for virtual users with different neighborhood types: 1.0 for homogeneous neighborhood users, 0.4 for spatiotemporal neighborhood users, and 0.2 for causal neighborhood users. Within each neighborhood, the value normalized by the cosine similarity of the behavioral representation vector is used as the user's internal weight. The preset retention ratio of the target user's own parameters is set to 0.2, and the remaining 0.8 is allocated according to the effective weights of each virtual neighborhood user. This strategy ensures that homogeneous neighborhoods contribute the most to aggregation, spatiotemporal neighborhoods play a supporting role due to behavioral synchronicity, and causal neighborhoods contribute less due to time delay, thus achieving differentiated aggregation. For example, among 208 virtual neighborhood users, 127 homogeneous neighborhood users, 68 spatiotemporal neighborhood users, and 43 causal neighborhood users are assigned aggregate weights based on their respective coefficients and internal similarity. The most active homogeneous high-similarity users have a weight of up to 0.012, spatiotemporal neighborhood users have a weight of about 0.005, and causal neighborhood users have a weight of about 0.002. This yields comprehensive weighted group decision model parameters.
[0044] Furthermore, in order to transform the continuous decision-making capability of the group decision-making model into a transferable knowledge form, reasoning needs to be performed on a pre-defined query scenario set. The query scenario set covers 500 typical scenarios with different combinations of time, site, SOC, and electricity price. The output logits of the group decision-making model in each scenario are scaled by the temperature parameter T=2.0 and then normalized by Softmax to obtain the probability distribution of three types of behavior: charging, discharging, and not participating. As group decision-making knowledge, temperature scaling can smooth the probability distribution and preserve the relative size relationship between categories.
[0045] Furthermore, to protect the privacy of all participants, privacy-preserving noise with a Laplace distribution needs to be injected into the group decision knowledge. The noise scale parameter is calculated according to the differential privacy budget ε=0.5. After adding independent Laplace noise to each probability value, it is truncated, with probability values less than 0 set to 0 and those greater than 1 set to 1, and then normalized again to obtain secure group decision knowledge. This ensures that even if an attacker obtains the knowledge, they cannot infer the original model parameters of any specific user.
[0046] Finally, in order to transfer the collective knowledge to the target user's local model, the knowledge of safe collective decision-making needs to be distributed to the target user's local system. Combined with the target user's own local behavioral data, a knowledge distillation method is used to train the local behavior prediction model. This model can not only fit the local real behavior, but also approximate the soft decision distribution of the group in the same scenario. At the same time, personalized preference features are maintained through preference alignment loss, so that a local behavior prediction model that combines collective wisdom and individual differences can be obtained.
[0047] For example, the pre-training process of the current local decision-making model is as follows: Each V2G user, from the moment they connect to the system, undergoes local pre-training based on their historical V2G behavior data from the previous 30 days. The model structure is uniformly a 3-layer fully connected network. The input layer has 7 neurons corresponding to 7 dimensions of scene features: time features, site type, current SOC, electricity price level, weather conditions, vehicle type, and historical behavior frequency. The hidden layer dimensions are 64 and 32, respectively. ReLU activation function and Dropout regularization are used, with a Dropout ratio of 0.2. The output layer has 3 neurons that are normalized by Softmax and output the probability distributions of charging, discharging, and not participating in the three behavior categories. The loss function is cross-entropy loss, the optimizer is Adam, the initial learning rate is 0.001, the batch size is 32, and the maximum training epochs are 80. If the loss on the validation set does not decrease for 8 consecutive epochs during training, the training is terminated early. Taking the target user as an example, the model was trained based on 98 behavioral events from the previous 30 days. After 47 rounds of training, early termination was triggered. The classification accuracy on the validation set was 0.847, and the model had approximately 12,000 parameters, which were stored as a local model parameter file. The local decision-making models for the current status of each user in the virtual neighborhood were all independently trained with the same structure and process. The model structures were the same, but the parameters were different, which provided a homogeneous parameter basis for subsequent federated aggregation.
[0048] Step S300 in the method provided in this application embodiment further includes: The loss function for distillation training includes a weighted sum of local behavior prediction loss, distillation loss, and preference alignment loss, wherein: the local behavior prediction loss is the cross-entropy loss of the local behavior prediction model on local behavior data; the distillation loss is defined as the KL divergence between the output of the local behavior prediction model on a preset query scenario set and the safety group decision knowledge, and the weight of the distillation loss decreases linearly with the increase of training rounds; the preference alignment loss is defined as the cosine similarity loss between the intermediate hidden layer representation of the local behavior prediction model after linear projection and the multidimensional preference vector.
[0049] In this embodiment, the local behavior prediction loss refers to the cross-entropy loss of the local behavior prediction model on the target user's local real behavior labels, used to maintain the model's ability to fit the user's actual behavior. The distillation loss refers to the KL divergence between the local model output and the safety group decision-making knowledge, used to transfer group decision-making knowledge to the local model, making the local model's output probability distribution close to the group consensus in the same scenario. The preference alignment loss refers to the cosine similarity loss between the intermediate hidden layer representation of the local model after passing through a learnable linear projection matrix and the multidimensional preference vector obtained in the aforementioned steps, used to constrain the model to retain the user's personalized preference direction in the representation space. The intermediate hidden layer selects the output of the first hidden layer of the local behavior prediction model, with a dimension of 64, which is linearly projected to 256 dimensions before calculating the cosine distance with the preference vector.
[0050] To balance the three objectives of group knowledge transfer, local data fitting, and personalized preference preservation during distillation training, the three loss components are weighted and summed to obtain the total loss. Specifically, the weight of the local behavior prediction loss is fixed at 1.0; the weight of the distillation loss is initially set to 0.5, and then linearly decays from 0.5 to 0.05 with each training epoch to prevent early dominance of group knowledge from being completely ignored later; and the weight of the preference alignment loss is fixed at 0.3. After training, the local behavior prediction model can predict possible user behaviors based on scene features, absorb consensus knowledge from the group in the same scene, and ultimately maintain consistency between the hidden layer representation and the user's multidimensional preferences.
[0051] For example, taking the target user as an example, its local distillation training uses 327 of its own historical behavioral events as local data, and calculates group decision-making knowledge as the distillation target on 500 preset query scenarios. The local behavior prediction model structure is the same as the current local decision-making model, which is a 3-layer fully connected network. The initial parameters are taken from the target user's own current local decision-making model. Training uses the AdamW optimizer with a learning rate of 0.0005, a batch size of 16, and 30 training epochs. In the total loss, the local behavior prediction loss weight is 1.0, the distillation loss weight decreases linearly from 0.5 to 0.05, the preference alignment loss weight is 0.3, and the KL divergence temperature parameter in the distillation loss is set to 3.0. After training, the behavior prediction accuracy of the local behavior prediction model on the validation set reaches 0.892, which is 4.5 percentage points higher than before distillation. At the same time, the cosine similarity between its hidden layer representation and the multidimensional preference vector remains at 0.86, indicating that personalized preferences are not lost due to the transfer of group knowledge.
[0052] In summary, this step establishes a group decision-making model by weighting and aggregating parameters of similar neighborhood user models based on neighborhood relevance. Secure group decision-making knowledge is generated through temperature scaling and differential privacy processing. This knowledge is then trained locally on the target user's machine using local data and preference alignment loss. Compared to existing technologies, this step offers the following advantages: First, the weighted aggregation strategy guided by neighborhood type and relevance avoids information dilution caused by treating all participants equally in traditional federated averages, enabling the group model to more accurately reflect the decision-making patterns of the target user's behavioral group. Second, differential privacy noise injection and truncation normalization provide quantifiable privacy protection while achieving group knowledge transfer, mitigating the privacy leakage risk in traditional federated learning during knowledge transmission. Third, the introduction of preference alignment loss constrains the distillation process, preserving personalized user preference representations while absorbing collective wisdom, thus avoiding the problem of assimilating individual differences into single-group knowledge.
[0053] S400: Combine the local behavior prediction model with the causal neighborhood to construct evidence input, and use the multidimensional preference vector as conditional input. Then, use a scheduling decision-maker based on conditional variational autoencoder to make scheduling scheme decisions and obtain a set of recommended scheduling schemes.
[0054] In this embodiment, after group knowledge distillation and local training, the target user possesses the ability to predict behavior by incorporating collective wisdom. Simultaneously, the causal neighborhood within the multidimensional virtual neighborhood provides causal delay windows and relationship strength information regarding the transmission influence of other users' behaviors on the target user. However, effectively fusing the behavior prediction results with the preceding behavioral signals already occurring in the causal neighborhood, and generating specific and executable scheduling recommendation schemes under the guidance of the user's multidimensional preferences, remains a pressing issue. Traditional recommendation methods typically generate deterministic suggestions based solely on a single prediction model, neglecting the uncertainty inherent in the prediction probability itself, and failing to incorporate recent behaviors of others in the causal neighborhood as auxiliary evidence. This results in insufficient response to environmental changes and group behavioral trends in the recommendation results. Furthermore, the generated candidate schemes often lack systematic consideration of physical feasibility, user preference matching, and scheme diversity, making it difficult to meet the dual requirements of grid flexibility scheduling for response accuracy and user participation willingness. Therefore, a generative scheduling decision-making method capable of synergistically fusing individual predictions, causal signals, and preference conditions is needed to address the problems of incomplete information from a single evidence source, the disconnect between recommended schemes and actual user preferences, and the homogenization of candidate schemes.
[0055] Step S400 in the method provided in this application embodiment includes: An encoder for a conditional variational autoencoder is constructed, and the sample evidence input is concatenated with the sample condition input, which is randomly set to zero according to a preset probability. This concatenation is defined as the encoder input. The sample mean vector and sample log-variance vector of the latent space's sample posterior distribution are defined as the encoder output. Based on a reparameterized sampling strategy, the sample mean vector is element-wise added to the sample standard deviation and standard normal distribution random noise calculated based on the sample log-variance vector to obtain the sample latent variables. A decoder for the conditional variational autoencoder is constructed, and the sample latent variables are concatenated with the sample condition input. This decoder is defined as the decoder. The encoder input, after being mapped through a shared hidden layer, branches into multiple output heads, wherein the multiple output heads respectively output the probability distribution of recommendation time, the probability distribution of recommendation site, the probability distribution of recommendation behavior type, and the target value of recommendation state of charge; a conditional prior distribution driven by a linear mapping of the sample multidimensional preference vector is defined; the training loss is calculated, and the parameters of the conditional variational autoencoder, the shared hidden layer, and the multiple output heads are updated to obtain the scheduling decision-maker; wherein, the training loss is at least a weighted sum of reconstruction loss, KL divergence loss, hidden layer distillation loss, and conditional consistency loss. A detailed explanation follows: In this embodiment, the conditional variational autoencoder refers to a generative probabilistic model guided by conditional information, consisting of an encoder, a latent space, and a decoder. The encoder maps the input to a latent variable probability distribution, and the decoder samples from the latent variables and combines them with conditions to generate the output, used to generate diverse recommendation schemes that conform to the distribution law under given conditions. Evidence input refers to a fixed-dimensional vector formed by concatenating the output probability distribution of the local behavior prediction model in the current scenario and causal neighborhood signals, used to provide a comprehensive information basis for scheduling decisions. Sample latent variables refer to low-dimensional random vectors obtained by reparameterizing the posterior distribution of the encoder output, carrying a compressed representation of the input information and introducing randomness to make the generated schemes diverse. The reparameterization sampling strategy refers to using a differentiable method in the sampling process to combine random noise with the mean and standard deviation of the distribution, allowing gradients to propagate backward. The conditional prior distribution refers to the latent variable prior distribution defined by linear mapping with a multidimensional preference vector as input; during training, the encoder's posterior distribution is approximated to this prior distribution. Shared hidden layers refer to several fully connected network layers shared by all output heads in the decoder, used to extract shared feature representations that are highly relevant to the recommendation task.
[0056] First, in order to compress multi-source information into a unified latent variable representation, the sample evidence input and the sample condition input after random zeroing need to be concatenated and then input into the encoder. The encoder outputs the mean and log-variance of the latent space. Then, the sample latent variables are obtained through reparameterized sampling, so that the latent variables simultaneously encode evidence information and condition preferences. The random zeroing operation enhances the robustness of the model when the condition part is missing.
[0057] Furthermore, in order to generate a multi-dimensional recommendation scheme under given conditions and latent variables, the latent variables and the original condition inputs need to be concatenated and mapped through a shared latent layer, and then branched into four output heads to predict recommendation time, recommendation site, recommendation behavior type and recommendation SOC target value respectively. The shared latent layer can extract common features across tasks, and multiple output heads ensure the independent expression of each recommendation dimension.
[0058] Furthermore, in order to establish the correct generative prior during the training phase, a conditional prior distribution driven by a linear mapping of multidimensional preference vectors needs to be defined. The encoder posterior distribution is constrained to the vicinity of this prior through KL divergence loss, so that latent variables can be directly sampled from the conditional prior for generation during the inference phase.
[0059] Finally, to train an effective scheduling decision-maker, a weighted total loss consisting of reconstruction loss, KL divergence loss, hidden layer distillation loss, and conditional consistency loss needs to be calculated. These four losses respectively guarantee output accuracy, latent space regularity, knowledge connection with the behavior prediction model, and semantic consistency with user preferences. The hidden layer distillation loss refers to the mean squared error between the 128-dimensional vector of the decoder's shared hidden layer output and the 128-dimensional vector of the first hidden layer of the local behavior prediction model after linear projection. It aims to transfer the behavior discrimination knowledge inherent in the behavior prediction model to the intermediate representation layer of the scheduling decision-maker, ensuring feature-level consistency between the generation process and behavior prediction.
[0060] For example, the pre-training process of the scheduling decision-maker is as follows: 120,000 samples are collected from 5000 users. Each sample contains evidence input, conditional input, and actual execution labels. These samples are divided into training, validation, and test sets in an 8:1:1 ratio. The encoder is a 4-layer fully connected network with a 265-dimensional input and a 16-dimensional latent space. The decoder shares 3 hidden layers with a 272-dimensional input and 4 output heads. The conditional prior follows a 16-dimensional Gaussian distribution, with the mean linearly mapped from the preference vector. The training loss includes reconstruction loss, KL divergence, hidden layer distillation loss, and conditional consistency loss, with weights of 1.0, 0.1, 0.05, and 0.05, respectively. The optimizer, Adam, has a learning rate of 0.0003, a batch size of 64, a maximum of 120 epochs, 10 early stopping epochs, and a 0.1 probability of randomly setting the conditional input to zero during training. The training was stopped early after 98 rounds. The test set showed a recommendation time accuracy of 0.834, a site recall of 0.791, a behavior type accuracy of 0.892, and a SOC target value MAE of 0.067.
[0061] Step S400 in the method provided in this application embodiment further includes: The process involves: obtaining the output probability distribution of the local behavior prediction model in the current decision-making scenario; obtaining the recent V2G behavior statistical features of the virtual neighborhood users corresponding to the causal neighborhood within the causal delay window to generate a causal neighborhood signal; concatenating the causal neighborhood signal with the output probability distribution to form evidence input; using the multidimensional preference vector as conditional input, inputting it along with the evidence input into the encoder of the scheduling decision-maker to generate a probability distribution in the latent space, and sampling latent variables from the probability distribution; and the decoder in the scheduling decision-maker, based on the latent variables and the conditional input, decoding and generating V2G recommendation schemes through a shared hidden layer and multiple output heads, outputting a set of recommended scheduling schemes. A detailed explanation follows: In this embodiment, the causal neighborhood signal refers to a vector composed of V2G behavior statistics of each virtual neighborhood user within the causal delay window, such as the number of charging events, average SOC change, and number of discharging events in the past 2 hours, used to capture the dynamic behavior of groups that have a causal impact on the target user. The current decision-making scenario refers to the specific scenario in which a recommendation solution needs to be generated, including environmental information such as current time, site location, vehicle SOC, and electricity price level.
[0062] In this step, firstly, in order to capture the recent changing trends of group behavior, it is necessary to obtain the V2G behavior statistical characteristics of causal neighborhood users within the causal delay window and generate a causal neighborhood signal, which reflects the group state that may have behavioral transmission to the target user.
[0063] Furthermore, in order to integrate individual predictions with group signals, the output probability distribution of the local behavior prediction model needs to be spliced with the causal neighborhood signal to form evidence input, so that the scheduling decision-maker can simultaneously refer to individual prediction tendencies and group dynamics.
[0064] Finally, in order to generate personalized recommendation schemes, the evidence input and multidimensional preference vector are input into the encoder of the scheduling decision-maker. After sampling the latent variables, the decoder generates V2G recommendation schemes, thereby outputting a set of recommended scheduling schemes.
[0065] Step S400 in the method provided in this application embodiment further includes: Multiple latent variables are decoded using a decoder to generate multiple candidate scheduling schemes. Feasibility filtering is performed on each candidate scheduling scheme, and a multi-dimensional score is applied to the candidate scheduling schemes that pass the feasibility filtering. The multi-dimensional score includes at least the expected economic benefit of normalized weighted fusion and the contribution value to grid flexibility. Based on the multi-dimensional score results, multiple candidate scheduling schemes are serialized in descending order. A predetermined number of candidate scheduling schemes are selected from top to bottom of the sequence, constrained by the difference between any two recommended schemes not being lower than a preset difference threshold, to form the recommended scheduling scheme set. Detailed explanation follows: In this embodiment, feasibility filtering refers to physical and rule-based constraint checks on candidate scheduling schemes, eliminating schemes that do not meet basic conditions such as SOC upper and lower limits, site capacity limitations, and user travel conflicts, ensuring that the recommended schemes are practically feasible. Multidimensional scoring refers to quantitatively evaluating schemes that pass feasibility filtering from multiple perspectives, including at least two dimensions: expected economic benefits and grid flexibility contribution. These two dimensions are normalized and weighted to reflect the comprehensive value of the scheme. The difference threshold refers to the minimum difference constraint set when selecting the final recommended scheme. When two schemes are too similar in time, site, or behavior type, only the scheme with the higher score is retained to ensure the diversity of the recommendation set.
[0066] In this step, firstly, in order to improve the diversity of the schemes, the latent variables need to be sampled multiple times, and multiple candidate scheduling schemes are generated by the decoder. Each sampling introduces randomness so that the generated schemes vary within a reasonable range.
[0067] Furthermore, to ensure the feasibility of the solution, the candidate solutions need to be filtered for feasibility, eliminating solutions that violate physical constraints or are not feasible for the user, and retaining the candidate solutions that pass the filter.
[0068] Furthermore, in order to evaluate the value of the schemes from multiple perspectives, it is necessary to calculate a normalized weighted fusion score for the expected economic benefits and grid flexibility contribution value of the remaining schemes. The weights of the two can be flexibly adjusted according to the actual dispatching tasks.
[0069] The difference between candidate solutions is measured using a multi-dimensional weighted distance, including three dimensions: time difference, site difference, and behavior type difference. Time difference is defined as the ratio of the absolute difference between the recommended times of two solutions to 24 hours; site spatial distance is defined as the ratio of the Euclidean distance between the latitude and longitude of two recommended sites to the maximum distance between charging stations within the city; behavior type difference is set to 0 when the recommended behaviors are the same and 1 when they are different. The weights of the three dimensions are set to 0.4, 0.3, and 0.3, respectively, with the overall difference value ranging from 0 to 1. A preset difference threshold of 0.3 is used to ensure that the selected solutions have sufficient differentiation in terms of time, location, or behavior type, avoiding overly similar solutions in the recommendation set.
[0070] Finally, in order to output diverse and high-quality recommendations, the solutions need to be sorted in descending order of scores, and then a predetermined number of solutions need to be selected from the top down with the difference threshold as a constraint to form the final set of recommendation scheduling solutions.
[0071] For example, taking the target user as an example, the current decision-making scenario is 8:00 PM on a weekday, the vehicle's SOC is 45%, the user is located at a residential charging station, and the off-peak electricity price is 0.35 yuan / kWh. The local behavior prediction model outputs a charging probability of 0.72, a discharging probability of 0.06, and a non-participation probability of 0.22. Within the causal neighborhood, 43 users have a total of 12 charging events and 3 discharging events within a causal delay window of 2 hours, with an average charging SOC increment of 28%, generating a 3-dimensional raw statistic [12, 3, 28%). To unify the dimensions, the number of charging events and the number of discharging events are normalized by dividing by the total number of users in the causal neighborhood, 43, respectively, resulting in a charging frequency ratio of 0.279 and a discharging frequency ratio of 0.070; the SOC increment ratio of 28% is converted to 0.28. The normalized causal neighborhood signal is [0.279, 0.070, 0.28], which is concatenated with the local prediction probability [0.72, 0.06, 0.22] to form a 6-dimensional evidence vector [0.72, 0.06, 0.22, 0.279, 0.070, 0.28].
[0072] The multidimensional preference vector uses time preference (evening 0.78), location preference (residential area 0.64), price sensitivity 0.85, convenience preference 0.60, and environmental awareness 0.72 as 256-dimensional conditional inputs. The encoder encodes the evidence and conditions to generate 16-dimensional mean and variance. Ten latent variables are extracted through 10 repeated samplings, and after decoding, 10 candidate scheduling schemes are generated. After feasibility filtering to remove two schemes exceeding the SOC limit, 8 remain. In the multidimensional scoring, economic benefit and grid contribution are weighted at 0.5. The scores of each scheme are calculated and sorted in descending order, with a difference threshold set to 0.3. The top 5 schemes are selected as the recommended scheduling scheme set, covering recommended charging time (22:00-23:00), recommended station number #0321 (residential area station), and recommended charging to 80% SOC.
[0073] In summary, this step constructs a conditional variational autoencoder (CVAE) as the scheduling decision-maker, fusing local behavior prediction output with causal neighborhood signals as evidence input. Using multidimensional preference vectors as conditions, it generates diverse V2G scheduling recommendation schemes, which are then filtered through feasibility filtering, multidimensional scoring, and difference constraints to obtain the final recommendation set. Compared with existing technologies, this step has the following advantages: First, by fusing individual predictions and group causal signals through the latent variable sampling mechanism of CVAE, it models the uncertainty of recommendations in a probabilistic manner, solving the problem of poor adaptability of traditional deterministic recommendations to fluctuations in user behavior. Second, by introducing causal neighborhood signals as evidence, it dynamically supplements the deficiencies of individual predictions with group behavior within the causal delay window during scheme generation, improving the response speed of recommendations to recent group behavior trends. Third, through a combination of feasibility filtering, economic and grid dual-dimensional scoring, and difference constraints, it ensures the executability, comprehensive value, and diversity of the recommended schemes, providing decision support for grid flexibility resource scheduling that takes into account both individual preferences and global benefits.
[0074] The embodiments of this application, through the specific implementation methods described above, achieve the following technical effects: This application proposes an intelligent scheduling method for power grid flexibility resources. First, historical V2G behavior data of target users is acquired and analyzed to obtain behavior representation vectors and multidimensional preference vectors. Then, a multidimensional virtual neighborhood, including homogeneous neighborhoods, causal neighborhoods, and spatiotemporal neighborhoods, is constructed based on the behavior representation vectors. Next, a group decision-making model is established using the homogeneous neighborhood, and after differential privacy protection, the group decision-making knowledge is distributed and trained locally to obtain a local behavior prediction model. Finally, the evidence input is constructed by combining the local behavior prediction model and causal neighborhood signals, using the multidimensional preference vectors as conditions, and a recommended scheduling scheme is generated through a conditional variational autoencoder. This method achieves accurate generation of personalized scheduling schemes for V2G users and secure transfer of group experience through multidimensional neighborhood collaboration and privacy-preserving federated distillation.
[0075] Example 2, as shown in the appendix Figure 3 As shown, based on the inventive concept of the intelligent scheduling method for power grid flexibility resources provided in Embodiment 1, this application also provides an intelligent scheduling system for power grid flexibility resources, specifically including: Behavior analysis module 11 is used to acquire historical V2G behavior data of target users and analyze and acquire the behavioral representation vector and multidimensional preference vector of target users; The neighborhood construction module 12 is used to construct a multidimensional virtual neighborhood of the target user based on the behavior representation vector, wherein each virtual neighborhood user in the multidimensional virtual neighborhood is marked with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. The group distillation module 13 is used to establish a group decision model corresponding to the target user based on the homogeneous neighborhood in the multidimensional virtual neighborhood, and to obtain the group decision knowledge in the form of probability distribution generated by the group decision model on a preset query scenario set, and send it to the target user's local distillation training to obtain a local behavior prediction model. The scheme recommendation module 14 is used to combine the local behavior prediction model and the causal neighborhood to construct evidence input, and use the multidimensional preference vector as condition input to make scheduling scheme decisions through a scheduling decision-maker based on conditional variational autoencoder to obtain a set of recommended scheduling schemes.
[0076] In one embodiment, the behavior analysis module 11 is further configured to: acquire historical V2G behavior data of the target user and structure it into behavior events containing timestamps, site numbers, behavior types, and state vectors, forming a behavior event sequence; perform time feature embedding, site feature embedding, behavior type embedding, and state vector mapping on each behavior event and concatenate them to obtain an event embedding vector; generate a sinusoidal position code based on the real time interval between adjacent behavior events, and superimpose the sinusoidal position code onto the corresponding event embedding vector to obtain a location-aware event sequence; input the location-aware event sequence into a pre-trained temporal encoder, perform multi-layer multi-head self-attention encoding, and then pool along the sequence dimension to obtain the behavior representation vector.
[0077] Furthermore, the behavior analysis module 11 is also used to: based on the historical V2G behavior data, identify actual adoption behavior as positive examples, identify non-adoption behavior as negative examples, and identify user browsing time exceeding a preset time threshold and not executed as hard negative examples; construct contrast sample pairs based on the positive examples, the negative examples, and the hard negative examples, and train the preference encoder through a contrastive learning method, wherein the attention weight of the hard negative examples in the loss function of the contrastive learning is higher than the attention weight of the negative examples; and through multiple independent decoding heads, structurally decompose the preference vector output by the preference encoder into a time preference subspace, a location preference subspace, a price sensitivity subspace, a convenience preference subspace, and an environmental awareness subspace to obtain the multidimensional preference vector, wherein the output of each subspace is a behavior probability.
[0078] In one embodiment, the neighborhood construction module 12 is further configured to: filter users whose behavior representation vectors exceed a first preset threshold as homogeneous virtual neighborhood users based on the cosine similarity between the behavioral representation vectors of multiple users, and obtain the homogeneous neighborhood; calculate the transfer entropy based on the V2G participation rate time series of each user, obtain the transfer entropy value between multiple user pairs, and make a significance judgment through permutation test, filter users whose transfer entropy value exceeds a second preset threshold and whose significance level meets preset requirements as causal virtual neighborhood users, and determine the causal delay window, and obtain the causal neighborhood; filter users whose co-occurrence frequency exceeds a third preset threshold within a preset time window and a preset spatial distance based on the spatiotemporal co-occurrence relationship between the behavioral events of each user, and obtain the spatiotemporal virtual neighborhood users; and combine the homogeneous neighborhood, the causal neighborhood, and the spatiotemporal neighborhood to form the multidimensional virtual neighborhood.
[0079] In one embodiment, the group distillation module 13 is further configured to: obtain the local model parameters of the current status local decision-making model of each virtual neighborhood user in the multidimensional virtual neighborhood; calculate the aggregation weight based on the neighborhood type and correlation degree corresponding to each virtual neighborhood user, and perform weighted aggregation of each local model parameter and the target user's own model parameters to obtain a group decision-making model; perform inference decision-making on a preset query scenario set based on the group decision-making model, perform temperature scaling on the inference output, and normalize it to generate group decision-making knowledge in the form of a probability distribution; inject privacy-preserving noise of Laplace distribution into the group decision-making knowledge, and perform truncation and normalization post-processing to obtain secure group decision-making knowledge; distribute the secure group decision-making knowledge to the target user's local environment, and combine it with the target user's local behavior data to perform distillation training on the target user's own current status local decision-making model to obtain a local behavior prediction model.
[0080] Furthermore, the group distillation module 13 is also used for: the loss function of the distillation training includes a weighted sum of local behavior prediction loss, distillation loss and preference alignment loss, wherein: the local behavior prediction loss is the cross-entropy loss of the local behavior prediction model on local behavior data; the distillation loss is defined as the KL divergence between the output of the local behavior prediction model on a preset query scenario set and the safe group decision knowledge, and the weight of the distillation loss decreases linearly with the increase of training rounds; the preference alignment loss is defined as the cosine similarity loss between the intermediate hidden layer representation of the local behavior prediction model after linear projection and the multidimensional preference vector.
[0081] In one embodiment, the scheme recommendation module 14 is further configured to: construct an encoder for a conditional variational autoencoder, and concatenate the sample evidence input with the sample condition input randomly set to zero according to a preset probability, defining it as the encoder input; define the sample mean vector and sample log variance vector of the sample posterior distribution in the latent space as the encoder output; based on a reparameterized sampling strategy, add the sample mean vector to the sample standard deviation and standard normal distribution random noise calculated based on the sample log variance vector element-wise to obtain the sample latent variables; construct a decoder for the conditional variational autoencoder, and combine the sample latent variables with the sample data... The input is concatenated and defined as the decoder input. After being mapped by the shared hidden layer, it branches into multiple output heads, wherein the multiple output heads respectively output the probability distribution of recommendation time, the probability distribution of recommendation site, the probability distribution of recommendation behavior type, and the target value of recommendation charge state; the conditional prior distribution driven by the multidimensional preference vector of the sample through linear mapping is defined; the training loss is calculated, and the parameters of the conditional variational autoencoder, the shared hidden layer, and the multiple output heads are updated to obtain the scheduling decision-maker; wherein the training loss is at least a weighted sum of reconstruction loss, KL divergence loss, hidden layer distillation loss, and conditional consistency loss.
[0082] Furthermore, the scheme recommendation module 14 is also used to: obtain the output probability distribution of the local behavior prediction model in the current decision-making scenario; obtain the recent V2G behavior statistical features of the virtual neighborhood users corresponding to the causal neighborhood within the causal delay window, and generate a causal neighborhood signal; concatenate the causal neighborhood signal with the output probability distribution to form evidence input; use the multidimensional preference vector as a conditional input, and input it together with the evidence input into the encoder in the scheduling decision-maker to generate a probability distribution in the latent space, and sample latent variables from the probability distribution; the decoder in the scheduling decision-maker, based on the latent variables and the conditional input, decodes and generates V2G recommendation schemes by sharing a hidden layer and multiple output heads, and outputs a set of recommended scheduling schemes.
[0083] Furthermore, the scheme recommendation module 14 is also used to: decode the multiple latent variables respectively through a decoder to generate multiple candidate scheduling schemes; perform feasibility filtering on each of the candidate scheduling schemes, and perform multi-dimensional scoring on the candidate scheduling schemes that pass the feasibility filtering, wherein the multi-dimensional scoring includes at least the expected economic benefits and grid flexibility contribution value of normalized weighted fusion; based on the multi-dimensional scoring results, serialize the multiple candidate scheduling schemes in descending order, and select a preset number of candidate scheduling schemes from top to bottom from the top of the sequence, constrained by the difference between any two recommended schemes not being lower than a preset difference threshold, to form the recommended scheduling scheme set.
[0084] The intelligent scheduling system for power grid flexibility resources provided in this application can achieve intelligent flexibility resource scheduling management in scenarios such as urban virtual power plant operation and management, electric vehicle charging network optimization scheduling, and power demand response resource aggregation. This management process involves everything from vectorizing user V2G behavior and decoupling multi-dimensional preferences to constructing homogeneous causal spatiotemporal multi-dimensional virtual neighborhoods, and then to recommending scheduling schemes based on collective experience federated distillation and multi-source evidence fusion. It can be integrated into virtual power plant aggregation platforms, charging operator scheduling systems, or vehicle-to-everything (V2X) cloud service platforms. This effectively improves the matching accuracy between V2G scheduling schemes and actual user behavior preferences, reduces the risk of response participation rate attenuation due to recommended schemes deviating from user habits, and provides structured decision support for power grid scheduling plans and user travel planning, including recommended time, recommended sites, recommended behavior types, and recommended state of charge target values. For the specific scheduling process and identification details of this system, please refer to Example 1.
[0085] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are possible or may be advantageous.
Claims
1. A method for intelligent scheduling of power grid flexibility resources, characterized in that, The method includes: Acquire historical V2G behavior data of target users, and analyze to obtain behavioral representation vectors and multidimensional preference vectors of target users; Based on the behavior representation vector, a multidimensional virtual neighborhood of the target user is constructed, wherein each virtual neighborhood user in the multidimensional virtual neighborhood is labeled with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. Based on the homogeneous neighborhood in the multidimensional virtual neighborhood, a group decision model corresponding to the target user is established, and the group decision knowledge in the form of probability distribution generated by the group decision model on the preset query scenario set is obtained and sent to the target user's local area for distillation training to obtain a local behavior prediction model. The evidence input is constructed by combining the local behavior prediction model with the causal neighborhood, and the multidimensional preference vector is used as the condition input. The scheduling scheme decision is made by the scheduling decision-maker based on the conditional variational autoencoder to obtain a set of recommended scheduling schemes.
2. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, Obtain historical V2G behavior data of target users, and analyze it to obtain behavioral representation vectors and multidimensional preference vectors of target users, including: Acquire historical V2G behavior data of target users and record it in a structured manner as behavior events containing timestamps, site numbers, behavior types and state vectors, forming a sequence of behavior events; For each behavioral event, time feature embedding, site feature embedding, behavior type embedding, and state vector mapping are performed and concatenated to obtain an event embedding vector; Based on the actual time interval between adjacent behavioral events, a sinusoidal position code is generated, and the sinusoidal position code is superimposed on the corresponding event embedding vector to obtain a position-aware event sequence; The location-aware event sequence is input into a pre-trained temporal encoder, and after multi-layer multi-head self-attention encoding, it is pooled along the sequence dimension to obtain the behavior representation vector.
3. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, This includes acquiring historical V2G behavior data of target users, analyzing and obtaining behavioral representation vectors and multidimensional preference vectors of target users, and also: Based on the historical V2G behavior data, actual adoption behavior is identified as positive example samples, non-adoption behavior is identified as negative example samples, and user browsing time exceeding a preset time threshold and not executed behavior is identified as hard negative example samples. Based on the positive sample, the negative sample, and the hard negative sample, a contrastive sample pair is constructed, and a preference encoder is trained through a contrastive learning method, wherein the attention weight of the hard negative sample in the loss function of the contrastive learning is higher than the attention weight of the negative sample. By using multiple independent decoding heads, the preference vector output by the preference encoder is structurally decomposed into time preference subspace, location preference subspace, price sensitivity subspace, convenience preference subspace, and environmental awareness subspace to obtain the multidimensional preference vector, wherein the output of each subspace is the behavioral probability.
4. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, Based on the behavioral representation vector, a multidimensional virtual neighborhood of the target user is constructed. Each virtual neighborhood user in the multidimensional virtual neighborhood is labeled with a neighborhood type, and the neighborhood types include at least homogeneous neighborhoods, causal neighborhoods, and spatiotemporal neighborhoods, including: Based on the cosine similarity between the behavioral representation vectors of multiple users, users exceeding a first preset threshold are selected as homogeneous virtual neighborhood users, and the homogeneous neighborhood is obtained. The transfer entropy is calculated based on the V2G participation rate time series of each user, the transfer entropy values between multiple user pairs are obtained, and the significance is judged by permutation test. Users whose transfer entropy values exceed the second preset threshold and whose significance level meets the preset requirements are selected as causal virtual neighborhood users, and the causal delay window is determined to obtain the causal neighborhood. Based on the spatiotemporal co-occurrence relationship between the behavioral events of each user, users whose co-occurrence frequency exceeds a third preset threshold within a preset time window and preset spatial distance are selected as spatiotemporal virtual neighborhood users, and the spatiotemporal neighborhood is obtained. The homogeneous neighborhood, the causal neighborhood, and the spatiotemporal neighborhood are combined to form the multidimensional virtual neighborhood.
5. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, Based on the homogeneous neighborhoods in the multidimensional virtual neighborhood, a group decision-making model corresponding to the target user is established, and the group decision-making knowledge in the form of a probability distribution generated by the group decision-making model on a preset query scenario set is obtained. This knowledge is then sent to the target user's local machine for distillation training to obtain a local behavior prediction model, including: Obtain the local model parameters of the current status local decision model for each virtual neighbor in the multidimensional virtual neighborhood; Based on the neighborhood type and association degree of each virtual neighborhood user, the aggregation weight is calculated, and the local model parameters of each virtual neighborhood user are weighted and aggregated with the model parameters of the target user to obtain the group decision model. Based on the aforementioned group decision-making model, reasoning and decision-making are performed on a preset set of query scenarios. After temperature scaling of the reasoning output, it is normalized to generate group decision-making knowledge in the form of a probability distribution. Injecting privacy-preserving noise with a Laplace distribution into the group decision knowledge, and then performing truncation and normalization post-processing, yields secure group decision knowledge; The security group decision-making knowledge is distributed to the target user's local machine. Combined with the target user's local behavior data, the target user's current local decision-making model is distilled and trained to obtain a local behavior prediction model.
6. The intelligent scheduling method for power grid flexibility resources as described in claim 5, characterized in that, The loss function for distillation training includes a weighted sum of the local behavior prediction loss, distillation loss, and preference alignment loss, where: The local behavior prediction loss is the cross-entropy loss of the local behavior prediction model on local behavior data. The distillation loss is defined as the KL divergence between the output of the local behavior prediction model on the preset query scenario set and the security group decision knowledge, and the weight of the distillation loss decreases linearly with the increase of training rounds. The preference alignment loss is defined as the cosine similarity loss between the intermediate hidden layer representation of the local behavior prediction model after linear projection and the multidimensional preference vector.
7. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, The construction of the scheduling decision-maker includes: Construct the encoder of the conditional variational autoencoder, and concatenate the sample evidence input with the sample condition input that is randomly set to zero according to a preset probability, which is defined as the encoder input. Define the sample mean vector and sample log variance vector of the latent space as the encoder output. Based on the reparameterized sampling strategy, the sample mean vector is added element-wise to the sample standard deviation and standard normal random noise calculated based on the sample log-variance vector to obtain the sample latent variables; A decoder for a conditional variational autoencoder is constructed by concatenating the latent variables of the samples with the conditional input of the samples, which is defined as the decoder input. After mapping through a shared hidden layer, the decoder branches into multiple output heads, wherein the multiple output heads respectively output the probability distribution of recommendation time, the probability distribution of recommendation site, the probability distribution of recommendation behavior type, and the target value of recommendation state of charge. Define the conditional prior distribution driven by a linear mapping of the sample multidimensional preference vector; Calculate the training loss and update the parameters of the conditional variational autoencoder, the shared hidden layer, and the multiple output heads to obtain the scheduling decision-maker; The training loss is at least a weighted sum of the reconstruction loss, KL divergence loss, hidden layer distillation loss, and conditional consistency loss.
8. The intelligent scheduling method for power grid flexibility resources as described in claim 1, characterized in that, The system combines the local behavior prediction model with the causal neighborhood to construct evidence input, and uses the multidimensional preference vector as conditional input. A scheduling decision-maker based on a conditional variational autoencoder makes scheduling scheme decisions to obtain a recommended scheduling scheme set. The system also includes: Obtain the output probability distribution of the local behavior prediction model in the current decision-making scenario; Obtain the recent V2G behavior statistics of the virtual neighbor users corresponding to the causal neighbor within the causal delay window, and generate a causal neighbor signal; The causal neighborhood signal and the output probability distribution are combined to form the evidence input; The multidimensional preference vector is used as a conditional input and, together with the evidence input, is input into the encoder of the scheduling decision-maker to generate a probability distribution in the latent space, and latent variables are sampled from the probability distribution. The decoder in the scheduling decision-maker, based on the hidden variables and the conditional input, decodes and generates V2G recommended schemes by sharing a hidden layer and multiple output heads, and outputs a set of recommended scheduling schemes.
9. The intelligent scheduling method for power grid flexibility resources as described in claim 8, characterized in that, The decoder in the scheduling decision-maker, based on the latent variables and the conditional input, decodes and generates V2G recommendation schemes through a shared hidden layer and multiple output heads, outputting a set of recommended scheduling schemes. It also includes: The hidden variables are decoded by a decoder to generate multiple candidate scheduling schemes; Each of the candidate scheduling schemes is subjected to feasibility filtering, and the candidate scheduling schemes that pass the feasibility filtering are given a multi-dimensional score, wherein the multi-dimensional score includes at least the expected economic benefits of normalized weighted fusion and the contribution value of grid flexibility. Based on the multidimensional scoring results, multiple candidate scheduling schemes are serialized in descending order. With the constraint that the difference between any two recommended schemes is not lower than a preset difference threshold, a preset number of candidate scheduling schemes are selected from top to bottom of the sequence to form the recommended scheduling scheme set.
10. A smart scheduling system for power grid flexibility resources, characterized in that, The system is used to execute the intelligent scheduling method for power grid flexibility resources according to any one of claims 1-9, and the system includes: The behavior analysis module is used to acquire historical V2G behavior data of target users and analyze and obtain the behavioral representation vector and multidimensional preference vector of target users. The neighborhood construction module is used to construct a multidimensional virtual neighborhood of the target user based on the behavior representation vector. Each virtual neighborhood user in the multidimensional virtual neighborhood is marked with a neighborhood type, and the neighborhood type includes at least homogeneous neighborhood, causal neighborhood and spatiotemporal neighborhood. The group distillation module is used to establish a group decision model corresponding to the target user based on the homogeneous neighborhood in the multidimensional virtual neighborhood, and to obtain the group decision knowledge in the form of probability distribution generated by the group decision model on a preset query scenario set, and send it to the target user's local distillation training to obtain a local behavior prediction model. The scheme recommendation module is used to combine the local behavior prediction model with the causal neighborhood to construct evidence input, and use the multidimensional preference vector as condition input to make scheduling scheme decisions through a scheduling decision-maker based on conditional variational autoencoder to obtain a set of recommended scheduling schemes.