Vehicle network interaction resource optimization scheduling method, system and device suitable for multiple adjustment requirements, and storage medium

By establishing a unified graph structured representation of multi-source data and Rainbow's improved DQN algorithm, the problems of system operation stability and resource utilization efficiency in vehicle-to-grid interactive resource scheduling are solved, and the coordinated optimization of power supply reliability, economy and user satisfaction is achieved.

CN121546727APending Publication Date: 2026-02-17NARI NANJING CONTROL SYSTEM CO LTD +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511730038.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing vehicle-to-grid (V2G) interactive resource scheduling technologies struggle to accurately match grid control demands in complex scenarios, leading to decreased system stability and unbalanced equipment loads. This results in an inability to effectively coordinate grid safety, user charging experience, and resource utilization efficiency.

Method used

By establishing a unified graph structured representation of multi-source data, and utilizing a Rainbow-based improved Deep Q-Network (DQN) algorithm, a high-level state representation of user charging and discharging behavior is generated. A Markov decision process is then constructed to optimize charging resource scheduling decisions, thereby synergistically optimizing power supply reliability, economy, and user satisfaction.

Benefits of technology

It has achieved a balance between charging station operating costs and user charging experience while ensuring the safe and stable operation of the power distribution network, thereby improving the efficiency of new energy utilization and optimizing resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121546727A_ABST
    Figure CN121546727A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle network interaction resource optimization scheduling method, system and device suitable for multiple adjustment demands and a storage medium, and the method comprises the steps: obtaining the feature information of an electric vehicle, energy storage, photovoltaic and power distribution network from historical charging and discharging data, so as to build the unified graph structured representation of multi-source data; analyzing the non-linear influence of the external influence factors on the charging and discharging behaviors of all nodes in the unified graph structured representation, and generating advanced state representation for predicting the charging and discharging behaviors of the user; based on advanced state representation, an optimization problem for realizing an optimal scheduling decision of the charging resources in a next scheduling period is established, the optimization problem is converted into a Markov decision process, the Markov decision process is solved by using an improved DQN algorithm, and the optimal scheduling decision of the charging resources in a specific state is obtained. The method overcomes the defect of a single optimization target in the prior art, and can effectively balance the operation cost of the charging station, the charging experience of a user and the utilization efficiency of new energy on the premise of guaranteeing the safe and stable operation of a power grid, thereby achieving the collaborative optimization scheduling of resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system technology, and in particular to a method, system, device, and storage medium for vehicle-grid interactive resource optimization scheduling applicable to multiple regulation needs. Background Technology

[0002] In vehicle-grid interaction scenarios, the dynamic aggregation characteristics of electric vehicle charging loads in time and space alter the original operating logic of the power system. Furthermore, the integration of distributed photovoltaic (PV) and energy storage resources further enhances the multi-dimensional dynamic coupling characteristics of the "source-grid-load-storage" system. Against this backdrop, the disorderly integration of charging loads can lead to increased grid load fluctuations and greater operational pressure on distribution equipment. Simultaneously, the spatiotemporal mismatch between PV / storage resources and charging demand further complicates power system regulation. Current single-objective-based scheduling strategies are no longer adequate to meet the practical needs of multi-resource collaborative optimization in vehicle-grid interaction systems.

[0003] Existing vehicle-to-grid (V2G) resource scheduling technologies have significant shortcomings when dealing with complex scenarios. On the one hand, static scheduling models fail to fully consider the randomness and adjustability of user charging behavior, leading to large prediction errors in charging load demand and difficulty in accurately matching grid regulation needs. On the other hand, a framework with a single optimization objective cannot effectively coordinate the contradictory relationships between multiple dimensions of indicators such as grid safety, user charging experience, and resource utilization efficiency, which can easily lead to problems such as decreased system stability and unbalanced equipment load in practical applications. Therefore, there is an urgent need to research a V2G resource scheduling method applicable to multiple regulation needs to achieve collaborative and optimized resource scheduling. Summary of the Invention

[0004] Purpose of the invention: The purpose of this invention is to provide a method, system, device, and storage medium for optimizing the scheduling of vehicle-to-everything (V2X) interactive resources with multiple adjustment requirements, so as to achieve collaborative optimization scheduling of resources.

[0005] Technical solution: To achieve the above objectives, the present invention provides a vehicle-to-grid (V2G) interactive resource optimization scheduling method applicable to multiple adjustment needs, comprising: obtaining characteristic information of electric vehicles, energy storage, photovoltaics and distribution networks from historical charging and discharging data to establish a unified graph structured representation of multi-source data, analyzing the nonlinear influence of external influencing factors on the charging and discharging behavior of all nodes in the unified graph structured representation, and generating a high-level state representation for predicting user charging and discharging behavior.

[0006] Based on high-level state representation, an optimization problem is established to achieve the optimal scheduling decision of charging resources within a future scheduling cycle. The optimization problem is then transformed into a Markov decision process. The Markov decision process is solved using a deep Q-network based on the Rainbow algorithm (a DQN algorithm based on Rainbow), to obtain the optimal scheduling decision of charging resources under a specific state.

[0007] Preferably, the unified graph structured representation is expressed as: ,

[0008] in, This represents a time series, a sequence of moments, where each moment t corresponds to a snapshot G(t). Represents the set of external influencing factors. V represents the set of mapping functions used to map, transform, and fuse multi-source heterogeneous information; V represents the set of nodes, including electric vehicle nodes, energy storage nodes, photovoltaic nodes, and distribution network nodes; E represents the set of edges, containing all physical or logical connections between nodes.

[0009] Preferably, the method for establishing a unified graph structured representation of multi-source data includes: first constructing a multi-source dataset based on feature information, and then constructing a unified graph structured representation based on the multi-source dataset.

[0010] Preferably, the method for constructing a multi-source dataset is as follows: first, identify the nodes and edges in the feature information, align all the data contained in the feature information according to the timestamp, and convert them into feature vectors of nodes and edges in a graph structure to construct a multi-source dataset aligned with timestamps regarding the interaction between user charging and discharging behavior and the power distribution network.

[0011] Preferably, for each timestamp t ∈ T, the dynamic feature vectors of all nodes and edges are obtained to form a graph snapshot: G(t) = (X v(t) X e(t) ), where X v(t) X represents the node feature matrix, which is the set of feature vectors of all nodes V at time stamp t. e(t) Let represent the edge feature matrix, which is the set of feature vectors of all edges E at time stamp t.

[0012] Preferably, the method for generating a high-level state representation for predicting user charging and discharging behavior is as follows: for each node, the node's own feature vector, the neighbor aggregated feature vector, and the external influencing factors are concatenated, and all timestamps are traversed to construct the node's feature sequence; the feature sequence is input into a bidirectional long short-term memory network to output a high-level feature vector containing complete temporal patterns, and then the high-level feature vector is input into a Bayesian output layer to generate a predicted value for the node's charging and discharging behavior in the next time period.

[0013] Preferably, the optimization problem of the optimal scheduling decision for charging resources includes an objective function and constraints, wherein the objective function is expressed as:

[0014] ,

[0015] In the formula, These are weighting coefficients used to reflect the priority of each sub-objective; These are the normalized power supply reliability target, economic target, and user satisfaction target, respectively.

[0016] The constraints include power balance constraints, charge and discharge load regulation constraints, battery state of charge limits, electric vehicle charge and discharge mutual exclusion constraints, energy storage constraints, and photovoltaic output constraints.

[0017] Preferably, the transformation of the optimization problem into a Markov decision process includes: constructing a state space containing the electric vehicle state, photovoltaic output, and energy storage capacity; using the charging pile start-up time as the action space; and designing a reward function that incorporates user satisfaction, operating costs, and curtailment penalties. The state space is represented as follows:

[0018] ,

[0019] In the formula, and The time it takes for an electric vehicle to arrive at a charging station and its SOC; and These represent the departure time of electric vehicle i and the expected SOC; Contribute to photovoltaic power; The remaining power of the energy storage system; Let t be the total charging load in the charging station at time t; Real-time electricity price for the distribution network; Dispatch instructions issued to the distribution network;

[0020] The action space is represented as:

[0021] ,

[0022] In the formula, Indicates the charging pile's operating time; This indicates the time that electric vehicle i stays in the charging parking space;

[0023] The reward function is expressed as:

[0024] ,

[0025] In the formula, Reward users for charging satisfaction. As a reward for the operating costs of charging stations, Penalties for curtailment of solar power.

[0026] Preferably, the Rainbow-based improved DQN algorithm includes a dual-Q network, a competition network, a priority replay buffer mechanism, and a learning rate decay strategy. Solving the Markov decision process using the Rainbow-based improved DQN algorithm includes: constructing a dual-Q network, replaying state transition tuples consisting of state space, action space, and reward function in the experience pool according to the time-series difference error, selecting the optimal action using one Q network, evaluating the value of the action using the other Q network to alleviate Q-value overestimation, splitting state value and action advantage using a competition network to optimize network parameters, and dynamically decaying the learning rate for strategy iteration until the output is the optimal charging resource scheduling decision that can achieve objective functions such as power supply reliability, economy, and user satisfaction under the constraints.

[0027] The present invention discloses a vehicle-to-grid (V2G) interactive resource optimization and scheduling system applicable to multiple adjustment needs, comprising the following modules:

[0028] Multi-source data graph modeling and state prediction module: used to obtain characteristic information of electric vehicles, energy storage, photovoltaic and distribution networks from historical charging and discharging data, in order to establish a unified graph structured representation of multi-source data, analyze the nonlinear influence of external influencing factors on the charging and discharging behavior of all nodes in the unified graph structured representation, and generate a high-level state representation for predicting user charging and discharging behavior.

[0029] The reinforcement learning-based resource optimization scheduling module is used to establish an optimization problem for achieving the optimal scheduling decision of charging resources within a future scheduling cycle based on high-level state representation. The optimization problem is transformed into a Markov decision process, and the Markov decision process is solved using the Rainbow-based improved DQN algorithm to obtain the optimal scheduling decision of charging resources under a specific state.

[0030] A computing device according to the present invention includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for performing the methods described above.

[0031] The present invention provides a computer-readable storage medium for storing one or more programs, the one or more programs including instructions that, when executed by a computing device, cause the computing device to perform the method described above.

[0032] Beneficial effects: The present invention has the following advantages: The present invention overcomes the drawbacks of the single optimization objective in the prior art, and can effectively balance the operating cost of charging stations, user charging experience and new energy utilization efficiency while ensuring the safe and stable operation of the power distribution network, so as to achieve resource collaborative optimization scheduling. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the method flow of the present invention;

[0034] Figure 2 A schematic diagram of a vehicle charging solution. Detailed Implementation

[0035] The technical solution of the present invention will be described in detail below with reference to the embodiments and accompanying drawings.

[0036] like Figure 1 As shown, the vehicle-to-grid (V2G) interactive resource optimization scheduling method applicable to multiple adjustment needs, as described in this invention, includes the following:

[0037] S1. Collect historical charging and discharging data of electric vehicles, power distribution networks, energy storage and photovoltaic systems to obtain characteristic information of electric vehicles, energy storage, photovoltaic and power distribution networks.

[0038] Electric vehicle characteristic information includes the time when each electric vehicle enters and is expected to leave the charging station, vehicle SOC (state of charge), charging and discharging power, user charging time, etc.; energy storage characteristic information includes remaining capacity, energy storage SOC, upper and lower limits of charging and discharging power, and cycle efficiency, etc.; photovoltaic characteristic information includes real-time output, rated capacity, and irradiance, etc.; distribution network characteristic information includes time-of-use pricing, base load curve, dispatch instructions, and power, etc.

[0039] All feature information data is aligned by timestamp and transformed into feature vectors of nodes and edges in a graph structure, constructing a timestamp-aligned multi-source dataset about user charging and discharging behavior and its interaction with the power distribution network. The construction process includes: first, identifying the nodes and edges in the feature information, and then using the aligned data as dynamic feature vectors for each node and edge.

[0040] S2. Establish a unified graph structured representation of multi-source data.

[0041] S201. First, define the static graph structure: G = (V, E), where V represents the set of nodes, including electric vehicle nodes, energy storage nodes, photovoltaic nodes, and distribution network nodes; E represents the set of edges, containing all physical or logical connections between nodes.

[0042] S202. For each timestamp t ∈ T, obtain the dynamic feature vectors of all nodes and edges to form a graph snapshot: G(t) = (X v(t) X e(t) ), where X v(t) X represents the node feature matrix, which is the set of feature vectors of all nodes V at time stamp t. e(t) Let represent the edge feature matrix, which is the set of feature vectors of all edges E at time stamp t.

[0043] S203. Construct a unified graph structured representation: ,

[0044] in, This represents a time series, a sequence of moments, where each moment t corresponds to a given value G(t). This represents a set of external influencing factors (such as weather data and global dispatch instructions). This represents a set of mapping functions used to map, transform, and fuse multi-source heterogeneous information. For example, it can associate and transform node features (such as electric vehicle SOC and photovoltaic output), edge features, time series information, and external influencing factors, thereby integrating these scattered information into a unified graph structured representation, providing structured and computable support for subsequent analysis and power system resource scheduling optimization decisions.

[0045] S3. Based on graph structured representation, analyze the nonlinear impact of external influencing factors on the charging and discharging behavior of all nodes, and generate a high-level state representation that includes predictions of future user charging and discharging behavior.

[0046] S301, For each node V j At time t, using the graph structure G and the graph snapshot G(t), aggregate all data related to V. j The feature vectors of directly connected neighboring nodes, and the connecting node V j The feature vectors of the edges between node V and its neighboring nodes are used to generate a neighbor aggregate feature vector, which represents the feature vector of node V. j The status and interaction information of all surrounding neighbors; Node V j The node V is formed by concatenating its own feature vector, the neighbor aggregated feature vector, and the external influence factor Ω(t). j The context-aware feature vector Z at time t j(t) Iterate through all timestamps and construct node V. j Feature sequence Z j = {Z j(t1) Z j(t2) ,..., Z j(tn)}, where j represents the index of a node within the node set.

[0047] S302, The feature sequence Z j The input is a bidirectional long short-term memory network (Bi-LSTM), which processes the feature sequence Z through two processes: forward and backward. j Capture feature sequence Z j It relies on long-term and bidirectional contextual information to output a high-level feature vector containing complete temporal patterns;

[0048] S303. Input the high-level feature vector into the Bayesian output layer. The Bayesian output layer models its own weights as a probability distribution and samples multiple sets of weights from the weight distribution. It then performs forward computation on the input high-level feature vector to generate a predicted value about the charging and discharging behavior of node Vj in the next time period.

[0049] S4. Considering the coordination of three factors—power supply reliability, economy, and user satisfaction—establish an optimization problem to achieve optimal scheduling of charging resources within a future scheduling cycle, including the objective function and constraints. Charging resource scheduling refers to the coordinated optimization of multiple objectives such as power supply reliability, economy, and user satisfaction through reasonable scheduling of charging and discharging resources of electric vehicles, distribution networks, energy storage, and photovoltaic systems, under the premise of meeting constraints.

[0050] S401, Objective Function:

[0051] ,

[0052] In the formula: These are weighting coefficients used to reflect the priority of each sub-objective; These are the normalized power supply reliability target, the economic target, and the user satisfaction target.

[0053] S4021, based on the distribution network load variance As a power supply reliability target, it is expressed as:

[0054] ,

[0055] After normalization:

[0056] ,

[0057] In the formula, For the load variance of the distribution network, T1 represents the maximum and minimum values ​​of the load variance; T1 is the number of time periods in a day. This represents the base load of the distribution network during time period t1. This represents the total number of charging stations. Let be the charging load and discharging load of electric vehicles at the u-th charging station at time t; This represents the average load of the distribution network over a day.

[0058] Reducing load fluctuations in the distribution network can significantly reduce system losses and improve the operational stability of the distribution network.

[0059] S4022, Revenue from charging stations As an economic objective, it is specifically expressed as: (Charging station electricity revenue minus user subsidies).

[0060] ,

[0061] After normalization:

[0062] ,

[0063] In the formula, The revenue from charging electricity at the charging station at time t. The discharge subsidy given to electric vehicle users by the charging station at time t; These represent the maximum and minimum net revenue of the charging station.

[0064] S4023, User satisfaction objectives include user satisfaction at the time level. User SOC status level satisfaction The normalized user satisfaction target is expressed as:

[0065] ,

[0066] In the formula: Weighted by time satisfaction.

[0067] Among them, user satisfaction at the time level Represented as:

[0068] ,

[0069] In the formula, The estimated time for electric vehicle i to leave the charging station; The moment when charging of electric vehicle i ends; The actual time that electric vehicle i leaves the charging station; The time it takes for an electric vehicle to enter a charging station. The greater the difference between the time it takes to finish charging and the expected time to leave, the higher the user's satisfaction in terms of time.

[0070] User SOC status level satisfaction Represented as:

[0071] ,

[0072] In the formula, The battery SOC state of electric vehicle i when it actually leaves the charging station. To achieve the desired battery SOC state, and The smaller the gap, the higher the user satisfaction.

[0073] S402. Construct scenario-based constraints.

[0074] The constraints include power balance constraints, charge and discharge load regulation constraints, battery state of charge limits, electric vehicle (EV) charge and discharge mutual exclusion constraints, energy storage constraints, and photovoltaic output constraints.

[0075] S4021, Power Balance Constraint

[0076] At any given moment during system operation, the power input to the distribution network, the power output of photovoltaic systems, the charging and discharging power of energy storage devices, and the charging and discharging power of electric vehicles must satisfy the law of conservation of energy, meaning the algebraic sum of the power of each component equals the total load demand of the system. The power balance constraint is expressed as:

[0077] ,

[0078] In the formula: To input active power into the distribution network; To output active power for photovoltaics; Active power for energy storage / EV charging and discharging; M represents the total active power load demand of the system; M represents the total number of EVs participating in the scheduling. Input reactive power to the distribution network; Reactive power regulation for energy storage / EVs; This represents the total reactive power load requirement of the system.

[0079] S4022, Charge / Discharge Load Regulation Constraint

[0080] Considering the capacity limitations of charging and discharging facilities and the load fluctuation tolerance of the distribution network, the charging and discharging power... The power must be limited to the equipment's permissible range to ensure safe operation. The charge / discharge load regulation constraint is expressed as:

[0081] ,

[0082] In the formula: This is the maximum charging power; This represents the maximum discharge power.

[0083] S4023, Battery State of Charge Limitation

[0084] The battery's state of charge (SCC) must be maintained within a safe operating range, meaning it must not fall below the minimum allowable value to avoid over-discharge damage and must not exceed the maximum allowable value to prevent overcharging risks, thus ensuring the battery's reliability and safety throughout its entire lifespan. Battery SCC limits are expressed as follows:

[0085] ,

[0086] In the formula: Let be the state of charge of electric vehicle i at time t; Minimum / maximum permissible SOC for the battery of an electric vehicle.

[0087] S4024, EV charging and discharging mutual exclusion constraint: Electric vehicles cannot simultaneously charge and discharge, therefore a charging and discharging mutual exclusion constraint is introduced. The EV charging and discharging mutual exclusion constraint is expressed as:

[0088] ,

[0089] In the formula: At that time, the EV is in a charging state. At this time, the EV is in a discharging state.

[0090] S4025, Energy Storage Constraints

[0091] The charging and discharging process of an energy storage system must meet power boundary conditions and capacity constraints, while also considering charging and discharging efficiency losses, to ensure that the energy storage device operates under reasonable conditions and achieves smooth regulation of the distribution network load and timely energy distribution. The energy storage constraints are expressed as:

[0092] ,

[0093] In the formula: The remaining capacity of the energy storage at time t, This refers to the minimum allowable capacity / rated capacity of energy storage. For energy storage charging and discharging power, This represents the maximum charging / discharging power of the energy storage.

[0094] S4026, Photovoltaic Output Constraints

[0095] The real-time output power of a photovoltaic (PV) array is limited by the current irradiance, temperature, and equipment conversion efficiency, and its maximum value does not exceed the rated power of the PV modules. Actual output is determined through a PV power prediction model or real-time measurement data, requiring that the PV output during system scheduling not exceed the available power. The PV output constraint is expressed as:

[0096] ,

[0097] In the formula: The photovoltaic system outputs active power at time t. The rated active power of photovoltaic power. This is a dimensionless correction factor for photovoltaic output.

[0098] S5. The optimization problem is transformed into a Markov decision process, and a reinforcement learning environment is constructed by defining the state space, action space, and reward function. Specifically, this involves constructing a state space that includes the electric vehicle's state, photovoltaic output, and energy storage capacity, using the charging pile's charging start-up time as the action space, and designing a reward function that incorporates user satisfaction, operating costs, and curtailment penalties.

[0099] S501, State Space: The constructed state space needs to accurately describe the environmental information of the electric vehicle, charging station, and power distribution network. The state space is represented as follows:

[0100] ,

[0101] In the formula, and The time it takes for an electric vehicle to arrive at a charging station and its SOC; and These represent the departure time of electric vehicle i and the expected SOC; Contribute to photovoltaic power; The remaining power of the energy storage system; Let t be the total charging load in the charging station at time t; Real-time electricity price for the distribution network; Dispatch instructions issued to the distribution network.

[0102] S502, Action Space: This space, combined with status information, controls the charging start-up time of the charging pile. The action space is represented as follows:

[0103] ,

[0104] In the formula: Indicates the charging pile's operating time; This indicates the time that electric vehicle i stays in the charging parking space.

[0105] S503, the reward function is expressed as follows:

[0106] ,

[0107] In the formula, Reward users for charging satisfaction. As a reward for the operating costs of charging stations, Penalties for curtailment of solar power.

[0108] S6. Solve the Markov decision process using the Rainbow-based improved DQN algorithm to obtain the optimal scheduling decision for charging resources under a specific state.

[0109] S601. Design an improved mechanism for Rainbow's DQN algorithm.

[0110] The DQN algorithm demonstrates strong problem-solving capabilities for time-series decision-making problems, but it suffers from shortcomings in generalization performance, learning ability, computational efficiency, and convergence performance. Therefore, this invention improves the DQN algorithm using the Rainbow algorithm, introducing a learning rate decay strategy from deep learning. This allows the agent to increase its exploration capabilities in the early stages of learning and decrease the learning rate in later stages, effectively utilizing previous experience.

[0111] The Rainbow-based improved DQN algorithm includes a dual-Q network, a competition network, a priority replay caching mechanism, and a learning rate decay strategy. By integrating multiple optimization techniques, it improves the policy search efficiency and decision stability of the DQN algorithm in vehicle-to-everything (V2X) interaction scenarios.

[0112] The dual-Q network separates the functions of the evaluation network and the target network: the evaluation network selects the action with the maximum Q value, and the target network calculates the Q value of the action, which alleviates the problem of Q value overestimation in the original DQN algorithm and improves the accuracy of distribution network load forecasting and scheduling strategies.

[0113] The competitive network modifies the structure of the evaluation network by splitting the evaluation network into state values. and the advantages of movement This allows the evaluation network to focus more on the importance of the state itself, rather than solely relying on action selection, thus enhancing its ability to represent multi-dimensional features such as distribution network state and user needs. When situations frequently arise where agents take different actions but the corresponding value functions differ only slightly, the competitive network can remove redundant degrees of freedom, improving algorithm stability.

[0114] The priority replay caching mechanism uses time-series error (TD-error) to weight the sampling of experience samples. Samples with large TD-error have a higher probability of participating in training, which accelerates the learning efficiency of key information such as real-time data of distribution network and charging load fluctuations.

[0115] The learning rate decay strategy dynamically adjusts the learning rate based on the number of iterations. In the early stages, it increases the exploration capability to adapt to dynamic topology changes in vehicle-to-network interaction, and in the later stages, it decreases the learning rate to stabilize the strategy output.

[0116] S602, the solution process is as follows:

[0117] S6021. Network Parameters and Experience Pool Configuration: Construct the evaluation network and the target network, and separate the state-value stream. With the advantage of action flow Output the Q-value matrix; initialize the priority experience replay pool, and set sample weights based on TD-error.

[0118] S6022, State Space Initialization: Collect real-time electricity prices, photovoltaic output, remaining energy storage capacity, charging load, SOC, and other data from the distribution network to form the initial state. ;

[0119] S6023, Action Selection: Inject random noise into the evaluation network output, dynamically adjust the exploration strategy, and generate actions. To balance the needs of power distribution network dispatching with users' charging preferences;

[0120] S6024, Status and Reward Update: Based on Action Control the charging pile start-up time, update the EV charging status, calculate real-time data such as photovoltaic utilization rate and load fluctuation, and generate the next state. ;

[0121] The state-action value function is expressed as:

[0122] .

[0123] S6025, Dual-Q Network Optimization and Experience Replay: The state-action-reward quadruple is stored in the experience pool, and the sample priority is determined according to the temporal difference error; the evaluation network selects actions, and the target network calculates the corresponding Q value to generate the target value, thereby optimizing the evaluation network parameters, and periodically synchronizing the evaluation network parameters to the target network.

[0124] S6026, Dynamically Adjusting Learning Rate and Policy Iteration: The learning rate is dynamically decayed according to the initial learning rate and decay coefficient, and the state-action value function is optimized by combining a multi-step reward mechanism to balance policy exploration capability and stability.

[0125] S6027. Termination of Iteration and Output of Results: When the preset maximum number of training rounds is reached or the improvement of system indicators is less than 0.5%, the iteration is terminated and the optimized vehicle-to-grid interactive resource scheduling strategy is output to achieve the goals of reducing load peak-valley difference, improving photovoltaic utilization rate and user satisfaction.

[0126] This invention was simulated in an integrated charging station with photovoltaic and energy storage systems. All simulations were verified in an environment with an INTER I7 7700HQ server, an NVIDIA GTX 1050, 16GB of RAM, and the simulation software Matlab2024a.

[0127] like Figure 2 As shown in Table 1, this is a partial vehicle charging scheme for the period from 8:00 to 10:30. The comparison of various indicators before and after resource optimization scheduling is illustrated in Table 1. Under this strategy, the photovoltaic utilization rate of the charging station increased to 92%, improving user satisfaction and alleviating the pressure of charging load on the power distribution network.

[0128] Table 1 Comparison of various indicators before and after scheduling

[0129] Is it scheduled? Peak-to-valley difference in charging load (kW) Photovoltaic utilization rate (%) User satisfaction (%) Before scheduling 312 82 69 After scheduling 242 92 91

Claims

1. A method for optimal scheduling of vehicle-to-grid interactive resources for multi-regulation demand, characterized in that, The application relates to a method for realizing optimal charging resource scheduling of electric vehicles and photovoltaic power generation. The method comprises the following steps: obtaining characteristic information of electric vehicles, energy storage and photovoltaic power generation from historical charging and discharging data to establish a unified graph structured representation of multi-source data; analyzing nonlinear influences of external influencing factors on charging and discharging behaviors of all nodes in the unified graph structured representation; and generating a high-level state representation for predicting the charging and discharging behaviors of the nodes. The method for establishing the unified graph structured representation of multi-source data comprises the following steps: constructing a multi-source data set based on the characteristic information; and constructing the unified graph structured representation based on the multi-source data set. 2.The method of claim 1, wherein, The unified graph structured representation is expressed as: , wherein, denotes a time series, a sequence of a series of time instants, for each time instant t, a graph snapshot G(t) is associated, denotes a set of external influencing factors, denotes a set of mapping functions for realizing the mapping, conversion and fusion of multi-source heterogeneous information; V denotes a set of nodes, including electric vehicle nodes, energy storage nodes, photovoltaic nodes, and power distribution network nodes; E denotes a set of edges, including all physical or logical connections between nodes. 3.The method of claim 2, wherein, The method for constructing the multi-source data set comprises the following steps: defining nodes and edges in the characteristic information; aligning all data contained in the characteristic information according to time stamps; and converting the data into feature vectors of the nodes and edges in a graph structure to construct a multi-source data set about the charging and discharging behaviors of the nodes and the interaction between the nodes and a power grid. 4.The method of claim 3, wherein, The method for generating the high-level state representation for predicting the charging and discharging behaviors of the nodes comprises the following steps: for each node, splicing a self feature vector of the node, a neighbor aggregated feature vector and external influencing factors, and simultaneously traversing all time stamps to construct a feature sequence of the node; inputting the feature sequence into a bidirectional long short-term memory network to output a high-level feature vector containing complete time sequence rules; and inputting the high-level feature vector into a Bayesian output layer to generate a prediction value about the charging and discharging behaviors of the node in a next time period.

5. The method of claim 2, wherein, For each timestamp t ∈ T, obtain the dynamic feature vectors of all nodes and edges, forming a snapshot of the graph: G(t) = (X v(t) , X e(t) ), where X v(t) represents the node feature matrix, i.e., the set of feature vectors of all nodes V at timestamp t, and X e(t) represents the edge feature matrix, i.e., the set of feature vectors of all edges E at timestamp t.

6. The method of claim 2, wherein, The optimization problem of the optimal charging resource scheduling decision comprises a target function and constraint conditions.

7. The method of claim 1, wherein, The constraint conditions comprise power balance constraints, charging and discharging load adjustment constraints, battery state of charge limits, electric vehicle charging and discharging exclusion constraints, energy storage constraints and photovoltaic output constraints. , In the formula, is a weight coefficient, used to reflect the priority of each sub-target; are normalized power supply reliability target, economic target, and user satisfaction target, respectively. The method for converting the optimization problem into a Markov decision process comprises the following steps: constructing a state space comprising electric vehicle states, photovoltaic output and energy storage capacity; taking a charging pile charging start time as an action space; designing a reward function in combination with user satisfaction, operation cost and light curtailment penalty; wherein the state space is represented as: the action space is represented as: and the reward function is represented as: 8.The method of claim 1, wherein, The deep Q network based on the improved rainbow algorithm comprises a double Q network, a competition network, a priority replay buffer mechanism and a learning rate decay strategy. , wherein, is the time of arrival of the electric vehicle i at the charging station and the SOC; is the time of arrival of the electric vehicle i at the charging station and the SOC; is the time of departure of the electric vehicle i from the station and the desired SOC, respectively; is the time of departure of the electric vehicle i from the station and the desired SOC, respectively; is the photovoltaic output; is the remaining amount of energy of the energy storage system; is the total charging load in the charging station at time t; is the real-time electricity price of the distribution network; is the dispatching instruction issued by the distribution network. The method for solving the Markov decision process by using the deep Q network based on the improved rainbow algorithm comprises the following steps: constructing a double Q network; replaying state transition tuples composed of the state space, the action space and the reward function in an experience pool according to time sequence difference errors; selecting an optimal action by using one of the Q networks; evaluating the value of the action by using the other Q network to relieve Q value overestimation; splitting state values and action advantages by using the competition network to optimize network parameters; dynamically decaying the learning rate to perform policy iteration; and outputting an optimal charging resource scheduling decision which can realize power supply reliability, economy and user satisfaction under the condition of meeting the constraint conditions. , In the formula, denotes the charging pile action time; denotes the time of the electric vehicle i staying in the charging parking space; ​ , wherein a user charging satisfaction reward, a charging station operating cost reward, a photovoltaic light waste penalty. 9.The method of claim 1, wherein, ​ 10. A vehicle-network interaction resource optimal scheduling system suitable for multi-regulation demand, characterized in that, The method comprises the following modules: A multi-source data graph modeling and state prediction module is configured to obtain feature information of an electric vehicle, energy storage, photovoltaic and a power distribution network from historical charging and discharging data, to establish a unified graph structured representation of multi-source data, to analyze nonlinear effects of external influencing factors on charging and discharging behaviors of all nodes in the unified graph structured representation, and to generate a high-level state representation for predicting user charging and discharging behaviors; A resource optimization scheduling module based on reinforcement learning is configured to establish an optimization problem for realizing optimal scheduling decisions of charging resources in a future scheduling period based on the high-level state representation, to convert the optimization problem into a Markov decision process, to solve the Markov decision process by using a deep Q network improved based on a rainbow algorithm, and to obtain optimal scheduling decisions of the charging resources in a specific state.

11. A computing device, comprising: One or more programs stored in the one or more memories and configured to be executed by the one or more processors, the one or more programs comprising instructions for performing any one of the methods according to claims 1-9.

12. A computer-readable storage medium storing one or more programs, the one or more programs comprising instructions for: The one or more programs comprise instructions that, when executed by a computing device, cause the computing device to perform any one of the methods according to claims 1-9.