Charging pile pricing method, device and equipment, readable storage medium and program product

By pricing charging piles using reinforcement learning methods, the problem of low charging station utilization is solved, and the profitability and user satisfaction of charging stations are improved.

CN120707225APending Publication Date: 2025-09-26SHENZHEN POWER SUPPLY BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510684444.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The utilization rate of charging stations is low, mainly due to overtime parking. Existing technology makes it difficult to accurately price multiple types of charging piles, affecting user satisfaction and profitability.

Method used

A reinforcement learning-based method is adopted to divide charging piles into different nests, determine their initial utility and target utility, and dynamically adjust the charging price to optimize the operating profit of the charging station by combining the selection probability and charging demand.

Benefits of technology

It improves the utilization rate of charging stations, reduces the impact of overtime parking, ensures user satisfaction, and optimizes the profitability of charging stations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707225A_ABST
    Figure CN120707225A_ABST
Patent Text Reader

Abstract

The invention relates to a charging pile pricing method, device and equipment, a readable storage medium and a program product. The method comprises the following steps: determining an initial utility representation of each charging pile based on a convenience degree representation of the charging pile, a charging price representation of the charging pile at a moment t and a waiting time representation of the charging pile at the moment t; dividing the charging piles into different nests, and determining target utility representations of the charging piles based on the initial utility representations of the charging piles and the total error of the charging piles in the nests; determining the selection probability representation of each nest based on the utility representation of each nest, and determining the selection probability representation of each charging pile in the nest based on the selection probability representation of each nest; and determining the target price of each charging pile at the moment t through reinforcement learning based on the selection probability representation, the residual charging demand representation and the parking time representation of each charging pile and the charging price representation of the charging pile at the moment t. By adopting the method, each charging pile in multiple types of charging stations can be accurately priced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a charging pile pricing method, apparatus, device, readable storage medium, and program product. Background Art

[0002] Electric vehicles (EVs), a smart and environmentally friendly mode of transportation with low emissions and low operating costs, are experiencing rapid growth. This growth has significantly boosted the commercial prospects of the public EV charging industry. However, the construction of charging infrastructure requires high initial costs, and the profit potential of charging stations (CSs) has not yet been fully tapped.

[0003] A key factor contributing to the lack of profitability of charging stations is their low utilization rate. This is partly due to a phenomenon known as "overparking," where electric vehicle owners fail to remove their fully charged vehicles promptly. Charging stations are often self-service, and station staff cannot remove the chargers plugged into electric vehicles with interlocking charging plugs, even when the vehicles are fully charged. Electric vehicle owners often park their vehicles to charge and wait until their next trip to pick them up, which almost inevitably leads to overparking. As a solution, fines for overparking are often used to encourage owners to remove fully charged vehicles promptly, especially during peak demand periods. However, fines increase charging costs and reduce customer satisfaction, a top concern among users as electric vehicles become more widely adopted.

[0004] To address the issue of overtime parking and improve charging plug utilization, a new type of charging pile, the Single Output Multiple Plug (SOMP) charging pile, has been proposed for charging station expansion. To improve the utilization of Single Output Single Plug (SOSP) charging piles, the Multiple Output Multiple Plug (MOMP) charging pile model first introduced the capability of simultaneous multi-output through multiple ports in the charging base. However, the rated power of a Multiple Output Multiple Plug (MOMP) charging pile is equal to the sum of the power outputs of all its ports, making its construction investment cost much higher than that of a Single Output Single Plug (SOSP) charging pile. In contrast, a Single Output Multiple Plug (SOMP) charging pile is equipped with multiple ports but limits power output to only one port at a time, resulting in a power rating and investment cost similar to that of a Single Output Single Plug (SOSP) charging pile. After one vehicle is fully charged, the charging pile can automatically begin charging another connected vehicle, reducing the impact of overtime parking without compromising user satisfaction.

[0005] However, single output multiple plug (SOMP) charging stations face obstacles such as user doubts about their reliability and the need to overcome the inertia brought by users' previous habits. Therefore, there is an urgent need for a method that can accurately price charging for multiple types of charging stations, such as single output single plug (SOSP) charging piles and single output multiple plug (SOMP) charging piles. Summary of the Invention

[0006] Based on this, it is necessary to provide a charging pile pricing method, device, computer equipment, computer-readable storage medium and computer program product that can accurately price each charging pile in multiple types of charging stations in response to the above technical problems.

[0007] In a first aspect, the present application provides a charging pile pricing method, the method comprising:

[0008] Determine an initial utility representation of each charging pile based on a representation of the convenience level of the charging pile, a representation of the charging price of the charging pile at time t, and a representation of the waiting time of the charging pile at time t;

[0009] Dividing each charging pile into different nests, and determining a target utility representation for each charging pile based on an initial utility representation of each charging pile and a total error of each charging pile in the nest;

[0010] Determining a selection probability representation of each nest based on the utility representation of each nest, and determining a selection probability representation of each charging post in the nest based on the selection probability representation of each nest;

[0011] The target price of each charging pile at time t is determined by reinforcement learning based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging pile at time t.

[0012] In one embodiment, dividing the charging piles into different nests and determining a target utility representation for each charging pile based on an initial utility representation of each charging pile and a total error representation of each charging pile in the nest includes:

[0013] Divide each charging pile into two nests, the two nests including a static optimal multi-price nest and a static optimal single-price nest;

[0014] Determining a static utility representation corresponding to each nest based on the initial utility representation of each charging pile;

[0015] A target utility representation of each charging post is determined based on the static utility representation and a total error representation of each charging post in the nest.

[0016] In one embodiment, determining the selection probability representation of each nest based on the utility of each nest includes:

[0017] Determining a selection probability representation of the static optimal multi-price nest based on a nest error representation and a maximum expected utility representation of the static optimal single-price nest;

[0018] Based on the selection probability representation of the static optimal multi-price nest, the selection probability representation of the static optimal single-price nest is obtained.

[0019] In one embodiment, determining the selection probability representation of each charging pile in the nest based on the selection probability representation of each nest includes:

[0020] Determining a selection probability representation of a static optimal multi-price charging pile based on the selection probability representation of the static optimal multi-price nest and the conditional probability;

[0021] The selection probability representation of the static optimal unit price charging pile is determined based on the selection probability representation of the static optimal unit price nest and the conditional probability.

[0022] In one embodiment, determining the target price of each charging pile at time t based on the selection probability representation, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging pile at time t by reinforcement learning includes:

[0023] Obtain the value function network based on reinforcement learning;

[0024] The target price of each charging pile at time t is determined by the value function network based on the selection probability representation, the remaining charging demand representation, the parking time representation and the charging price representation of the charging pile at time t.

[0025] In one embodiment, the training method of the value function network includes:

[0026] determining a status based on the remaining charging demand indication and the parking time indication;

[0027] Determine the action space based on the price representation of each charging station;

[0028] Obtaining a selected charging pile representation based on the selection probability representation of each charging pile by a random number method;

[0029] Determining a profit item for the charging station based on the selected charging pile representation and the charging price representation of the charging pile at time t, generating a penalty item based on the remaining charging demand and the parking time, and obtaining a reward based on the profit item and the penalty item;

[0030] generating an objective function based on the state, the action space, and the reward;

[0031] Track environmental changes through statistical data;

[0032] The objective function is trained based on environmental changes to obtain a trained value function network.

[0033] In a second aspect, the present application further provides a charging pile pricing device, the device comprising:

[0034] an initial utility determination module, configured to determine an initial utility representation of each charging pile based on a convenience representation of the charging pile, a charging price representation of the charging pile at time t, and a waiting time representation of the charging pile at time t;

[0035] a target utility determination module, configured to divide each charging pile into different nests, and determine a target utility representation of each charging pile based on the initial utility representation of each charging pile and the total error of each charging pile in the nest;

[0036] a probability determination module, configured to determine a selection probability representation of each nest based on the utility representation of each nest, and determine a selection probability representation of each charging post within the nest based on the selection probability representation of each nest;

[0037] The target price determination module is used to determine the target price based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation and the charging price representation of the charging pile at time t through reinforcement learning.

[0038] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in any one of the above embodiments when executing the computer program.

[0039] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of the method in any one of the above-mentioned embodiments when the computer program is executed by a processor.

[0040] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the steps of the method in any one of the above embodiments when executed by a processor.

[0041] The above-mentioned charging pile pricing method, device, computer equipment, computer-readable storage medium and computer program product determine the initial utility representation of each charging pile based on the convenience representation of the charging pile, the charging price representation of the charging pile at time t and the waiting time representation of the charging pile at time t; divide each charging pile into different nests, and determine the target utility representation of each charging pile based on the initial utility representation of each charging pile and the total error of each charging pile in the nest; determine the selection probability representation of each nest based on the utility representation of each nest, and determine the selection probability representation of each charging pile in the nest based on the selection probability representation of each nest; determine the target price based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation and the charging price representation of the charging pile at time t through reinforcement learning, so as to accurately price each charging pile in multiple types of charging stations. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0043] Figure 1 A diagram illustrating an application environment of a charging station pricing method according to an embodiment;

[0044] Figure 2 A flowchart of a charging station pricing method according to one embodiment is shown;

[0045] Figure 3 Schematic diagram of a charging station selection process considering user bounded rationality in one embodiment;

[0046] Figure 4 A schematic diagram of the profit of charging stations corresponding to the method of the present application in residential areas (RAs) and workplaces (WPs) in one embodiment;

[0047] Figure 5 A structural block diagram of a charging station pricing device in one embodiment;

[0048] Figure 6 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0050] The charging station pricing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. The regional management center communicates with the charging stations in each region through the network. The charging stations in each region upload real-time charging status to the regional management center. The regional management center executes the charging station pricing method of this application to obtain the target price of each charging station in the charging station. Figure 1 The figure also shows the bounded rationality of bounded rational users in selecting charging stations, which provide charging services to users. The regional management center can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0051] Multiple Single Outlet Multiple Plug (SOMP) and Single Outlet Single Plug (SOSP) charging stations are supervised by a regional management center, which is usually a public or private enterprise responsible for managing multiple charging stations within a specific geographical area. Figure 1 As shown, the management center can monitor the dynamic situation of its subordinate charging stations in real time and formulate pricing strategies based on this information. In this system, single output multi-plug (SOMP) charging stations and single output single plug (SOSP) charging stations are combined to provide charging services to users, and users will make bounded rational choices based on the service conditions of these charging stations. The goal of this application is to optimize the operating profits of these charging stations by coordinating the prices of these charging stations while ensuring that charging needs are met as much as possible. Since the charging process of a multi-output multi-plug (MOMP) charging station is the same as that of a single output single plug (SOSP) charging station, it can be regarded as a collection of multiple single output single plug (SOSP) chargers. Therefore, for the sake of simplicity, in the subsequent parts of this application, the model will not specifically discuss the case of a multi-output multi-plug (MOMP) charging station.

[0052] The charging station pricing method of this application provides a collaborative pricing method for multiple charging stations using single-output multi-plug (SOMP) charging stations. Each SOMP charging station is equipped with multiple ports, which can automatically switch to charging other electric vehicles after fully charging one electric vehicle. This reduces the impact of vehicles parked overtime and thus improves the profitability of the charging station. The optimal price for the collaborative pricing scheme for multiple charging stations (CPSMT) is determined based on a reinforcement learning method. This method uses a statistical model as an observation method to effectively compress information and maintain information dimensionality.

[0053] In an exemplary embodiment, Figure 2 As shown, a charging station pricing method is provided, which is applied to Figure 1Taking the regional management center in the example as an example, the method includes the following steps S202 to S208.

[0054] S202: Determine an initial utility representation of each charging station based on the convenience representation of the charging station, the charging price representation of the charging station at time t, and the waiting time representation of the charging station at time t.

[0055] The set of electric vehicles (EVs) arriving at time t is represented by I t The set of electric vehicles that arrive before time t and still have remaining charging needs is represented by J t For charging station z∈Z, this application will use E z,t It is defined as the charging price of electric vehicles at the charging station at time t. The waiting time of charging station z at time t is recorded as W z,t , which will be displayed in real time on the booking application of the charging station. When there is an available charging parking space, the waiting time is 0; otherwise, the waiting time is the time required to complete the existing order at the charging station.

[0056] In a charging pricing system based on multi-site coordination (CPSMT), guiding users to make balanced site selection is the key to improving system utilization. In order to effectively guide users' charging behavior, it is crucial to accurately model user utility, as this is the basis for users to make choices. In this application, the utility evaluation of user i at charging site z at time t will be affected. The factors are expressed as:

[0057]

[0058] The cost to electric vehicle users is determined by the service price E z,t and waiting time W z,t composition, they are related to the initial utility Inversely proportional. In addition, C z,i,t This term represents the convenience of charging stations, primarily influenced by the distance between the charging station and the EV user's parking destination. Because charging stations are assumed to be relatively close to each other and users are generally less likely to choose a charging station far from their parking destination, the impact of the distance the EV must travel to reach the charging station is omitted. The coefficients γ1, γ2, and γ3 in this equation represent the sensitivity of these factors.

[0059] S204: Divide each charging station into different nests, and determine a target utility representation for each charging station based on the initial utility representation of each charging station and the total error of each charging station in the nest.

[0060] In this paper, to address the correlation problem between options and relax the independent and identically distributed (IID) constraint, a hierarchical choice model is proposed. Inspired by the nested logit model, this hierarchical choice model conceptualizes the user's choice process for a charging station (CS) as a two-step process: first, selecting a nest that groups charging stations with certain common characteristics; then, selecting a specific charging station within the nest.

[0061] Therefore, the alternatives are divided into two nests, namely the static optimal multi-price (SOMP) nest and the static optimal single-price (SOSP) nest.

[0062] In an optional embodiment, each charging station is divided into different nests, and a target utility representation of each charging station is determined based on the initial utility representation of each charging station and the total error representation of each charging station in the nest, including: dividing each charging station into two nests, the two nests including a static optimal multi-price nest and a static optimal single-price nest; determining the static utility representation corresponding to each nest based on the initial utility representation of each charging station; and determining the target utility representation of each charging station based on the static utility representation and the total error representation of each charging station in the nest.

[0063] The error associated with each option can be viewed as the sum of the individual error and the nest-level error. When charging station z is a static optimal multi-price (SOMP) charging station, the initial utility simplifies to t and i are omitted. This simplification also applies to Taking into account the error term, the adjusted utility is the target utility U SOMP and U SOSP The definition is as follows:

[0064]

[0065]

[0066] and represents the total error of these options, that is, the sum of individual errors and nested errors, which obey the independent and identically distributed Gumbel distribution with parameters (0,θ0). ∈ SOSP and ∈ SOMP They follow Gumbel distributions with parameters (0, θ1) and (0, θ2), respectively. θ1 and θ2 are scale parameters ranging from 1 to ∞, representing the correlation between the options within the two nests. Furthermore, higher values ​​of θ1 and θ2 indicate a greater degree of substitution between the options within the nest, reflecting greater similarity or interchangeability between the options.

[0067] S206: Determine a selection probability representation for each nest based on the utility representation of each nest, and determine a selection probability representation for each charging station within the nest based on the selection probability representation of each nest.

[0068] In reality, electric vehicles (EVs) do not always choose the option with the highest utility, as decision-making inertia or personal preferences may affect their choices, especially when the utility difference is minimal. This extremely small difference interval is called the "indifference interval". Figure 3 As shown, Figure 3 Figure 1 is a schematic diagram of a charging station selection process that considers bounded user rationality in one embodiment. Because each SOMP charging station can only supply power through one port at a time, electric vehicles may need to wait for charging. This may cause users to question the reliability of SOMP CSs, thereby increasing their preference for SOSP CSs.

[0069] To this end, a selection probability representation of each nest is determined based on the utility representation of each nest, and a selection probability representation of each charging station within the nest is determined based on the selection probability representation of each nest.

[0070] In one of the optional embodiments, the selection probability representation of each nest is determined based on the utility of each nest, including: determining the selection probability representation of the static optimal multi-price nest based on the indifference interval boundary and the maximum expected utility representation of each nest; and obtaining the selection probability representation of the static optimal single-price nest based on the selection probability representation of the static optimal multi-price nest.

[0071] In one of the optional embodiments, the selection probability representation of each charging station in the nest is determined based on the selection probability representation of each nest, including: determining the selection probability representation of the static optimal multi-price charging station based on the selection probability representation of the static optimal multi-price nest and the conditional probability; determining the selection probability representation of the static optimal single-price charging station based on the selection probability representation of the static optimal single-price nest and the conditional probability.

[0072] The selection probability of the static optimal multi-price nest is expressed as:

[0073]

[0074] where Δ is the boundary of the indifference interval and τ is the probability that the user chooses the static optimal multi-price charging station (SOMP CS) within the interval. and represents the utility of the nest, which is determined by the maximum utility within the nest. In order to calculate the utility of the nest, it is divided into two parts, namely the nest error and the other part, because the nest error is common to all options in the nest, and its expression is as follows:

[0075]

[0076] According to the maximum stability property of Gumbel distribution, is the maximum expected utility of the SOSP nest, which can be expressed as:

[0077]

[0078] To do this, you can define a new variable It obeys the Gumbel distribution with parameters (0,θ3), and after derivation, we can get thereby

[0079]

[0080] Then the probability of choosing SOSP nest is P(NEST SOSP )=1-P(NEST SOMP ).

[0081] The probability of choosing an option within a nest can be obtained using conditional probability.

[0082] P(SOSP)=P(NEST SOSP )·P(SOSP|NEST SOSP ), can be calculated by the following two formulas

[0083]

[0084]

[0085] S208: Determine the target price of each charging station at time t through reinforcement learning based on the selection probability representation of each charging station, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging station at time t.

[0086] In one of the optional embodiments, the target price of each charging station at time t is determined based on the selection probability representation, remaining charging demand representation, parking time representation and charging price representation of each charging station at time t through reinforcement learning, including: obtaining a value function network obtained based on reinforcement learning; determining the target price of each charging station at time t based on the selection probability representation, remaining charging demand representation, parking time representation and charging price representation of the charging station at time t through the value function network.

[0087] This application proposes a method for deriving the optimal decision for a multi-site collaborative charging pricing system (CPSMT) based on deep reinforcement learning. The decision-making process is described using states, actions, rewards, and objective functions. The objective function is the value function corresponding to the value function network. The value function network determines the target price of each charging station at time t based on the selection probability representation, remaining charging demand representation, parking time representation, and charging price representation of the charging station at time t.

[0088] Among them, the training method of the value function network includes: determining the state based on the remaining charging demand representation and the parking time representation; determining the action space based on the price representation of each charging station; obtaining the selected charging station representation based on the selection probability representation of each charging station through a random number method; determining the profit item of the charging station based on the selected charging station representation and the charging price representation of the charging station at time t, and generating a penalty item based on the remaining charging demand and the parking time, and obtaining a reward based on the profit item and the penalty item; generating an objective function based on the state, action space and reward; tracking environmental changes through statistical data; and training the objective function based on environmental changes to obtain a trained value function network.

[0089] Among them, the state can be expressed as:

[0090]

[0091] and is the remaining charging demand and parking time. The state changes as vehicles arrive and leave and charging progresses.

[0092] The action space is represents the price of all charging stations. When the algorithm makes a decision, it means that the actual price for each charging station is determined.

[0093] The reward is expressed as

[0094] The first term represents the profit of the charging station, ρ t is the real-time price of electricity purchased by the charging station (CS), z * is the charging station selected by electric vehicle i according to the probability described above. In order to simulate the user's choice of charging station, random numbers are used. As shown below,

[0095] z * =z,if RV∈I z

[0096] If the random variable RV falls within a certain interval, it indicates the user’s choice of a charging station.z According to the method described above, based on the selection probability P z To confirm, P z is the probability of selecting charging station z, which is expressed as follows:

[0097]

[0098] Assuming charging demand Related to the real-time price, it can be expressed as d i,z =σ1-σ2E z,t .e t is the total charging power of all charging stations in the time interval t, where σ1 and σ2 are the corresponding parameters representing the real-time price, which are used to obtain the charging demand.

[0099] The second term is a virtual penalty term to ensure that the priority of the scheme is to complete charging rather than to obtain profit. This penalty can be quantified as follows:

[0100] M is the penalty factor, and and They represent the remaining parking time and remaining charging demand of electric vehicle i in the next time interval t+1 respectively.

[0101] The action-value function Q(s,a) represents the expected total reward when taking action a in state s and is defined as follows:

[0102] where γ is the attenuation factor.

[0103] Q(o,a;θ)≈Q π (s,a)

[0104] o is the observed value of the system at time t. Since the number of EVs with various charging requirements is considerable and fluctuates, we use statistical data to track the changes in the environment.

[0105]

[0106] and are the average remaining parking time and average charging demand of the charging pile in each charging station (CS) at time t. and represents the cumulative remaining parking time and cumulative charging demand of electric vehicles waiting in each charging station at time t. is the number of available charging ports at each charging station at time t. It is worth emphasizing that in the static optimal multi-price (SOMP) scenario, each charger is equipped with multiple charging ports. Therefore, the remaining charging demand and remaining parking time of each charger are calculated by adding the charging demand and parking time of all EVs connected to its ports.

[0107] By interacting with the environment, the Deep Q Network (DQN) gathers experience to train its value function network, improving its ability to accurately estimate Q values. This optimized estimate enables the agent to identify and execute optimal actions to achieve the highest possible cumulative reward.

[0108] Among them, combined Figure 4 As shown, Figure 4 The figure is a schematic diagram of the profit of charging stations corresponding to the method of the present application in residential areas (RAs) and workplaces (WPs) in one embodiment, wherein experiments were carried out in residential areas (RAs) and workplaces (WPs), and the corresponding electric vehicle arrival distribution parameters were (75.802, 6.195, -282.236) and (0.829, 7.667, 7.0139) respectively. Figure 4 , Figure 4 The results show that the use of single-output multi-plug charging equipment has greatly improved the profit of charging stations. RA refers to residential areas, WP refers to workplaces, and bench refers to the situation where only single-output single-plug charging equipment is used.

[0109] In the benchmark experiment, the static optimal single price (SOSP) charging station was used to replace the static optimal multi-price (SOMP) charging station, while keeping the number of charging piles unchanged, and the probability of each EV choosing the nearest charging station was also adjusted accordingly to match the capacity of the charging station.

[0110] Figure 4 The results show that the benchmark experiment achieves significantly lower rewards than a system integrating static optimal multi-price (SOMP) charging stations, with many negative rewards occurring due to the severe penalty imposed on uncompleted charging requests. The problem of users staying overtime leads to significant waste of resources, and the overall charging capacity in the benchmark experiment is significantly reduced compared to a system integrating static optimal multi-price (SOMP) charging stations. This highlights the effectiveness of static optimal multi-price (SOMP) charging stations in mitigating the negative impact of overstays. Furthermore, because the power output of static optimal multi-price (SOMP) chargers is comparable to that of static optimal single-price (SOSP) chargers, they improve customer flow without placing additional strain on the local power grid.

[0111] Furthermore, the profitability of charging stations in residential areas is approximately 2.1 times that of workplace charging stations. Electric vehicles tend to arrive at workplaces more frequently than those returning home, making it challenging for charging stations to provide timely charging services. This results in significant penalties for unmet demand. Furthermore, charging times for electric vehicles in residential areas more often coincide with periods of low electricity prices, while workplace charging demand peaks during the day. This reduces the costs of charging stations in residential areas and allows for more time to meet charging demand, paving the way for higher profit margins.

[0112] The above-mentioned charging station pricing method determines the initial utility representation of each charging station based on the convenience representation of the charging station, the charging price representation of the charging station at time t, and the waiting time representation of the charging station at time t; divides each charging station into different nests, and determines the target utility representation of each charging station based on the initial utility representation of each charging station and the total error of each charging station in the nest; determines the selection probability representation of each nest based on the utility representation of each nest, and determines the selection probability representation of each charging station within the nest based on the selection probability representation of each nest; and determines the target price based on the selection probability representation of each charging station, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging station at time t through reinforcement learning, which can accurately price each charging station among multiple types of charging stations.

[0113] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0114] Based on the same inventive concept, embodiments of the present application also provide a charging station pricing device for implementing the aforementioned charging station pricing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more charging station pricing device embodiments provided below can be found in the limitations of the charging station pricing method described above and will not be further elaborated here.

[0115] In an exemplary embodiment, Figure 5As shown, a charging station pricing device is provided, including: an initial utility determination module 501, a target utility determination module 502, a probability determination module 503 and a target price determination module 504, wherein:

[0116] An initial utility determination module 501 is configured to determine an initial utility representation of each charging pile based on a convenience representation of the charging pile, a charging price representation of the charging pile at time t, and a waiting time representation of the charging pile at time t;

[0117] a target utility determination module 502 for grouping the charging piles into different nests and determining a target utility representation for each charging pile based on the initial utility representation of each charging pile and the total error of each charging pile in the nest;

[0118] a probability determination module 503 for determining a selection probability representation of each nest based on the utility representation of each nest, and determining a selection probability representation of each charging post within the nest based on the selection probability representation of each nest;

[0119] The target price determination module 504 is configured to determine a target price based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging pile at time t through reinforcement learning.

[0120] In one of the optional embodiments, the target utility determination module 502 is specifically used to divide each charging pile into two nests, the two nests including a static optimal multi-price nest and a static optimal single-price nest; determine the static utility representation corresponding to each nest based on the initial utility representation of each charging pile; and determine the target utility representation of each charging pile based on the static utility representation and the total error representation of each charging pile in the nest.

[0121] In one of the optional embodiments, the above-mentioned probability determination module 503 is specifically used to determine the selection probability representation of the static optimal multi-price nest based on the nest error representation and the maximum expected utility representation of the static optimal single-price nest; and obtain the selection probability representation of the static optimal single-price nest based on the selection probability representation of the static optimal multi-price nest.

[0122] In one of the optional embodiments, the above-mentioned probability determination module 503 is specifically used to determine the selection probability representation of the static optimal multi-price charging pile based on the selection probability representation of the static optimal multi-price nest and the conditional probability; and determine the selection probability representation of the static optimal single-price charging pile based on the selection probability representation of the static optimal single-price nest and the conditional probability.

[0123] In one of the optional embodiments, the target price determination module 504 is specifically used to obtain a value function network obtained based on reinforcement learning; the target price of each charging pile at time t is determined through the value function network based on the selection probability representation, remaining charging demand representation, parking time representation and charging price representation of each charging pile at time t.

[0124] In one of the optional embodiments, the target price determination module 504 is specifically used to determine the state based on the remaining charging demand representation and the parking time representation; determine the action space based on the price representation of each charging pile; obtain the selected charging pile representation based on the selection probability representation of each charging pile in a random number manner; determine the profit item of the charging station based on the selected charging pile representation and the charging price representation of the charging pile at time t, and generate a penalty item based on the remaining charging demand and the parking time, and obtain a reward based on the profit item and the penalty item; generate an objective function based on the state, action space and the reward; track environmental changes through statistical data; and train the objective function based on environmental changes to obtain a trained value function network.

[0125] Each module in the charging station pricing device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor within a computer device in hardware form, or may be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.

[0126] In an exemplary embodiment, a computer device is provided. The computer device may be a server. The server corresponds to the regional management center mentioned above. The internal structure diagram thereof may be as follows: Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store corresponding data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a charging site pricing method is implemented.

[0127] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0128] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0129] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0130] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0131] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0132] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.

[0133] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0134] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A charging pile pricing method, characterized in that: The method comprises: Determine an initial utility representation of each charging pile based on a representation of the convenience level of the charging pile, a representation of the charging price of the charging pile at time t, and a representation of the waiting time of the charging pile at time t; Dividing each charging pile into different nests, and determining a target utility representation for each charging pile based on an initial utility representation of each charging pile and a total error of each charging pile in the nest; Determining a selection probability representation of each nest based on the utility representation of each nest, and determining a selection probability representation of each charging post in the nest based on the selection probability representation of each nest; The target price of each charging pile at time t is determined by reinforcement learning based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging pile at time t.

2. The method according to claim 1, characterized in that The dividing the charging piles into different nests and determining the target utility representation of each charging pile based on the initial utility representation of each charging pile and the total error representation of each charging pile in the nest includes: Divide each charging pile into two nests, the two nests including a static optimal multi-price nest and a static optimal single-price nest; Determining a static utility representation corresponding to each nest based on the initial utility representation of each charging pile; A target utility representation of each charging post is determined based on the static utility representation and a total error representation of each charging post in the nest.

3. The method according to claim 2, characterized in that The determining of the selection probability representation of each nest based on the utility of each nest comprises: Determining a selection probability representation of the static optimal multi-price nest based on a nest error representation and a maximum expected utility representation of the static optimal single-price nest; Based on the selection probability representation of the static optimal multi-price nest, the selection probability representation of the static optimal single-price nest is obtained.

4. The method according to claim 3, characterized in that The determining the selection probability representation of each charging pile in the nest based on the selection probability representation of each nest includes: Determining a selection probability representation of a static optimal multi-price charging pile based on the selection probability representation of the static optimal multi-price nest and the conditional probability; The selection probability representation of the static optimal unit price charging pile is determined based on the selection probability representation of the static optimal unit price nest and the conditional probability.

5. The method according to any one of claims 1 to 4, characterized in that The method of determining the target price of each charging pile at time t based on the selection probability representation, the remaining charging demand representation, the parking time representation, and the charging price representation of the charging pile at time t by reinforcement learning includes: Obtain the value function network based on reinforcement learning; The target price of each charging pile at time t is determined by the value function network based on the selection probability representation, the remaining charging demand representation, the parking time representation and the charging price representation of the charging pile at time t.

6. The method according to claim 5, characterized in that The training method of the value function network includes: determining a status based on the remaining charging demand indication and the parking time indication; Determine the action space based on the price representation of each charging station; Obtaining a selected charging pile representation based on the selection probability representation of each charging pile by a random number method; Determining a profit item for the charging station based on the selected charging pile representation and the charging price representation of the charging pile at time t, generating a penalty item based on the remaining charging demand and the parking time, and obtaining a reward based on the profit item and the penalty item; generating an objective function based on the state, the action space, and the reward; Track environmental changes through statistical data; The objective function is trained based on environmental changes to obtain a trained value function network.

7. A charging pile pricing device, characterized in that: The device comprises: an initial utility determination module, configured to determine an initial utility representation of each charging pile based on a convenience representation of the charging pile, a charging price representation of the charging pile at time t, and a waiting time representation of the charging pile at time t; a target utility determination module, configured to divide each charging pile into different nests, and determine a target utility representation of each charging pile based on the initial utility representation of each charging pile and the total error of each charging pile in the nest; a probability determination module, configured to determine a selection probability representation of each nest based on the utility representation of each nest, and determine a selection probability representation of each charging post within the nest based on the selection probability representation of each nest; The target price determination module is used to determine the target price based on the selection probability representation of each charging pile, the remaining charging demand representation, the parking time representation and the charging price representation of the charging pile at time t through reinforcement learning.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.