Charging scheduling method and device based on deep reinforcement learning, equipment and medium
By optimizing charging scheduling through deep reinforcement learning, the problem of high computational pressure in large-scale vehicle charging scheduling by traditional methods is solved, and more efficient charging equipment scheduling is achieved.
Patent Information
- Application Number
- CN202410720423.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-05
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2044-06-05
AI Technical Summary
Traditional operations research optimization methods face high computational burdens in large-scale vehicle charging scheduling, making it difficult to find the optimal decision solution within a limited time, resulting in low scheduling efficiency of charging equipment.
By employing deep reinforcement learning, charging demand and equipment resource information are obtained through a pre-trained model, target charging equipment is selected and charging actions are determined, and a charging sequence is generated to optimize charging scheduling.
It improves the scheduling efficiency of charging equipment, maximizes the output of electricity or serves more vehicles, and optimizes the charging process.
Smart Images

Figure CN118735722B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and in particular to a charging scheduling method and device based on deep reinforcement learning, equipment and medium. BACKGROUND
[0002] At present, new energy vehicles are developing rapidly, especially electric vehicles driven by electric energy. However, as the number of electric vehicles increases, users have higher requirements for the timeliness of electric vehicle charging.
[0003] In related technologies, an operational optimization method is used to select an optimal decision scheme in a vehicle discrete decision space to determine the matching relationship between a plurality of intelligent charging devices and a plurality of vehicles to be charged. However, as the vehicle discrete decision space continues to expand and real-time solving personalization requirements, the traditional operational optimization method faces huge computing pressure and it is difficult to accurately find the optimal decision scheme within a limited time. That is, when there are a large number of vehicles to be charged, the related art has the problem of poor charging scheduling efficiency of intelligent charging devices. SUMMARY
[0004] The main purpose of the embodiments of the present application is to provide a charging scheduling method and device based on deep reinforcement learning, which can improve the charging scheduling efficiency of intelligent charging robots.
[0005] To achieve the above purpose, the first aspect of the embodiments of the present application provides a charging scheduling method based on deep reinforcement learning, the method comprising:
[0006] obtaining charging demand information of a plurality of candidate vehicles in a target area and device resource information of a plurality of candidate charging devices;
[0007] inputting the charging demand information and the device resource information into a pre-trained deep reinforcement learning model, when the candidate vehicles are not selected, selecting a target charging device from the plurality of candidate charging devices, and determining a target charging action corresponding to the target charging device, wherein the target charging action comprises one of inward charging and outward charging;
[0008] If the target charging action is outward charging, based on the current charging demand information and the device resource information, a corresponding target vehicle is selected for the target charging device from the plurality of remaining candidate vehicles, and the corresponding device resource information and charging demand information are updated;
[0009] until all candidate vehicles are selected, based on the selection order of a plurality of target charging actions under each target charging device, a corresponding target charging sequence is generated;
[0010] Based on each target charging sequence, the corresponding target charging equipment is scheduled to perform the target charging action in the selected order to complete the charging of the target vehicle.
[0011] In some embodiments, the device resource information includes the current remaining battery level, the current cumulative operating time, and the current charging sequence, wherein the current charging sequence is a subset of the target charging sequence;
[0012] Select the target charging device from a list of candidate charging devices, including:
[0013] Based on the preset first decoder parameters, all current residual power values and current cumulative working time are linearly transformed and activated to obtain the device status characteristics;
[0014] A high-dimensional mapping sequence is generated based on the current charging sequence corresponding to each candidate charging device. Based on the preset second decoder parameters, all high-dimensional mapping sequences are linearly transformed and activated to obtain sequence state features.
[0015] The result of concatenating the device state features and sequence state features is subjected to linear transformation and activation processing to obtain the device selection probability value corresponding to the candidate charging device. The target charging device is selected from multiple candidate charging devices based on the device selection probability value.
[0016] In some embodiments, generating a high-dimensional mapping sequence based on the current charging sequence corresponding to each candidate charging device includes:
[0017] Based on the preset encoder parameters, the charging demand information corresponding to each candidate vehicle and the equipment resource information corresponding to all candidate charging devices are spliced together to obtain the initial features;
[0018] Using a pre-defined first attention mechanism as a constraint, feature enhancement is performed on the initial features to obtain updated initial features;
[0019] The updated initial features are processed by feedforward propagation to obtain the high-dimensional mapping features corresponding to the candidate vehicles.
[0020] For each candidate charging device, the corresponding high-dimensional mapping features are sequentially concatenated to obtain a high-dimensional mapping sequence, based on the selection order of each candidate vehicle indicated in the current charging sequence.
[0021] In some embodiments, selecting a corresponding target vehicle for a target charging device based on current charging demand information and device resource information includes:
[0022] By fusing all high-dimensional mapping features, global feature information is obtained;
[0023] Determine the device resource information corresponding to the target charging device, and use a preset second attention mechanism as a constraint to perform feature enhancement on the result of splicing the device resource information, global feature information and high-dimensional mapping sequence to obtain vehicle state features;
[0024] The vehicle state features are extracted and activated to obtain the vehicle selection probability value corresponding to the candidate vehicles. Based on the vehicle selection probability value, the target vehicle corresponding to the target charging device is determined from multiple candidate vehicles.
[0025] In some embodiments, the charging demand information includes charging start and end values and dwell time;
[0026] Update the corresponding device resource information, including:
[0027] Obtain the equipment resource information corresponding to the currently selected target charging equipment, as well as the charging demand information corresponding to the target vehicle;
[0028] The power difference is calculated based on the current remaining power value and the start and end values of charging, and the total time value is calculated based on the current cumulative working time and dwell time.
[0029] Update the current remaining battery level based on the battery level difference, update the current cumulative working time based on the total time value, and update the current charging sequence based on the selected order of the target vehicles.
[0030] In some embodiments, the deep reinforcement learning model is trained through the following steps:
[0031] Obtain sample charging demand information for multiple candidate vehicles within the sample area, as well as sample equipment resource information for multiple candidate charging devices;
[0032] The sample charging demand information and sample equipment resource information are input into the deep reinforcement learning model. When the sample candidate vehicles are not all selected, the sample target charging equipment is selected from multiple sample candidate charging equipment, and the sample target charging action corresponding to the sample target charging equipment is determined. The sample target charging action includes one of inward charging and outward charging.
[0033] If the target charging action of the sample is to charge externally, select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current sample charging demand information and sample equipment resource information, and update the corresponding sample equipment resource information and sample charging demand information.
[0034] Until all candidate vehicles have been selected, a corresponding target charging sequence is generated based on the selection order of multiple target charging actions under each target charging device.
[0035] The sample reward value is determined based on the sample target charging sequence. The parameters of the deep reinforcement learning model are adjusted based on the sample reward value to obtain the trained deep reinforcement learning model.
[0036] In some embodiments, the sample target charging device in the sample target charging sequence is determined based on the sample device selection probability value, each sample target vehicle corresponding to the sample target charging device is determined based on the sample vehicle selection probability value, and the sample reward value includes a first sample reward value and a second sample reward value.
[0037] The sample reward value is determined based on the sample target charging sequence, including:
[0038] If the target charging device is determined based on the maximum selection probability value of the target device, and each target vehicle corresponding to the target charging device is determined based on the maximum selection probability value of the target vehicle, the first number of target vehicles matched with each target charging device and the first amount of charge released by each target vehicle are determined based on the target charging sequence.
[0039] The first sample reward value is obtained based on the total value of all first charging quantities and / or the total value of first charging power.
[0040] If the sample target charging device is determined based on the random sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on the random sample vehicle selection probability value, the second charging quantity of each sample target charging device and the second charging amount released by each sample target vehicle are determined based on the sample target charging sequence.
[0041] The second sample reward value is obtained based on the total value of all second charge quantities and / or the total value of second charge amounts.
[0042] To achieve the above objectives, a second aspect of this application proposes a charging scheduling device based on deep reinforcement learning, the device comprising:
[0043] The acquisition module is used to acquire charging demand information of multiple candidate vehicles and equipment resource information of multiple candidate charging devices within the target area.
[0044] The target charging action determination module is used to input charging demand information and equipment resource information into a pre-trained deep reinforcement learning model. When the candidate vehicles are not all selected, the target charging device is selected from multiple candidate charging devices, and the target charging action corresponding to the target charging device is determined. The target charging action includes one of inward charging and outward charging.
[0045] The target vehicle determination module is used to select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current charging demand information and equipment resource information if the target charging action is external charging, and update the corresponding equipment resource information and charging demand information.
[0046] The target charging sequence determination module is used to generate a corresponding target charging sequence based on the selection order of multiple target charging actions under each target charging device until all candidate vehicles have been selected.
[0047] The target execution module is used to schedule the corresponding target charging equipment to perform target charging actions in a selected order based on each target charging sequence, so as to complete the charging of the target vehicle.
[0048] To achieve the above objectives, a third aspect of the present application proposes an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the deep reinforcement learning-based charging scheduling method of the first aspect described above.
[0049] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the deep reinforcement learning-based charging scheduling method of the first aspect described above.
[0050] This application proposes a charging scheduling method, apparatus, device, and medium based on deep reinforcement learning. The method first acquires charging demand information for multiple candidate vehicles and equipment resource information for multiple candidate charging devices within a target area. Next, the charging demand information and equipment resource information are input into a pre-trained deep reinforcement learning model. When not all candidate vehicles have been selected, a target charging device is selected from the multiple candidate charging devices, and the corresponding target charging action is determined. The target charging action includes either inward charging or outward charging. By simultaneously considering both inward and outward charging actions, the target charging device can serve as many target vehicles as possible. Furthermore, if the target charging action is outward charging, a corresponding target vehicle is selected from the remaining candidate vehicles based on the current charging demand information and equipment resource information, and the corresponding equipment resource information and charging demand information are updated. This process continues until all candidate vehicles have been selected. Based on the selection order of multiple target charging actions under each target charging device, a corresponding target charging sequence is generated. Based on each target charging sequence, the corresponding target charging device is scheduled to execute the target charging action in the selected order to complete the charging of the target vehicles. With the aim of maximizing power output or serving as many vehicles as possible, a target charging sequence containing two types of charging actions is generated. This allows for better scheduling of each charging device based on the target charging sequence, thereby improving the charging scheduling efficiency of the charging devices. Attached Figure Description
[0051] Figure 1 This is a schematic diagram illustrating an application scenario of the charging scheduling device based on deep reinforcement learning provided in the embodiments of this application;
[0052] Figure 2 This is an optional flowchart of the charging scheduling method based on deep reinforcement learning provided in the embodiments of this application;
[0053] Figure 3 This is a schematic diagram of an optional functional module of the charging scheduling device based on deep reinforcement learning provided in the embodiments of this application;
[0054] Figure 4 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0056] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0057] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0058] First, let's analyze some of the terms used in this application:
[0059] Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0060] New energy vehicles refer to automobiles that use unconventional vehicle fuels as their power source (or use conventional vehicle fuels but employ new onboard power devices), integrating advanced technologies in vehicle power control and drive, resulting in vehicles with advanced technical principles and new technologies and structures. New energy vehicles include pure electric vehicles, range-extended electric vehicles, hybrid electric vehicles, fuel cell electric vehicles, and hydrogen engine vehicles.
[0061] Currently, new energy vehicles are developing rapidly, especially electric vehicles powered by electricity. However, as the number of electric vehicles increases, users are placing higher demands on the timeliness of charging.
[0062] In related technologies, operations research optimization methods are used to select the optimal decision scheme in the vehicle discrete decision space to determine the matching relationship between multiple smart charging devices and multiple vehicles to be charged. However, with the continuous expansion of the vehicle discrete decision space and the personalized requirements of real-time solution, traditional operations research optimization methods face huge computational pressure and find it difficult to accurately find the optimal decision scheme within a limited time. That is, when there are a large number of vehicles to be charged, the related technologies have the problem of poor charging scheduling efficiency for smart charging devices.
[0063] Based on this, this application proposes a charging scheduling method, device, equipment and medium based on deep reinforcement learning, which can improve the charging scheduling efficiency of intelligent charging robots.
[0064] It should be noted first that in this application embodiment, when information such as the user's vehicle charging demand information or user geographical location information is required, the user's permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when this application embodiment needs to obtain the user's sensitive personal information, the user's individual permission or consent will be obtained first. Only after obtaining the user's individual permission or consent will the necessary data for the normal operation of this application embodiment be obtained. For example, in order to charge vehicles parked in a target area, the deep reinforcement learning-based charging scheduling device in this application needs to obtain the geographical location of each vehicle and its corresponding charging demand information. The consent of the relevant user will be obtained first, and the charging demand information required by this application embodiment will only be obtained after obtaining the user's permission or consent.
[0065] The charging scheduling method, apparatus, device, and medium based on deep reinforcement learning provided in this application are specifically illustrated through the following embodiments. First, the application scenarios of the charging scheduling apparatus based on deep reinforcement learning in this application embodiment are described, such as... Figure 1 As shown, Figure 1This is a schematic diagram of an application scenario for a charging scheduling device based on deep reinforcement learning provided in this application embodiment. The charging scheduling method based on deep reinforcement learning provided in this application embodiment (hereinafter referred to as "charging scheduling method" for ease of description) can be applied to a charging scheduling device based on deep reinforcement learning (hereinafter referred to as "charging scheduling device" for ease of description). In one application scenario, the target area includes candidate vehicles A, B, and C waiting for charging. These three candidate vehicles send their respective charging demand information to the charging scheduling device. At the same time, the charging scheduling device obtains the equipment resource information corresponding to the candidate charging equipment. Then, the charging scheduling device generates a target charging sequence based on the charging demand information and equipment resource information. The target charging sequence indicates that the selected candidate charging equipment is the target charging equipment and indicates the charging action of the target charging equipment. The charging action includes inward charging and outward charging. When the charging action is outward charging, the target charging sequence represents that the target charging equipment charges the selected candidate vehicles in a certain charging order. In this way, each target charging device can perform the corresponding charging action based on the target charging sequence, outputting electricity to the maximum extent and meeting the charging needs of each candidate vehicle in the target area.
[0066] Next, the charging scheduling method based on deep reinforcement learning in the embodiments of this application will be specifically described through the following examples.
[0067] In this embodiment, the charging scheduling device will be described from the perspective of the charging scheduling device itself. The charging scheduling device can be integrated into a computer device, such as a server. Figure 2 As shown, Figure 2 This is an optional flowchart of the charging scheduling method based on deep reinforcement learning provided in the embodiments of this application. Figure 2 The method may include, but is not limited to, the following steps 101 to 105. When the charging scheduling device executes the charging scheduling method, the specific process is as follows. It should be noted that this embodiment... Figure 2 The order of steps 101 to 105 is not specifically limited. The order of steps can be adjusted or some steps can be reduced or added according to actual needs.
[0068] Step 101: Obtain charging demand information for multiple candidate vehicles and equipment resource information for multiple candidate charging devices within the target area.
[0069] Step 101 will be described in detail below.
[0070] In some embodiments, in order to determine the charging order, charging value, and charging time of multiple candidate vehicles in the target area by the selected charging equipment, the charging scheduling device first needs to obtain the charging demand information of multiple candidate vehicles in the target area, as well as the equipment resource information corresponding to each candidate charging equipment.
[0071] The size of the target area can be preset by the charging scheduling device. For example, a parking lot includes two areas, area A and area B. The charging scheduling device can obtain the charging demand information of each candidate vehicle a in area A and then generate a target charging sequence a. The target charging sequence a indicates that the selected candidate charging equipment will serve each candidate vehicle a according to the target charging sequence a. The charging scheduling device can also obtain the charging demand information of each candidate vehicle b in area B and then generate a target charging sequence b. The target charging sequence b indicates that the selected candidate charging equipment will serve each candidate vehicle b according to the target charging sequence b.
[0072] Among them, the candidate vehicles (EVs) are electrically powered. After obtaining the consent of the relevant users, the charging scheduling device usually also obtains the geographical location of the candidate vehicles (hereinafter referred to as "vehicles" for ease of description) so that the target charging equipment can find the corresponding target vehicle and charge it based on the geographical location. There are usually multiple candidate vehicles. When there are multiple candidate vehicles to be charged in the target area, the charging scheduling device can schedule a limited number of candidate charging equipment (hereinafter referred to as "charging equipment" for ease of description) based on the obtained information to maximize the output of electricity or maximize the number of vehicles serving the target area. Of course, there can also be only one candidate vehicle. In this case, the charging scheduling device can select the most suitable one from the multiple candidate charging equipment to charge it based on the obtained information.
[0073] The charging demand information includes charging start and end values and dwell time. The charging start and end values indicate the current remaining battery power of the corresponding vehicle and the expected battery level after charging is completed. For example, the charging start and end values for vehicle A are [20, 100], which indicates that vehicle A currently has 20% remaining battery power and expects to reach full charge after the charging equipment completes charging. Alternatively, the charging start and end values can indicate the expected input of battery power for the corresponding vehicle. For example, when the charging start and end value is 20%, it means that 20% of the electrical energy needs to be charged to the target vehicle.
[0074] The candidate charging device can be an Intelligent Charging Robot (ICR), a robot with autonomous navigation and intelligent charging capabilities, typically designed to provide convenient charging services for electric vehicles, mobile devices, or robots. The ICR can be vehicle-mounted or flying, depending on the specific requirements. Regardless of the type, it can complete the charging operation for candidate vehicles within the target area based on the target charging sequence output by the charging scheduling device provided in this application embodiment.
[0075] The equipment resource information includes the current remaining power value, the current cumulative working time, and the current charging sequence. Furthermore, the current remaining power value indicates the amount of electricity remaining in the charging equipment. Based on the specific value of the current remaining power value, the charging equipment will take different actions. It may charge the corresponding target vehicle according to its charging demand information. Alternatively, if charging the vehicle is not the best option for the charging equipment at this time, the charging equipment will first perform self-charging and then provide service to the vehicle to maximize the satisfaction of the charging needs of each vehicle.
[0076] Furthermore, the scheduling conditions that need to be met in the embodiments of this application are explained as follows:
[0077] (1) Each ICR can only charge EVs whose charging demand does not exceed the current remaining power of the ICR;
[0078] (2) The charging completion time for each EV should not exceed the dwell time of the EV;
[0079] (3) Each EV can only be served by one ICR once. Once the charging process starts, it cannot be interrupted in order to improve the overall service efficiency and service quality of the charging equipment and avoid the waste of computing resources caused by scheduling and planning the charging equipment.
[0080] (4) ICR can self-charge during working intervals until the maximum battery capacity E is reached;
[0081] (5) During the scheduling process, the time and power consumption required for ICR to move are ignored.
[0082] Step 102: Input the charging demand information and equipment resource information into the pre-trained deep reinforcement learning model. When the candidate vehicles are not all selected, select the target charging device from multiple candidate charging devices and determine the target charging action corresponding to the target charging device. The target charging action includes one of internal charging and external charging.
[0083] Step 102 is described in detail below.
[0084] In some embodiments, the charging scheduling device may be equipped with a pre-trained deep reinforcement learning model. After obtaining charging demand information and equipment resource information, these two are input into the pre-trained deep reinforcement learning model. Since the purpose of this solution is to charge all candidate vehicles in the target area, the model first determines whether each candidate vehicle has a corresponding charging device for charging. When there are still candidate vehicles that have not been selected, it first selects one of the multiple candidate charging devices as the target charging device and determines the target charging action corresponding to the target charging device.
[0085] The target charging action includes either internal charging or external charging. Internal charging refers to the charging device's self-charging operation, that is, inputting electricity into the charging device. Internal charging indicates that the charging device's current remaining power is too low to charge the vehicle, or that the charging device can achieve the maximum charging effect by self-charging first and then servicing the vehicle. External charging refers to the charging device releasing electricity to the corresponding vehicle to charge it. When the target charging action is external charging, it is necessary to determine the number of vehicles that the target charging device needs to serve and the service order of each vehicle, so that the target charging device can charge the corresponding vehicles in sequence according to the charging demand information.
[0086] The maximum charging effect refers to the maximum total power output of the charging equipment within a certain period of time, or the maximum number of vehicles that the charging equipment can serve within a certain period of time. Specific service indicators can be set according to actual conditions, and this application embodiment does not make specific requirements.
[0087] It should be noted that the target charging sequence formed by the relevant technology based on the charging demand information of the candidate vehicle and the equipment resource information of the candidate charging equipment only indicates the order in which the charging equipment serves the vehicle. That is, the target charging action of the charging equipment in the relevant technology only involves charging externally. As a result, the relevant technology cannot achieve the maximum charging effect.
[0088] For example, Figure 1The scenario includes charging device A and charging device B, vehicle A, vehicle B, and vehicle C. Charging device A and charging device B both have a charging power of 6 kW / h. Charging device A has: current remaining power of 10 kWh and accumulated working time of 0; charging device B has: current remaining power of 4 kWh and accumulated working time of 0; vehicle A has: charging start / end value of 5 kWh and dwell time of 90 minutes; vehicle B has: charging start / end value of 7 kWh and dwell time of 120 minutes; and vehicle C has: charging start / end value of 2 kWh and dwell time of 60 minutes. In this scenario, under the relevant scheduling conditions proposed in the embodiments of this application, the target charging sequence given by the relevant technology is charging device A-vehicle B (service time 70min) and charging device B-vehicle C (service time 20min). However, the target charging sequence given by the charging scheduling method in the embodiments of this application is charging device A-vehicle C (service time 20min)-vehicle B (service time 70min) and charging device B-self-charging (self-charging time 40min)-vehicle A (service time 50min). When the desired amount of electricity is charged to each vehicle during its dwell time, the total amount of electricity output by the relevant technology is 9kWh, while the total amount of electricity output by the embodiments of this application is 14kWh. That is, the charging scheme of the embodiments of this application is superior.
[0089] In some embodiments, selecting a target charging device from a plurality of candidate charging devices includes the following steps 201 to 203:
[0090] Step 201: Based on the preset first decoder parameters, perform linear transformation and activation processing on all current residual power values and current cumulative working time to obtain device status characteristics.
[0091] Step 201 is described in detail below.
[0092] In some embodiments, the deep reinforcement learning model (hereinafter referred to as the "model" for ease of description) includes two parts: an encoder and a decoder. The decoder further includes an ICR selection decoder and an EV selection decoder. To select a target charging device from multiple candidate charging devices, the ICR selection decoder needs to determine the device state characteristics and sequence state characteristics of each charging device based on charging demand information and device resource information.
[0093] Furthermore, the charging demand information and equipment resource information are vectorized. For example, the total equipment resource information can be represented as... Where r represents a corresponding device resource information, m represents the total number of charging devices, t represents the t-th selection process, and each selection process includes the selection of the target charging device and the corresponding target vehicle; E represents the current remaining battery value; T represents the current cumulative working time; and G represents the current charging sequence. In this way, the model can perform efficient data processing operations based on the vectorized information.
[0094] Furthermore, the ICR selection decoder obtains the first context state feature based on the device resource information corresponding to each charging device.
[0095]
[0096] in, Indicates that r is at step t-1 m The corresponding current remaining power value, Indicates that r is at step t-1 m The corresponding current cumulative working time is used to concatenate the E and T values of m charging devices to obtain the first context state feature.
[0097] Furthermore, based on the preset first decoder parameters, the first context state features are linearly transformed and activated to obtain the device state features, specifically as shown in equation (1) below:
[0098]
[0099] Wherein, W2 and b1 are both preset first decoder parameters, FF is the feedforward layer (FF layer) in the model, and the FF layer is a 512-dimensional neural network processing layer with ReLU activation function. The FF layer typically includes fully connected layers, convolutional layers, pooling layers, etc. Furthermore, the FF layer mentioned in the embodiments of this application has the same meaning as that here, and will not be repeated hereafter.
[0100] Among them, the rectified linear near unit (ReLU) is a commonly used activation function for neural networks. It improves the expressive power and training efficiency of neural network models by introducing nonlinear mapping and alleviating the gradient vanishing problem.
[0101] Furthermore, The device state features are obtained by performing a linear transformation on the first context state features based on the first decoder parameters, followed by activation processing based on the result of the linear transformation. In this embodiment, the linear transformation can be a linear projection, and the activation processing of the data is performed by the FF layer to obtain the device state features.
[0102] Step 202: Generate a high-dimensional mapping sequence based on the current charging sequences corresponding to each candidate charging device, and perform linear transformation and activation processing on all the high-dimensional mapping sequences according to the preset second decoder parameters to obtain sequence state features.
[0103] The following provides a detailed description of Step 202.
[0104] In some embodiments, to implement the selection of the target charging device, in addition to the device state features, it is also necessary to obtain the current charging sequences corresponding to each charging device, so as to determine the sequence state features based on the current charging sequences later, and determine the target charging device from multiple candidate charging devices based on the device state features and the sequence state features.
[0105] Among them, the current charging sequence is used to indicate the target vehicles selected by each charging device during the previous t selection processes, and the charging order between these target vehicles. The current charging sequence can provide directional guidance for the selection of the target vehicle at the (t + 1)-th time later. For example, assume that all candidate vehicles need to be selected in T selection processes, the target charging sequence corresponding to charging vehicle a is [A, B, C], and at the t-th selection process, the current charging sequence corresponding to charging vehicle a is [A, B], where t < T. Then it can be understood that the current charging sequence is a subset of the target charging sequence.
[0106] Among them, the high-dimensional mapping sequence is a mapping and representation of the current charging sequence in a higher-dimensional space. To obtain the high-dimensional mapping sequence, first determine the vehicles included in the current charging sequence. Each vehicle has its corresponding high-dimensional mapping feature. Concatenate the corresponding high-dimensional mapping features of each vehicle in the order of selection of the vehicles in the current charging sequence to obtain the high-dimensional mapping sequence.
[0107] Further, the ICR selection decoder continues to obtain the second context state feature based on the device resource information corresponding to each charging device
[0108]
[0109] Among them, represents the high-dimensional mapping feature corresponding to the (t - 1)-th charging device in the current charging sequence. Connect all the high-dimensional mapping features in the current charging sequence to obtain the second context state feature.
[0110] Further, as shown in the following formula (2), after determining the second state context feature, perform linear transformation and activation processing according to the preset second decoder parameters to obtain sequence state features
[0111]
[0112] Where W2 and b2 are both preset second decoder parameters; The second context state features of the FF layer based on the input The results obtained after max pooling and aggregation are specifically shown in equations (3) and (4) below, after performing max pooling and aggregation operations:
[0113]
[0114]
[0115] Where max represents the max pooling operation. This is the aggregation result obtained by aggregating the results of max pooling of the context features of each second state.
[0116] In some embodiments, a high-dimensional mapping sequence is generated based on the current charging sequence corresponding to each of the candidate charging devices, including the following steps 301 to 304:
[0117] Step 301: Based on the preset encoder parameters, the charging demand information corresponding to each candidate vehicle and the equipment resource information corresponding to all candidate charging devices are spliced together to obtain the initial features.
[0118] Step 301 will be described in detail below.
[0119] In some embodiments, in order to better determine the relationship between each candidate vehicle and each charging device, it is necessary to obtain high-dimensional mapping features that characterize the features of the candidate charging devices and the candidate vehicles, so that a high-dimensional mapping sequence can be formed later based on the current charging sequence and the high-dimensional mapping features corresponding to each vehicle.
[0120] In some embodiments, the high-dimensional mapping features characterize the association between each candidate vehicle and all candidate charging devices. Therefore, in order to determine the high-dimensional mapping sequence, it is first necessary to determine the high-dimensional mapping features corresponding to each vehicle.
[0121] Furthermore, the purpose of splicing charging demand information and equipment resource information is to integrate vehicle characteristics and equipment characteristics. By fusing information from different sources, more comprehensive and richer characteristic data can be obtained. This allows for decision-making not only from the perspective of the vehicle or the equipment, but also by fusing information from multiple data sources to compensate for the shortcomings of a single data source and comprehensively consider the data correlation between the vehicle and the equipment. This enables a more reasonable and reliable decision-making process for selecting the target charging equipment.
[0122] For example, the encoder in the model embeds the acquired charging demand information and device resource information into a 128-dimensional vector to obtain the initial feature h. i As shown in equation (5):
[0123]
[0124] Here, Concat is the concatenation operation, and Σ is the summation operation.
[0125] in, E j T represents the current remaining power value of any candidate charging device. j This indicates the current cumulative working time corresponding to the selected charging device; t i d represents the dwell time of any candidate vehicle. i This indicates the start and end values for charging for the candidate vehicle. W evt W evd and W cr These are the preset encoder parameters.
[0126] Step 302: Using a preset first attention mechanism as a constraint, perform feature enhancement on the initial features to obtain updated initial features.
[0127] Step 302 will be described in detail below.
[0128] In some embodiments, the encoder includes a first attention module constrained by a first attention mechanism. After obtaining initial features, the initial features are augmented based on the first attention module to update the initial features. Based on the first attention mechanism, the model can selectively focus on more important parts of the initial data to improve the accuracy of subsequent selection of target charging devices and target vehicles.
[0129] Furthermore, as shown in equation (6), the initial feature h i The input is fed into the first attention module to perform feature enhancement on the initial features, resulting in updated initial features.
[0130]
[0131] Here, softmax is an activation function in the model; W Q W K and W V These are independent trainable parameters between different encoder sublayers in the model; typically, d k =16.
[0132] Step 303: Perform feedforward propagation on the updated initial features to obtain the high-dimensional mapping features corresponding to the candidate vehicles.
[0133] Furthermore, the model in this embodiment includes three autoencoder sub-layers, and each sub-layer contains an attention module and a feedforward propagation module that are interconnected. The attention module processes the data as shown in equation (7):
[0134]
[0135] Where MHA represents the first attention layer processing operation; the layer index η indicates that each encoder sublayer does not share parameters; BN represents batch normalization, and the encoder layer uses a skip connection processing method; h i i It is the first input h i It is understandable that h obtained in step 301 i Afterwards, the updated initial features are obtained through processing according to formula (6). Then, as shown in formula (8), all the updated initial features are concatenated, and then the first attention module performs feature processing to obtain the attention features.
[0136]
[0137] Concat is a concatenation operation.
[0138] Furthermore, in formula (7) It comes from the output of the previous feedforward propagation module connected to the current first attention module, and the attention features output by the first attention module. This data will be input into the feedforward propagation module of the next connection. Specifically, the feedforward propagation module processes the data as shown in equation (9):
[0139]
[0140] Furthermore, the feedforward processing operation in equation (9) is shown in equation (10) below:
[0141]
[0142] Among them, W F1 W F2 b F1 and b F2 represents the trainable parameters; the attention features, after being linearly processed based on the trainable parameters, are activated by the ReLU function on a linear feedforward propagation layer with a hidden sublayer dimension of 512.
[0143] Furthermore, through the cooperation between multiple attention modules and feedforward propagation modules, important information and contextual relationships of charging demand information and device resource information are captured, and the high-dimensional mapping features corresponding to each candidate vehicle are determined by the feedforward propagation module output of the last encoder sublayer.
[0144] It should be noted that, since the decoder of the model also includes an attention module, in order to better distinguish them, the attention module in the encoder of this application embodiment is the first attention module, while the attention module in the decoder is the second attention module.
[0145] Step 304: For each candidate charging device, the corresponding high-dimensional mapping features are sequentially concatenated to obtain a high-dimensional mapping sequence, based on the selection order of each candidate vehicle indicated in the current charging sequence.
[0146] In some embodiments, after determining the high-dimensional mapping features corresponding to each candidate vehicle, the high-dimensional mapping features of the corresponding candidate vehicles can be spliced together according to the current charging sequence to obtain a high-dimensional mapping sequence.
[0147] Furthermore, each high-dimensional mapping feature can be processed again according to preset weight parameters, and the processed results can be spliced together to obtain a high-dimensional mapping sequence.
[0148] Step 203: Perform linear transformation and activation processing on the result of splicing the device state features and sequence state features to obtain the device selection probability value corresponding to the candidate charging device, and select the target charging device from multiple candidate charging devices according to the device selection probability value.
[0149] Step 203 will be described in detail below.
[0150] In some embodiments, the current service status of the candidate charging device is comprehensively evaluated based on device state features characterizing the current power level and service time of the candidate charging device, and sequence state features characterizing the features of vehicles that are ready to be served by the candidate charging device, formed according to the current charging sequence, so as to better select the target charging device from multiple candidate charging devices.
[0151] Furthermore, as shown in equations (11) and (12) below, the splicing equipment state characteristics and sequence state features The spliced result is then subjected to linear transformation and activation processing:
[0152]
[0153]
[0154] Where W3 and b3 are trainable parameters; softmax is the activation function; p t Select a probability value for the device.
[0155] Furthermore, after obtaining the device selection probability value corresponding to each candidate charging device, the candidate charging device with the largest device selection probability value is selected as the target charging device; or, the candidate charging device with the largest device selection probability value is randomly selected as the target charging device; the specific selection can be adjusted according to different model selection strategies, and the embodiments of this application do not impose specific limitations.
[0156] Understandably, compared to related technologies that use only one encoder to simultaneously determine the target charging device and the target vehicle, the embodiments of this application subdivide the decoder into an ICR selection decoder and an EV selection decoder, making the model's task modular and making it easier for the model to understand and debug the decoder parameters during training. Furthermore, the ICR selection decoder can focus more on processing information related to the candidate charging device, while the EV selection decoder can focus more on processing information related to the candidate vehicle, thereby improving the accuracy, efficiency, and flexibility of the model's decision-making.
[0157] Step 103: If the target charging action is to charge externally, select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current charging demand information and equipment resource information, and update the corresponding equipment resource information and charging demand information.
[0158] Step 103 will be described in detail below.
[0159] Furthermore, if the target charging action is internal charging, the target charging device will not charge any of the candidate vehicles, but will instead replenish its own power so that it can better perform vehicle charging operations later.
[0160] In some embodiments, selecting a corresponding target vehicle for a target charging device based on current charging demand information and device resource information includes the following steps 401 to 403:
[0161] Step 401: Fuse all high-dimensional mapping features to obtain global feature information.
[0162] Step 401 will be described in detail below.
[0163] In some embodiments, if the target charging action is external charging, it is necessary to select the corresponding target vehicle for the already determined target charging device. To help the model better understand the data and improve the accuracy of target vehicle selection, in addition to the current device resource information of the target charging device and the charging demand information corresponding to each candidate vehicle, it is also necessary to obtain global feature information based on high-dimensional mapping features.
[0164] Furthermore, as shown in equation (13), all high-dimensional mapping features are fused, and the fused result is averaged to obtain the global feature information H. N :
[0165]
[0166] Step 402: Determine the device resource information corresponding to the target charging device, and use the preset second attention mechanism as a constraint to perform feature enhancement on the result of splicing the device resource information, global feature information and high-dimensional mapping sequence to obtain vehicle state features.
[0167] Step 402 is described in detail below.
[0168] In some embodiments, the encoder is further provided with a second attention module constrained by a second attention mechanism, as shown in equations (14) and (15) below, to obtain vehicle state features based on the second attention module.
[0169]
[0170]
[0171] in, and These are the current cumulative working time and current remaining power value in the equipment resource information corresponding to the target charging device; This is the last high-dimensional mapping feature in the current high-dimensional mapping sequence. It can also be the result of weighting multiple high-dimensional mapping features; H N This refers to global feature information; and These are trainable parameters; it is understandable that the vehicle state features are obtained by using the second attention mechanism (MHA) to enhance the features of the concatenated results of device resource information, global feature information, and high-dimensional mapping sequence.
[0172] Step 403: Perform feature extraction and activation processing on the vehicle state features to obtain the vehicle selection probability value corresponding to the candidate vehicle, and determine the target vehicle corresponding to the target charging device from multiple candidate vehicles based on the vehicle selection probability value.
[0173] Step 403 will be described in detail below.
[0174] Furthermore, as shown in equations (16) to (19) below, the vehicle selection probability value of the corresponding candidate vehicle is obtained based on the vehicle state characteristics:
[0175]
[0176]
[0177]
[0178]
[0179] Where C is the control of the first intermediate value u t The entropy is usually set to 10; the second intermediate value q t and the third intermediate value k t It can be obtained through training; combined with masking rules. t And obtain the first intermediate value u based on vehicle state features t The softmax function outputs the vehicle selection probability value for each candidate vehicle; where Z is usually a large negative number.
[0180] Furthermore, after obtaining the vehicle selection probability value corresponding to each candidate vehicle, the candidate vehicle with the largest vehicle selection probability value is selected as the target vehicle; or, the candidate vehicle with the largest vehicle selection probability value is randomly selected as the target vehicle; the specific selection can be adjusted according to different model selection strategies, and the embodiments of this application do not impose specific limitations.
[0181] In some embodiments, updating the corresponding device resource information includes the following steps 501 to 504:
[0182] Step 501: Obtain the device resource information corresponding to the currently selected target charging device, and the charging demand information corresponding to the target vehicle.
[0183] Step 501 is described in detail below.
[0184] In some embodiments, in order for the model to better understand the changes in the information status of each charging device and vehicle, so as to ensure the smooth execution of charging tasks and the effective utilization of resources, it is necessary to update the device resource information and charging demand information.
[0185] Furthermore, for each candidate vehicle, its own charging demand information remains unchanged; however, for the model in this embodiment, once a vehicle is selected, its charging demand information is no longer considered. Instead, the analysis continues based on the charging demand information of other candidate vehicles to achieve charging for all candidate vehicles within the target area. Therefore, updating the charging demand information in this embodiment involves deleting the charging demand information corresponding to the selected target vehicle from the total charging demand information. For example, if the target area includes candidate vehicles A, B, and C, and the total charging demand information formed based on the charging demand information of each vehicle is [A, B, C], then if vehicle C is selected as the target vehicle in a certain round, the updated charging demand information relative to the model is [A, B].
[0186] Furthermore, in order to update the equipment resource information, it is first necessary to obtain the equipment resource information corresponding to the target charging equipment selected in the current round, as well as the charging demand information corresponding to the target vehicle.
[0187] Step 502: Calculate the power difference based on the current remaining power value and the charging start and end values, and calculate the total time value based on the current cumulative working time and dwell time.
[0188] Step 503: Update the current remaining power value based on the power difference, update the current cumulative working time based on the total time value, and update the current charging sequence based on the selected order of the target vehicles.
[0189] The following describes steps 502 to 503 with examples.
[0190] For example, in a certain round, the selected target charging equipment is target charging equipment A and target vehicle A. The equipment resource information for target charging equipment A is: current remaining power value 70%, cumulative working time 70 minutes, and current charging sequence [B,C]. The charging demand information for target vehicle A is: charging start / end value 50%, and dwell time 150 minutes. Therefore, the updated current remaining power value can be determined as (70%-50%) = 20%. Based on the charging power, the charging time required for target charging equipment A to charge target vehicle A is calculated. Assuming the required charging time is 50 minutes, the updated current cumulative working time can be determined as (70+50) = 120 minutes. The updated current charging sequence is determined as [B,C,A].
[0191] Step 104: Until all candidate vehicles have been selected, generate the corresponding target charging sequence based on the selection order of multiple target charging actions under each target charging device.
[0192] Step 104 will be described in detail below.
[0193] In some embodiments, when all candidate vehicles have been selected, a target charging sequence is formed based on the target charging action corresponding to each target charging device and the order of each target charging action. Specifically, when the target charging action indicates that the charging device is charging externally, the target vehicle corresponding to the target charging and the selection order of each target vehicle are determined to form the target charging sequence.
[0194] For example, a charging device may select multiple target charging actions in a chosen order as [charging inward, charging outward, charging outward]. Further, based on the target vehicle selected when the charging device performs the charging outward action, the target charging sequence of the charging device is determined to be [charging inward, vehicle A, vehicle B]. Alternatively, when the charging device only performs the charging outward action, it outputs a target charging sequence containing the charging order of each target vehicle.
[0195] Step 105: Based on each target charging sequence, schedule the corresponding target charging equipment to perform the target charging action in the selected order to complete the charging of the target vehicle.
[0196] Step 105 is described in detail below.
[0197] In some embodiments, after determining the target charging sequence corresponding to each charging device, the target charging devices can be scheduled according to the target charging action indicated by the target charging sequence to complete the charging of the target vehicle. It is understood that the deep reinforcement learning model can intelligently output the target charging sequence based on charging demand information and device resource information. For the user, only the relevant charging start and end values and dwell time need to be provided to obtain a vehicle charged to the specified level within the expected time. For the charging device, it can maximize the output of electricity, thereby improving the overall charging efficiency.
[0198] Furthermore, for all charging devices within the target area, barring mechanical failures, all devices will act to execute their respective target charging actions within their target charging sequence. Each individual charging device will complete its current action before moving on to the next. In this way, each charging device can charge vehicles within the target area in an orderly manner, avoiding conflicts and delays caused by resource competition between multiple charging devices, ensuring a smooth charging process. Pre-determining the target actions for each charging device and the corresponding charging sequence for each target vehicle helps improve charging efficiency, reduce system operating costs, enhance user experience, and also contributes to system stability and optimized resource utilization.
[0199] In some embodiments, the deep reinforcement learning model is trained through the following steps 601 to 605:
[0200] Step 601: Obtain the sample charging demand information of multiple sample candidate vehicles within the sample area, as well as the sample equipment resource information of multiple sample candidate charging devices.
[0201] Step 602: Input the sample charging demand information and sample equipment resource information into the deep reinforcement learning model. When the sample candidate vehicles are not all selected, select the sample target charging equipment from multiple sample candidate charging equipment and determine the sample target charging action corresponding to the sample target charging equipment. The sample target charging action includes one of inward charging and outward charging.
[0202] Step 603: If the target charging action of the sample is to charge externally, select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current sample charging demand information and sample equipment resource information, and update the corresponding sample equipment resource information and sample charging demand information.
[0203] Step 604: Until all candidate vehicles have been selected, generate the corresponding target charging sequence based on the selection order of multiple target charging actions under each target charging device.
[0204] Step 605: Determine the sample reward value based on the sample target charging sequence, and adjust the parameters of the deep reinforcement learning model based on the sample reward value to obtain the trained deep reinforcement learning model.
[0205] Steps 601 to 605 are described in detail below.
[0206] In some embodiments, to enable the deep reinforcement learning model to output more accurate target charging sequences in practical applications, the deep reinforcement learning model needs to be pre-trained. Similar to the target region, the size of the sample region can also be set according to the actual training needs; the sample candidate vehicles are also electrically powered vehicles, and the sample candidate charging devices can also be intelligent charging robots. That is, the method for obtaining sample charging demand information and sample device resource information during the training process of the deep reinforcement learning model, and the method for processing the obtained sample charging demand information and sample device resource information to obtain the sample target charging sequence, is similar to steps 101 to 105, and will not be repeated here. The difference is that after determining the sample target charging sequence, a sample reward value is determined based on the sample target charging sequence. The parameters of the deep reinforcement learning model are adjusted according to the sample reward value, so that the trained deep reinforcement learning model can better adapt to specific tasks and environments, improve performance and generalization ability, and thus better solve the charging scheduling problem in practical applications.
[0207] In some embodiments, determining the sample reward value based on the sample target charging sequence includes the following steps 701 to 704:
[0208] Step 701: If the sample target charging device is determined based on the maximum sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on the maximum sample vehicle selection probability value, determine the first charging quantity of each sample target vehicle matched with each sample target charging device and the first charging amount released by each sample target vehicle according to the sample target charging sequence.
[0209] Step 702: Obtain the first sample reward value based on the total value of all first charging quantities and / or the total value of first charging power.
[0210] Step 703: If the sample target charging device is determined based on the random sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on the random sample vehicle selection probability value, determine the second charging quantity of each sample target charging device and the second charging amount released by each sample target vehicle according to the sample target charging sequence.
[0211] Step 704: Obtain the second sample reward value based on the total value of all second charging quantities and / or the total value of second charging power.
[0212] Steps 701 to 704 are described in detail below.
[0213] In some embodiments, similar to steps 201 to 203 and steps 401 to 403, the sample target charging device in the sample target charging sequence is determined based on the sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on the sample vehicle selection probability value.
[0214] Furthermore, the model includes two different policy networks, each determining the sample device selection probability and sample vehicle selection probability in different ways. One policy network uses a greedy algorithm to select the sample device and sample vehicle selection probabilities with the highest values; while the other policy network uses a random method, randomly selecting one from multiple sample device and sample vehicle selection probability values to determine the sample target charging device and sample target vehicle. The first sample reward value and the second sample reward value corresponding to each policy network are then obtained.
[0215] Furthermore, after obtaining multiple sample device selection probability values and vehicle selection probability values, if the charging device corresponding to the device selection probability value with the largest value is selected as the target charging device, and the charging device corresponding to the vehicle selection probability value with the largest value is selected as the target vehicle, then the total value of the first charging quantity and the total value of the first charging quantity under this strategy are determined.
[0216] Furthermore, the first sample reward value can be determined based on the total value of the first number of charges, for example, the larger the total value of the first number of charges, the larger the first sample reward value; or the first sample reward value can be determined based on the total value of the first charge amount, for example, the larger the total value of the first charge amount, the larger the first sample reward value; or the first sample reward value can be determined by summing the total value of the first number of charges and the total value of the first charge amount according to a preset weight value.
[0217] Furthermore, after obtaining multiple sample device selection probability values and vehicle selection probability values, if a charging device corresponding to a device selection probability value is randomly selected as the target charging device, and a charging device corresponding to a vehicle selection probability value is randomly selected as the target vehicle, then the total value of the second charging quantity and the total value of the second charging quantity under this strategy are determined.
[0218] Furthermore, the second sample reward value can be determined based on the total value of the second charging quantity, for example, the larger the total value of the second charging quantity, the larger the second sample reward value; or the second sample reward value can be determined based on the total value of the second charging power, for example, the larger the total value of the second charging power, the larger the second sample reward value; or the second sample reward value can be determined by summing the total value of the second charging quantity and the total value of the second charging power according to a preset weight value.
[0219] Furthermore, the parameters of different policy networks are adjusted according to the reward values of the first and second samples. The parameters of the policy networks include the weight values of the connections between neurons, the bias terms of each neuron, the parameters of each hidden layer, and the parameters of each activation function. The parameters that need to be adjusted can be determined according to the actual situation, and this application embodiment does not impose specific limitations.
[0220] Furthermore, to accelerate parameter adjustment, parameters from a policy network with higher sample reward values can be used to replace those parameters in another policy network. Additionally, after each round of training based on instance data, a new round of evaluation based on new instance data is performed to prevent overfitting. Thus, the iterative updates of the two policy networks allow the model to be trained towards achieving a better solution.
[0221] like Figure 3 As shown, Figure 3This is a schematic diagram of an optional functional module of the charging scheduling device based on deep reinforcement learning provided in the embodiments of this application, wherein the charging scheduling device may include:
[0222] The acquisition module 801 is used to acquire charging demand information of multiple candidate vehicles and equipment resource information of multiple candidate charging devices within the target area.
[0223] The target charging action determination module 802 is used to input charging demand information and equipment resource information into a pre-trained deep reinforcement learning model. When the candidate vehicles are not all selected, the target charging device is selected from multiple candidate charging devices, and the target charging action corresponding to the target charging device is determined. The target charging action includes one of inward charging and outward charging.
[0224] The target vehicle determination module 803 is used to select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current charging demand information and equipment resource information if the target charging action is external charging, and to update the corresponding equipment resource information and charging demand information.
[0225] The target charging sequence determination module 804 is used to generate a corresponding target charging sequence based on the selection order of multiple target charging actions under each target charging device until all candidate vehicles have been selected.
[0226] The target execution module 805 is used to schedule the corresponding target charging equipment to perform target charging actions in a selected order based on each target charging sequence, so as to complete the charging of the target vehicle.
[0227] This application proposes a charging scheduling method, apparatus, device, and medium based on deep reinforcement learning. The method first acquires charging demand information for multiple candidate vehicles and equipment resource information for multiple candidate charging devices within a target area. Next, the charging demand information and equipment resource information are input into a pre-trained deep reinforcement learning model. When not all candidate vehicles have been selected, a target charging device is selected from the multiple candidate charging devices, and the corresponding target charging action is determined. The target charging action includes either inward charging or outward charging. By simultaneously considering both inward and outward charging actions, the target charging device can serve as many target vehicles as possible. Furthermore, if the target charging action is outward charging, a corresponding target vehicle is selected from the remaining candidate vehicles based on the current charging demand information and equipment resource information, and the corresponding equipment resource information and charging demand information are updated. This process continues until all candidate vehicles have been selected. Based on the selection order of multiple target charging actions under each target charging device, a corresponding target charging sequence is generated. Based on each target charging sequence, the corresponding target charging device is scheduled to execute the target charging action in the selected order to complete the charging of the target vehicles. With the aim of maximizing power output or serving as many vehicles as possible, a target charging sequence containing two types of charging actions is generated. This allows for better scheduling of each charging device based on the target charging sequence, thereby improving the charging scheduling efficiency of the charging devices.
[0228] The specific implementation of the deep reinforcement learning-based charging scheduling device is basically the same as the specific implementation of the deep reinforcement learning-based charging scheduling method described above, and will not be repeated here.
[0229] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the aforementioned charging scheduling method based on deep reinforcement learning. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.
[0230] like Figure 4 As shown, Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. The electronic device includes:
[0231] The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.
[0232] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to implement the deep reinforcement learning-based charging scheduling method of the embodiments of this application.
[0233] The input / output interface 903 is used to implement information input and output;
[0234] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0235] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);
[0236] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.
[0237] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described deep reinforcement learning-based charging scheduling method.
[0238] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0239] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.
[0240] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.
[0241] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0242] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.
[0243] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0244] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0245] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0246] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0247] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0248] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0249] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.
Claims
1. A charging scheduling method based on deep reinforcement learning, characterized in that, include: Obtain charging demand information for multiple candidate vehicles within the target area, as well as equipment resource information for multiple candidate charging devices; The charging demand information and the equipment resource information are input into a pre-trained deep reinforcement learning model. When the candidate vehicles are not all selected, a target charging device is selected from the multiple candidate charging devices, and the target charging action corresponding to the target charging device is determined. The target charging action includes one of inward charging and outward charging. Inward charging means inputting electrical energy into the target charging device, and outward charging means releasing electrical energy outward from the target charging device. If the target charging action is external charging, select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current charging demand information and the device resource information, and update the corresponding device resource information and the charging demand information. Until all the candidate vehicles have been selected, a corresponding target charging sequence is generated based on the selection order of the multiple target charging actions under each target charging device; Based on each of the target charging sequences, the corresponding target charging devices are scheduled to perform the target charging actions in the selected order to complete the charging of the target vehicle; The device resource information includes the current remaining power, the current cumulative working time, and the current charging sequence, wherein the current charging sequence is a subset of the target charging sequence; Selecting a target charging device from a plurality of candidate charging devices includes: Based on the preset first decoder parameters, all the current remaining power values and the current cumulative working time are linearly transformed and activated to obtain the device status characteristics; A high-dimensional mapping sequence is generated based on the current charging sequence corresponding to each of the candidate charging devices. A linear transformation and activation process are performed on all the high-dimensional mapping sequences according to the preset second decoder parameters to obtain the sequence state features. The result of concatenating the device state features and the sequence state features is subjected to linear transformation and activation processing to obtain the device selection probability value corresponding to the candidate charging device. The target charging device is selected from multiple candidate charging devices based on the device selection probability value.
2. The method according to claim 1, characterized in that, The step of generating a high-dimensional mapping sequence based on the current charging sequence corresponding to each of the candidate charging devices includes: Based on the preset encoder parameters, the charging demand information corresponding to each candidate vehicle and the equipment resource information corresponding to all candidate charging devices are spliced together to obtain the initial features; Using a preset first attention mechanism as a constraint, feature enhancement is performed on the initial features to obtain the updated initial features; The updated initial features are subjected to feedforward propagation to obtain the high-dimensional mapping features corresponding to the candidate vehicles; For each of the candidate charging devices, the corresponding high-dimensional mapping features are sequentially concatenated to obtain a high-dimensional mapping sequence, based on the selection order of each candidate vehicle indicated in the current charging sequence.
3. The method according to claim 2, characterized in that, The step of selecting a corresponding target vehicle for the target charging device based on the current charging demand information and the device resource information includes: By fusing all the aforementioned high-dimensional mapping features, global feature information is obtained; The device resource information corresponding to the target charging device is determined, and the device resource information, the global feature information and the result of splicing the high-dimensional mapping sequence are enhanced with a preset second attention mechanism as a constraint to obtain vehicle state features. The vehicle state features are subjected to feature extraction and activation processing to obtain the vehicle selection probability value corresponding to the candidate vehicle. Based on the vehicle selection probability value, the target vehicle corresponding to the target charging device is determined from multiple candidate vehicles.
4. The method according to claim 1, characterized in that, The charging demand information includes charging start and end values and dwell time; The updated device resource information includes: Obtain the device resource information corresponding to the currently selected target charging device, and the charging demand information corresponding to the target vehicle; The power difference is calculated based on the current remaining power value and the charging start and end values, and the total time value is calculated based on the current cumulative working time and the dwell time. The current remaining battery value is updated based on the battery difference, the current cumulative working time is updated based on the total time value, and the current charging sequence is updated based on the selected order of the target vehicles.
5. The method according to claim 1, characterized in that, The deep reinforcement learning model is trained through the following steps: Obtain sample charging demand information for multiple candidate vehicles within the sample area, as well as sample equipment resource information for multiple candidate charging devices; The sample charging demand information and the sample device resource information are input into the deep reinforcement learning model. When the sample candidate vehicles are not all selected, a sample target charging device is selected from multiple sample candidate charging devices, and the sample target charging action corresponding to the sample target charging device is determined. The sample target charging action includes one of inward charging and outward charging. If the target charging action is external charging, select the corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current target charging demand information and the target device resource information, and update the corresponding target device resource information and target charging demand information. Until all the selected vehicles have been selected, a corresponding target charging sequence is generated based on the selection order of multiple target charging actions under each target charging device. The sample reward value is determined based on the sample target charging sequence, and the parameters of the deep reinforcement learning model are adjusted based on the sample reward value to obtain the trained deep reinforcement learning model.
6. The method according to claim 5, characterized in that, The sample target charging device in the sample target charging sequence is determined based on the sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on the sample vehicle selection probability value. The sample reward value includes a first sample reward value and a second sample reward value. The step of determining the sample reward value based on the sample target charging sequence includes: If the target charging device is determined based on the maximum selection probability value of the target device, and each target vehicle corresponding to the target charging device is determined based on the maximum selection probability value of the target vehicle, the first charging quantity of each target vehicle matched with each target charging device and the first charging quantity released by each target vehicle are determined according to the target charging sequence. The first sample reward value is obtained based on the total value of all the first charging quantities and / or the total value of the first charging power. If the sample target charging device is determined based on a random sample device selection probability value, and each sample target vehicle corresponding to the sample target charging device is determined based on a random sample vehicle selection probability value, the second charging quantity of each sample target vehicle matched with each sample target charging device and the second charging amount released by each sample target vehicle are determined based on the sample target charging sequence. The second sample reward value is obtained based on the total value of all the second charging quantities and / or the total value of the second charging power.
7. A charging scheduling device based on deep reinforcement learning, characterized in that, The device includes: The acquisition module is used to acquire charging demand information of multiple candidate vehicles and equipment resource information of multiple candidate charging devices within the target area. The target charging action determination module is used to input the charging demand information and the equipment resource information into a pre-trained deep reinforcement learning model. When the candidate vehicles are not all selected, the module selects a target charging device from the multiple candidate charging devices and determines the target charging action corresponding to the target charging device. The target charging action includes one of inward charging and outward charging. Inward charging means inputting electrical energy into the target charging device, and outward charging means releasing electrical energy outward from the target charging device. The target vehicle determination module is used to select a corresponding target vehicle for the target charging device from the remaining multiple candidate vehicles based on the current charging demand information and the device resource information if the target charging action is external charging, and update the corresponding device resource information and the charging demand information. The target charging sequence determination module is used to generate a corresponding target charging sequence based on the selection order of multiple target charging actions under each target charging device until all the candidate vehicles have been selected. The target execution module is used to schedule the corresponding target charging device to perform the target charging action in the selected order based on each target charging sequence, so as to complete the charging of the target vehicle; The device resource information includes the current remaining power, the current cumulative working time, and the current charging sequence, wherein the current charging sequence is a subset of the target charging sequence; Selecting a target charging device from a plurality of candidate charging devices includes: Based on the preset first decoder parameters, all the current remaining power values and the current cumulative working time are linearly transformed and activated to obtain the device status characteristics; A high-dimensional mapping sequence is generated based on the current charging sequence corresponding to each of the candidate charging devices. A linear transformation and activation process are performed on all the high-dimensional mapping sequences according to the preset second decoder parameters to obtain the sequence state features. The result of concatenating the device state features and the sequence state features is subjected to linear transformation and activation processing to obtain the device selection probability value corresponding to the candidate charging device. The target charging device is selected from multiple candidate charging devices based on the device selection probability value.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the charging scheduling method based on deep reinforcement learning as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the charging scheduling method based on deep reinforcement learning as described in any one of claims 1 to 6.
Citation Information
Patent Citations
WRSN target k coverage charging scheduling method based on deep reinforcement learning
CN117979305A
Methods, computer programs and systems for assigning vehicles to vehicular tasks and for providing a machine-learning model
EP3806007A1