Battery replacement scheduling method, device, equipment, computer storage medium and program product
Patent Information
- Application Number
- CN202610233430.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-27
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-02-27
AI Technical Summary
[0004]然而,换电模式在实际应用中仍存在显著瓶颈:例如,在港口、矿山及大型封闭物流园区等场景中,大量重卡同时执行动态指派的任务,作业结束时间与电量消耗均具有较强不确定性,容易导致车辆换电需求在时间与空间上集中产生,进而引发换电站排队拥堵
[0034]第四方面,本申请提供了一种计算机可读存储介质,计算机可读存储介质上存储有计算机程序指令,计算机程序指令被处理器执行时实现如第一方面的换电调度方法。
Smart Images

Figure CN121766717B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of battery swapping station technology, and in particular to a battery swapping scheduling method, apparatus, equipment, computer-readable storage medium, and computer program product. Background Technology
[0002] Due to current limitations in battery technology, electric heavy-duty trucks still generally face practical problems such as limited driving range and long charging time, which poses new challenges to scenarios such as ports and mines that require high continuity and high load operation.
[0003] To shorten energy replenishment time, battery swapping is becoming increasingly popular. This model allows vehicles to resume operation within minutes by quickly replacing fully charged battery packs at battery swapping stations, significantly improving operational continuity and vehicle uptime.
[0004] However, battery swapping still faces significant bottlenecks in practical applications. For example, in scenarios such as ports, mines, and large enclosed logistics parks, a large number of heavy trucks simultaneously perform dynamically assigned tasks. The completion time and power consumption of these tasks are highly uncertain, easily leading to a concentrated surge in battery swapping demand in both time and space, resulting in queues and congestion at battery swapping stations. This phenomenon not only prolongs the time vehicles are away from operations due to battery swapping but may also cause delays in subsequent tasks, reducing overall operational efficiency. Therefore, how to efficiently and adaptively schedule and make real-time decisions for battery swapping of electric heavy trucks in dynamic and uncertain environments has become a critical issue that urgently needs to be addressed. Summary of the Invention
[0005] This application provides a battery swapping scheduling method, apparatus, equipment, computer-readable storage medium, and computer program product, which can perform efficient and adaptive battery swapping scheduling and real-time decision-making for vehicles in dynamic and uncertain environments, thereby reducing queuing time and improving vehicle uptime and operational continuity.
[0006] In a first aspect, this application provides a battery swapping scheduling method, which includes: responding to a reporting request from a first vehicle in a closed operation scenario to complete a task at a first moment, acquiring vehicle data of all operating vehicles and battery swapping station data of all battery swapping stations in the closed operation scenario; constructing a first multi-dimensional state vector of the closed operation scenario at the first moment based on the vehicle data and the battery swapping station data; inputting the first multi-dimensional state vector into a pre-trained battery swapping action prediction model in an intelligent agent to obtain reward function values corresponding to different candidate battery swapping actions of the first vehicle; making a decision on the battery swapping action of the first vehicle based on multiple reward function values through a decision module in the intelligent agent to obtain the target battery swapping action of the first vehicle; and scheduling the battery swapping of the first vehicle based on the target battery swapping action.
[0007] As described above, in this embodiment, by responding to vehicle operation completion requests in real time, vehicle data and battery swapping station data of all operating vehicles in the closed operation scenario are dynamically collected. A multi-dimensional state vector is constructed and input into the battery swapping action prediction model in the trained agent to obtain multiple reward function values. Based on these reward function values, the decision module in the agent selects target battery swapping actions, ensuring that the decision results take into account both individual vehicle needs and global optimization of battery swapping resources. This achieves real-time, adaptive scheduling of battery swapping demand. This method overcomes the shortcomings of rigid rule-based scheduling and poor real-time performance of operations optimization in traditional methods. It can make intelligent scheduling decisions based on the global state in dynamic and uncertain closed operation scenarios (such as ports and mines), thereby reducing queuing time and improving vehicle attendance and operational continuity.
[0008] In some implementations, the first multidimensional state vector includes the multidimensional state vector of the first vehicle and the global state vector in the closed operation scenario; the multidimensional state vector of the first vehicle includes: the first state of charge (SOC) value of the first vehicle and the time required for the first vehicle to travel to each battery swapping station; the global state vector includes: the next idle time of each battery swapping station, the duration of a single battery swap, the SOC distribution statistics of all operating vehicles, and the predicted demand information within a preset time period.
[0009] By fusing local vehicle information with global information from a closed operation scenario using multi-dimensional state vectors, the system accurately captures the vehicle's core attributes (SOC, driving time) and integrates the battery swapping station's resource status (idle time, battery swapping duration) with the global supply and demand status (SOC distribution, predicted demand). This provides a complete context for decision-making, enabling the agent to not only consider the current vehicle's battery level and location when making decisions based on multi-dimensional state vectors, but also to perceive the battery swapping station's busy / idle status, the overall battery distribution of the fleet, and short-term demand predictions, thereby making more systematic and better scheduling decisions.
[0010] In some embodiments, the first multidimensional state vector is input into a pre-trained intelligent agent's battery swapping action prediction model to obtain reward function values corresponding to different candidate battery swapping actions of the first vehicle. This includes: inputting the first multidimensional state vector into a pre-trained intelligent agent's battery swapping action prediction model to obtain different candidate battery swapping actions of the first vehicle; when a candidate battery swapping action is used as a battery swapping destination, for each candidate battery swapping action: based on a first moment, the time required for the first vehicle to travel to the corresponding battery swapping station, and the battery swapping duration of the battery swapping station, determining the total offline time from the first moment to the completion of battery swapping for the first vehicle at the battery swapping station; and based on the first state of charge (SOC) value, the total offline time, the preset average offline time, the time reward coefficient, and the energy reward coefficient, determining the reward function value of the candidate battery swapping action.
[0011] By quantifying the total time a vehicle takes to go offline from the current moment until the battery swap is completed, and combining it with factors such as the preset average offline time and vehicle battery level, a highly guiding reward signal is constructed. This enables the intelligent agent to learn how to reduce the time a vehicle is away from operation while rationally arranging the timing of battery swaps and site selection, thereby improving overall scheduling efficiency.
[0012] In some embodiments, the power reward coefficient includes a first power reward coefficient, a second power reward coefficient, and a third power reward coefficient. The first power reward coefficient is greater than the second power reward coefficient, and the third power reward coefficient is between the first power reward coefficient and the second power reward coefficient, and varies with the first state of charge (SOC). The value decreases linearly with increasing value. Based on the first State of Charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and energy reward coefficient, the reward function value for candidate battery swapping actions is determined, including: when the first SOC value is less than a preset first SOC, the reward function value for candidate battery swapping actions is determined based on the first SOC value, total offline time, preset average offline time, time reward coefficient, and first energy reward coefficient; when the first SOC value is greater than or equal to a preset second SOC, the reward function value for candidate battery swapping actions is determined based on the first SOC value, total offline time, preset average offline time, time reward coefficient, and second energy reward coefficient; when the first SOC value is greater than or equal to a preset first SOC and less than a preset second SOC, the reward function value for candidate battery swapping actions is determined based on the first SOC value, total offline time, preset average offline time, time reward coefficient, and third energy reward coefficient.
[0013] By setting the first power reward coefficient ( ) greater than the second electricity reward coefficient ( This approach prioritizes battery swapping for low-battery vehicles, ensuring that low-battery vehicles receive higher rewards and preventing task interruptions or safety risks due to depleted battery power. Simultaneously, it reduces the reward weight for battery swapping for high-battery vehicles, minimizing the resource consumption of swapping stations due to ineffective swapping by high-battery vehicles and ensuring that swapping resources are allocated to low-battery vehicles that genuinely need them, thus improving resource utilization. A dynamic third reward coefficient (β) is established, falling between the first and second battery reward coefficients and decreasing linearly with increasing State of Charge (SOC) value. iThis approach balances the necessity of battery swapping with the rationality of resource utilization by allowing vehicles with intermediate battery levels to neither be forced to swap batteries (avoiding safety risks associated with low battery levels) nor prohibited from swapping batteries (reserving flexibility for battery swapping to cope with fluctuations in subsequent tasks). By differentiating reward coefficients based on battery level thresholds, unreasonable behaviors such as "not swapping batteries when low" or "frequent swapping when high," are avoided. This ensures that the scheduling strategy conforms to vehicle energy replenishment patterns and adapts to the high-continuity operational needs of closed work scenarios.
[0014] In some embodiments, the first multidimensional state vector is input into the battery swapping action prediction model in the pre-trained agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The method further includes: when the candidate battery swapping action is not to swap batteries, determining the reward function value for not swapping batteries based on the first state of charge (SOC) value.
[0015] For the "no battery swapping" action, the reward is determined directly based on the vehicle's SOC value, ensuring both assessment efficiency (no complex calculations required) and alignment with actual operational needs (continuing to operate vehicles with high battery levels is more valuable). By directly linking the SOC value to the reward, vehicles with high battery levels are encouraged to continue operating (receiving higher rewards), while vehicles with low battery levels are prevented from forced operation (receiving lower rewards), thus promoting a reasonable behavior pattern of "operating with high battery levels and swapping batteries with low battery levels." At the same time, there is no need to calculate additional parameters such as driving time and battery swapping time, allowing for rapid reward assessment of the "no battery swapping" action. Combined with the quantitative assessment of the "go to battery swapping" action, this ensures decision-making quality while improving real-time response speed.
[0016] In some embodiments, the first multidimensional state vector is input into the battery swapping action prediction model in the pre-trained agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The method further includes scaling and smoothing the reward function values to obtain the processed reward function values.
[0017] By scaling and smoothing (such as using the tanh function), the numerical range of the reward function value is limited, avoiding model training oscillations caused by extreme reward values and accelerating model convergence. At the same time, the smoothed reward value is more continuous and reasonable, avoiding misjudging the value of actions by the agent due to excessive fluctuations in the original reward value, and improving the model's ability to recognize the merits of different actions. Furthermore, the processed reward function value is comparable within a unified numerical range, ensuring that the reward evaluation standard is consistent across different scenarios and vehicle states, and avoiding decision-making biases caused by differences in reward scale.
[0018] In some embodiments, the decision on the battery swapping action of the first vehicle is made based on multiple reward function values to obtain the target battery swapping action of the first vehicle, including: using an action masking mechanism to filter out legal candidate battery swapping actions from multiple candidate battery swapping actions; and filtering out the battery swapping action with the largest reward function value from the legal candidate battery swapping actions to obtain the target battery swapping action of the first vehicle.
[0019] By using an action masking mechanism to block invalid or dangerous candidate battery swapping actions, the agent is prevented from choosing unreasonable behaviors such as "not swapping batteries when the battery is low" or "swapping batteries when the battery is high," thus reducing scheduling risks. At the same time, invalid options other than legal actions can be excluded, reducing the decision-making computation load of the agent and ensuring that the target action is output within seconds / minutes, meeting the real-time scheduling requirements of closed operation scenarios. Furthermore, by selecting the action with the largest reward function value within the range of legal actions, system failures caused by illegal actions are avoided, while ensuring the optimality of the decision result, thus balancing safety and efficiency.
[0020] In some embodiments, an action masking mechanism is used to filter out legitimate candidate battery swapping actions from multiple candidate battery swapping actions, including: when the first state of charge (SOC) value is less than a first SOC threshold, filtering out non-battery swapping actions from multiple candidate battery swapping actions to obtain battery swapping actions that do not require battery swapping; and when the first state of charge (SOC) value is greater than a second SOC threshold, filtering out battery swapping actions from multiple candidate battery swapping actions to obtain non-battery swapping actions.
[0021] By using a first SOC threshold (safety threshold) and a second SOC threshold (battery swapping prohibition threshold), the boundaries of action legality are quantified, avoiding subjective judgments in the action masking mechanism and ensuring the consistency of the screening logic. At the same time, vehicles with low battery levels (below the first SOC threshold) are forced to choose "go for battery swapping," fundamentally eliminating safety issues such as work interruptions and breakdowns caused by vehicles running out of power. Vehicles with high battery levels (above the second SOC threshold) are forced to choose "not to swap batteries," preventing high-battery vehicles from occupying battery swapping station resources, freeing up battery swapping space for low-battery vehicles, and reducing queuing congestion at battery swapping stations.
[0022] In some embodiments, the method further includes: after the first vehicle completes battery swapping, updating the state vector of the closed operation scenario at a second time moment and the second multidimensional state vector; the second time moment is the moment when the first vehicle completes battery swapping; and constructing training samples for training the agent based on the first multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second multidimensional state vector.
[0023] On the one hand, by recording the state transitions of the entire battery swapping scheduling process (first multi-dimensional state vector → target battery swapping action → reward value → second multi-dimensional state vector), high-quality training samples from real-world scenarios are continuously generated, providing data support for agent optimization. On the other hand, samples from real-world scenarios contain uncertainties in dynamic environments (such as fluctuations in power consumption and changes in tasks), making them more practical than simulation samples. This allows the agent to continuously learn and adapt to the dynamic changes in the scenario. At the same time, the scheduling process also serves as a sample collection process, enabling the agent to continuously iterate and upgrade in practical applications, gradually improving decision-making accuracy and adapting to the long-term operational needs of closed operation scenarios.
[0024] In some embodiments, the method further includes: sampling from a training sample set to obtain target training samples; each training sample in the training sample set includes: a first multidimensional state vector sample, a battery swapping action sample, a reward function value sample, and a second multidimensional state vector sample corresponding to a closed operation scenario sample under multiple different first time-time samples, wherein the second time-time sample is the time corresponding to the closed operation scenario sample when the battery swapping action sample is completed; inputting the first multidimensional state vector sample into the agent to be trained to obtain a battery swapping prediction action; determining a first reward function prediction value based on the battery swapping prediction action and the first multidimensional state vector sample; determining a loss function value based on the first reward function prediction value and the reward function value sample; adjusting the model parameters of the agent to be trained based on the loss function value if the iteration stopping condition is not met; resampling from the training sample set to obtain newly sampled target training samples; and returning to input the first multidimensional state vector sample into the agent to be trained until the iteration stopping condition is met.
[0025] By adjusting model parameters based on the loss function value, the agent continuously learns the mapping relationship between "state-action-reward," gradually approaching the optimal scheduling strategy and ensuring the effectiveness and convergence of the agent's training. Simultaneously, the training sample set covers multiple state vectors and action combinations at different times, enabling the agent to learn scheduling patterns in diverse scenarios, avoiding overfitting to a single scenario, and improving the model's generalization ability and robustness. Furthermore, by controlling the training pace through iteration stopping conditions (such as the loss function value falling below a threshold or the number of iterations reaching a target), overtraining or undertraining is avoided, ensuring the agent's stable performance.
[0026] In some embodiments, the training sample set includes a first training sample set and a second training sample set. The training samples in the first training sample set are samples constructed based on expert experience, and the training samples in the second training sample set are samples generated during simulated battery swapping in a simulated closed operation scenario. Resampling from the training sample set to obtain newly sampled target training samples includes: counting the current iteration number; determining the ratio of the current iteration number to a preset iteration number; determining a mixed sampling ratio of the first training sample set and the second training sample set based on the ratio and the minimum retention ratio of the first training sample set; and collecting target training samples corresponding to the mixed sampling ratio from the first training sample set and the second training sample set.
[0027] The first training sample set (expert experience samples) provides high-quality initial learning materials for the agent, enabling it to quickly master basic scheduling rules and avoid inefficiencies caused by random exploration in the early stages of training. By dynamically adjusting the mixed sampling ratio, the initial stage mainly uses expert experience samples (for quick learning) and the later stage mainly uses simulation samples (to overcome the limitations of experience), so that the agent has both basic scheduling capabilities and can learn better strategies for complex scenarios. At the same time, by combining the "determinism" of expert experience with the "uncertainty" of simulation scenarios, the training sample set covers both normal and extreme scenarios (such as high-concurrency battery swapping), improving the agent's coping ability.
[0028] In some embodiments, determining the mixed sampling ratio of the first training sample set and the second training sample set based on the ratio and the minimum retention ratio of the first training sample set includes: calculating the difference between 1 and the ratio; and determining the larger of the minimum retention ratio and the difference as the mixed sampling ratio of the first training sample set and the second training sample set.
[0029] By calculating the difference between "1 - number of iterations / preset number of iterations" and combining it with the minimum retention ratio, the sampling ratio of the first training sample set is gradually reduced to avoid model training oscillations caused by sudden changes in the sampling ratio. The minimum retention ratio (p_min) ensures that a certain proportion of expert experience samples are retained throughout the training process, preventing the agent from completely deviating from reasonable scheduling rules and ensuring the bottom-line safety of the strategy. Furthermore, the dynamically adjusted mixed sampling ratio makes full use of the guiding value of the initial expert samples and fully leverages the exploratory value of the simulation samples in the later stages, so that the sample resources at different stages are optimally utilized.
[0030] In some embodiments, the process of constructing the second training sample set includes: constructing a simulation environment for a closed operation sample scenario and initializing environmental state simulation data in the simulation environment; the environmental state data includes vehicle data and battery swapping station data in the closed operation scenario; when the simulation clock advances to the first simulation moment when the first simulated vehicle completes its operation task, obtaining the first simulation multidimensional state vector of the simulation environment at the first simulation moment from the environmental state simulation data; the first simulated vehicle is any one of multiple simulated vehicles; inputting the first simulation multidimensional state vector into the battery swapping action prediction model in the agent to be trained to obtain simulation reward function values corresponding to different candidate battery swapping actions of the first simulated vehicle; through the decision module in the agent, making a decision on the battery swapping action of the first simulated vehicle based on multiple simulation reward function values to obtain the battery swapping prediction action of the first simulated vehicle; after the first simulated vehicle completes the battery swapping prediction action, updating the state variables in the simulation environment to obtain the second simulation multidimensional state vector; and constructing a second training sample set for training the agent based on the first simulation multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second simulation multidimensional state vector.
[0031] By constructing a high-fidelity simulation environment, the battery swapping scheduling process under a large number of different scenarios (number of vehicles, task distribution, battery swapping demand) can be quickly simulated, and a massive amount of training samples can be efficiently generated, solving the problems of long sample collection cycles and high costs in real scenarios. On the one hand, the simulation environment can flexibly adjust parameters (such as the number of vehicles, battery swapping station capacity, and task frequency) to generate diverse samples such as normal scenarios, high-concurrency scenarios, and extreme power scenarios, enabling the agent to learn comprehensive scheduling rules. On the other hand, the simulation environment can safely test the effects of different battery swapping actions, avoiding operational losses caused by incorrect scheduling in real scenarios. At the same time, through iterative training of simulation samples, the performance of the agent can be optimized in advance, improving the reliability after actual deployment.
[0032] Secondly, this application also provides a battery swapping scheduling device, comprising: an acquisition module, configured to acquire vehicle data of all operating vehicles and battery swapping station data of all battery swapping stations in the closed operating scenario in response to a reporting request from a first vehicle in a closed operating scenario at a first moment to complete its operating task; a construction module, configured to construct a first multi-dimensional state vector of the closed operating scenario at the first moment based on the vehicle data and the battery swapping station data; an input module, configured to input the first multi-dimensional state vector into a pre-trained battery swapping action prediction model in an intelligent agent to obtain reward function values corresponding to different candidate battery swapping actions of the first vehicle; a decision module, configured to make a decision on the battery swapping action of the first vehicle based on multiple reward function values through the decision module in the intelligent agent to obtain the target battery swapping action of the first vehicle; and a scheduling module, configured to perform battery swapping scheduling for the first vehicle based on the target battery swapping action.
[0033] Thirdly, this application provides an electronic device, which includes: a processor and a memory storing computer program instructions; the processor executes the computer program instructions to implement the battery swapping scheduling method as described in the first aspect.
[0034] Fourthly, this application provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the battery swapping scheduling method as described in the first aspect.
[0035] Fifthly, this application provides a computer program product in which instructions, when executed by the processor of an electronic device, cause the electronic device to perform the battery swapping scheduling method as described in the first aspect.
[0036] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0037] The features, advantages, and technical effects of exemplary embodiments of this application will now be described with reference to the accompanying drawings.
[0038] Figure 1 This is a flowchart of a battery swapping scheduling method according to an embodiment of this application; Figure 2 This is a flowchart illustrating the first determination of the reward function value according to an embodiment of this application; Figure 3 This is a flowchart illustrating the second determination of the reward function value according to one embodiment of this application; Figure 4 This is a flowchart illustrating the scaling and smoothing of reward function values according to an embodiment of this application; Figure 5 This is a flowchart illustrating the determination of a target battery swapping action for a first vehicle according to an embodiment of this application. Figure 6 This is a flowchart illustrating the process of screening legitimate candidate battery swapping actions according to one embodiment of this application; Figure 7 This is a flowchart illustrating the construction process of training samples according to one embodiment of this application; Figure 8 This is a flowchart illustrating the training process of an intelligent agent according to an embodiment of this application; Figure 9 This is a flowchart illustrating the process of obtaining newly sampled target training samples according to one embodiment of this application; Figure 10 This is a flowchart illustrating the determination of the mixed sampling ratio according to one embodiment of this application; Figure 11 This is a flowchart illustrating the construction of a second training sample set according to one embodiment of this application; Figure 12 This is a schematic diagram of a battery swapping dispatching device according to an embodiment of this application; Figure 13 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application; the terms “comprising” and “having”, and any variations thereof, in the specification, claims, and foregoing description of the drawings are intended to cover non-exclusive inclusion.
[0041] In the description of the embodiments of this application, technical terms such as "first" and "second" are used only to distinguish different objects and should not be construed as indicating or implying relative importance or implicitly specifying the number, specific order, or primary and secondary relationship of the indicated technical features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly defined.
[0042] In this application, the term "embodiment" is used to mean that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.
[0043] In the description of the embodiments in this application, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " in this document generally indicates that the preceding and following related objects have an "or" relationship.
[0044] In the description of the embodiments of this application, the term "multiple" refers to two or more (including two), similarly, "multiple sets" refers to two or more (including two sets), and "multiple pieces" refers to two or more (including two pieces).
[0045] In the description of the embodiments of this application, the technical terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing the embodiments of this application and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the embodiments of this application.
[0046] In the description of the embodiments of this application, unless otherwise expressly specified and limited, technical terms such as "installation," "connection," "joining," and "fixing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0047] Due to current limitations in battery technology, electric heavy-duty trucks still generally face practical problems such as limited driving range and long charging time, which poses new challenges to scenarios such as ports and mines that require high continuity and high load operation.
[0048] To shorten energy replenishment time, battery swapping has become increasingly popular. This model allows vehicles to resume operation within minutes by quickly replacing fully charged battery packs at swapping stations, significantly improving operational continuity and vehicle uptime. One approach in related technologies is to determine the swapping order by setting fixed logic such as battery level thresholds and queuing rules. For example, vehicles with battery levels below a certain threshold will be given priority in the swapping queue, and swapping stations will operate on a "first-come, first-served" basis. This method is simple to implement and can ensure basic swapping order in small-scale, low-complexity scenarios. However, this method relies on preset rules and static logic, lacking adaptability to dynamic uncertainties in complex scenarios. In port operations, where there are numerous vehicles, randomly generated tasks, and highly dynamic states, simple rule-based scheduling cannot effectively avoid problems such as concentrated vehicle swapping and severe queuing, nor can it achieve globally optimal scheduling results.
[0049] One approach involves establishing a mathematical programming model that comprehensively considers factors such as vehicle battery power, task requirements, and battery swapping station capacity to seek a globally optimized scheduling scheme. Theoretically, this method can provide an optimized solution for battery swapping scheduling and improve overall efficiency at a certain scale. While this method, based on operations research algorithms, can obtain the theoretically optimal solution when the model size is small, its complexity and computation time increase dramatically with the number of vehicles and battery swapping stations. In environments like ports where task instructions are generated in real-time and states change frequently, this method often requires continuous reconstruction and re-solving of the optimization model, leading to excessive computational overhead and making it difficult to meet the real-time decision-making requirements at the second or minute level. Furthermore, this type of method lacks adaptability to the uncertainties of battery consumption and task arrival, and lacks the ability to handle dynamic disturbances.
[0050] In summary, existing battery swapping technologies still face significant bottlenecks in practical applications. For example, in scenarios such as ports, mines, and large enclosed logistics parks, a large number of heavy trucks simultaneously perform dynamically assigned tasks. The completion time and power consumption of these tasks are highly uncertain, easily leading to a concentrated surge in battery swapping demand in both time and space, resulting in queues and congestion at battery swapping stations. This phenomenon not only prolongs the time vehicles are away from operations due to battery swapping but may also cause delays in subsequent tasks, reducing overall operational efficiency. Therefore, how to efficiently and adaptively schedule and make real-time decisions for battery swapping of electric heavy trucks in dynamic and uncertain environments has become a critical issue that urgently needs to be addressed.
[0051] To enable efficient and adaptive battery swapping scheduling and real-time decision-making for vehicles in dynamic and uncertain environments, this application provides a battery swapping scheduling method, apparatus, device, computer-readable storage medium, and computer program product. In the solution provided by this application, vehicle data and battery swapping station data of all operating vehicles in a closed operating scenario are dynamically collected by responding to vehicle operation completion requests in real time. A multi-dimensional state vector is constructed and input into a trained intelligent agent's battery swapping action prediction model to obtain multiple reward function values. Based on these reward function values, the decision module in the intelligent agent filters target battery swapping actions, ensuring that the decision results take into account both individual vehicle needs and global optimization of battery swapping resources, thus achieving real-time and adaptive scheduling of battery swapping demand. This method overcomes the shortcomings of traditional rule-based scheduling, such as rigidity and poor real-time performance of operations optimization. It can make intelligent scheduling decisions based on the global state in dynamic and uncertain closed operating scenarios (such as ports and mines), thereby reducing queuing time and improving vehicle attendance and operational continuity.
[0052] The following describes the battery swapping scheduling method provided in the embodiments of this application. The server can serve as the execution entity for the method provided in the embodiments of this application.
[0053] In one embodiment, Figure 1 A flowchart of the battery swapping dispatch method is shown, such as... Figure 1 As shown, the method may include the following steps: S101 to S105: S101, in response to the first vehicle in the closed operation scenario reporting the completion of the operation task at the first moment, obtains vehicle data of all operation vehicles in the closed operation scenario and battery swapping station data of all battery swapping stations.
[0054] In S101, the closed operation scenario can be a fixed-range, large-scale, and high-frequency scenario such as ports, mines, and logistics parks. The first vehicle can be any electric heavy truck that completes its task and initiates a reporting request at the current moment. The first moment can be the specific time when the first vehicle completes its current task and reports it. Vehicle data can include information such as the remaining battery power, current location, energy consumption rate, and task completion time of all operating vehicles. Battery swapping station data can include operational status information such as the next idle time, single battery swap duration, and location of all battery swapping stations.
[0055] It should be noted that in closed operation scenarios, there are uncertainties in the vehicle operation completion time and power consumption, and the demand for battery swapping is prone to sudden surges. Relying solely on data from a single vehicle cannot achieve global optimization scheduling. Therefore, it is necessary to simultaneously acquire data from all vehicles and battery swapping stations to provide a complete basis for subsequent status characterization and decision-making, and to avoid scheduling imbalances due to incomplete information.
[0056] In practice, the operational status of all vehicles within the closed operation scenario can be monitored in real time. When a work completion report request is received from the first vehicle, the data collection process is initiated. For example, the real-time operating data of all operating vehicles can be retrieved through the vehicle management system, and the real-time operating data of all battery swapping stations can be collected through the battery swapping station management system, ensuring the completeness and timeliness of data collection and providing data support for subsequent steps.
[0057] S102, based on vehicle data and battery swapping station data, constructs the first multi-dimensional state vector of the closed operation scenario at the first moment.
[0058] In S102, the first multidimensional state vector is used to comprehensively characterize the standardized data set of the scene state at the first moment, which integrates individual vehicle attributes and global environmental information. The dimensions are fixed and do not change with the scale of the scene.
[0059] For example, the first multi-dimensional state vector may include the multi-dimensional state vector of the first vehicle and the global state vector in the closed operation scenario. The multi-dimensional state vector of the first vehicle may include: the first state of charge (SOC) value of the first vehicle and the time required for the first vehicle to travel to each battery swapping station. The global state vector may include: the next idle time of each battery swapping station, the duration of a single battery swap, statistical information on the SOC distribution of all operating vehicles, and predicted demand information within a preset time period.
[0060] To better understand the above multidimensional state vector, the following example illustrates a closed operation scenario involving three battery swapping stations.
[0061] In this example scenario, the state At the moment of decision Defined as The specific composition is as follows: The first state of charge (SOC) value H of the current vehicle (the first vehicle) (for example, H can be a 1-dimensional vector): H reflects the current state of charge of the vehicle's battery. The time it takes for the first vehicle to reach each battery swapping station. (For example, there are 3 battery swapping stations.) (Can be a 3D vector) in, For the current vehicle to reach the th Travel time of each battery swapping station and These are the preset global minimum and maximum travel times in a closed operation scenario, used for normalization.
[0062] Next idle time for each battery swapping station (For example, there are 3 battery swapping stations.) (Can be a 3D vector) This represents the offset of the next available service time for each battery swapping station relative to the current time, and is logarithmically transformed to compress the numerical range.
[0063] in, This indicates the next idle time of the k-th battery swapping station, that is, the specific time when the station can serve new vehicles after completing the battery swapping of the currently queued vehicles; The current decision-making moment is the moment when the first vehicle completes its operation and initiates the battery swapping decision. This is a preset time limit, usually the total daily operating time in a closed operation scenario (e.g., 1440 minutes of daily operation time at a port).
[0064] SOC distribution statistics F for all operating vehicles (for example, the distribution intervals can be 5, and F can be a 5-dimensional vector): The entire team The vehicles are divided into five equal-length intervals based on their State of Charge (SOC): 1-20%, 21-40%, 41-60%, 61-80%, and 81-100%. The number of vehicles in each interval is then counted. and divide by the total. Normalize.
[0065] Predicted demand information M within a preset time period (for example, the distribution intervals can be 5, and M can be a 5-dimensional vector): Count the number of vehicles that will complete their tasks and go offline within the next 10 minutes. Similarly, it is divided and normalized according to the SOC interval.
[0066] It should be noted that the original vehicle data and battery swapping station data are scattered and have inconsistent formats, making them unsuitable for direct input into the intelligent agent for computation. By constructing standardized multidimensional state vectors, the discrete data is structurally integrated, preserving key information about individual attributes and global states while ensuring a fixed input dimension. This adapts to the input requirements of the intelligent agent model, enabling the transformation of data into decision-making basis.
[0067] By fusing local vehicle information with global information from a closed operation scenario using multi-dimensional state vectors, the system accurately captures the vehicle's core attributes (SOC, driving time) and integrates the battery swapping station's resource status (idle time, battery swapping duration) with the global supply and demand status (SOC distribution, predicted demand). This provides a complete context for decision-making, enabling the agent to not only consider the current vehicle's battery level and location when making decisions based on multi-dimensional state vectors, but also to perceive the battery swapping station's busy / idle status, the overall battery distribution of the fleet, and short-term demand predictions, thereby making more systematic and better scheduling decisions.
[0068] S103, input the first multi-dimensional state vector into the pre-trained battery swapping action prediction model in the intelligent agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle.
[0069] In S103, candidate battery swapping actions can be a set of scheduling behaviors that the agent can choose from, including two categories: not swapping batteries and going to a designated battery swapping station. The reward function aims to guide the agent to learn efficient and safe scheduling strategies. The reward function value is a quantitative indicator used to evaluate the merits of candidate battery swapping actions, calculated by comprehensively considering factors such as battery swapping time and vehicle battery status. The agent is trained based on a deep reinforcement learning algorithm and has the ability to evaluate the value of actions based on the scene state and output a decision.
[0070] In practice, the constructed first multi-dimensional state vector can be input into the battery swapping action prediction model trained to a stable performance in the agent. The agent analyzes the state vector through the built-in network model (battery swapping action prediction model), and calculates the reward function value corresponding to each candidate battery swapping action according to the preset reward function calculation rules, ensuring that the value of each action is accurately quantified, providing data support for subsequent action decisions.
[0071] S104, through the decision-making module in the intelligent agent, the battery swapping action of the first vehicle is decided based on multiple reward function values to obtain the target battery swapping action of the first vehicle.
[0072] In S104, the target battery swapping action is the action with the optimal reward function value selected from all candidate battery swapping actions, and which meets the requirements for safe operation of the scenario.
[0073] For example, taking the above-mentioned closed operation scenario including 3 battery swapping stations as an example, the action space can be defined as a discrete set. ,in Indicates no battery replacement. Indicates going to the The battery swapping station is used for battery swapping.
[0074] In practice, an action masking mechanism can be used to first screen legitimate candidate battery swapping actions, eliminating unreasonable actions such as not swapping batteries for vehicles with low battery levels or swapping batteries for vehicles with high battery levels. Within the scope of legitimate actions, the reward function values corresponding to each action are compared, and the action with the largest reward function value is selected as the target battery swapping action, ensuring that the decision result is both legitimate and safe while achieving global optimization.
[0075] S105, based on the target battery swapping action, performs battery swapping scheduling for the first vehicle.
[0076] In S105, battery swapping scheduling is the process of sending a scheduling instruction to the first vehicle based on the target battery swapping action, guiding the vehicle to perform battery swapping or continue the operation.
[0077] In practice, when the target battery swapping is not scheduled for a swap, a new task instruction can be sent to the first vehicle, specifying the start time and content of the new task, and updating the vehicle's task end time. When the target battery swapping is to travel to a designated battery swapping station, a battery swapping instruction can be sent to the first vehicle, including the location of the target battery swapping station and the optimal driving route, while simultaneously updating the operational status data of the target battery swapping station to ensure the orderly connection of battery swapping services.
[0078] By precisely executing dispatch instructions, vehicles can be ensured to travel to the battery swapping station or continue operations along the optimal route, avoiding queuing and congestion at the battery swapping station, improving operational continuity and resource utilization, and ultimately achieving the dispatching goals.
[0079] Based on the solutions defined in S101 to S105 above, it can be understood that in this embodiment, by responding to the vehicle's job completion request in real time, vehicle data and battery swapping station data of all operating vehicles in the closed operating scenario are dynamically collected. A multi-dimensional state vector is constructed and input into the battery swapping action prediction model in the trained agent to obtain multiple reward function values. Based on these multiple reward function values, the decision module in the agent filters the target battery swapping action, ensuring that the decision result takes into account both the individual needs of the vehicle and the global optimization of battery swapping resources, thus achieving real-time and adaptive scheduling of battery swapping demand. This method overcomes the shortcomings of rigid rule scheduling and poor real-time performance of operations optimization in traditional methods. It can make intelligent scheduling decisions based on the global state in dynamic and uncertain closed operating scenarios (such as ports and mines), thereby reducing queuing time and improving vehicle attendance and operational continuity.
[0080] The implementation process of the method provided in the embodiments of this application is described below.
[0081] In one embodiment, Figure 2 The first step in determining the reward function value is illustrated, as follows: Figure 2 As shown, the process includes the following steps: S201, the first multi-dimensional state vector is input into the battery swapping action prediction model in the pre-trained agent to obtain different candidate battery swapping actions of the first vehicle.
[0082] S202, determine the type of candidate battery swapping action. If the candidate battery swapping action is to go to battery swapping, execute S203; if the candidate battery swapping action is to not swap, execute S205.
[0083] S203, for each candidate battery swapping action: based on the first moment, the time required for the first vehicle to travel to the corresponding battery swapping station, and the battery swapping duration of the battery swapping station, determine the total offline time from the first moment to the completion of the battery swapping when the first vehicle is swapping its battery at the battery swapping station.
[0084] In S203, when a candidate battery swapping station is selected as the destination, the time required for the first vehicle to travel to the corresponding battery swapping station is the estimated time for the first vehicle to travel from its current location to the target battery swapping station, which can be calculated based on the distance between the vehicle's location and the location of the battery swapping station, as well as the vehicle's speed. The battery swapping time at a battery swapping station is a fixed standard time for the station to complete a battery swap for a single vehicle, which can be determined by the performance of the battery swapping station's equipment and operating procedures. The total offline time is the total time taken for the first vehicle from the first moment until it completes the battery swap and restores its operational capability; it is the core indicator for evaluating the impact of the battery swapping operation on operational continuity.
[0085] In practice, the specific values at the first moment can be obtained first. Then, based on the current position of the first vehicle and the target battery swapping station, and combined with the preset driving speed, the travel time can be calculated. The fixed battery swapping duration parameters of the battery swapping station are retrieved, and the estimated start time of the battery swapping station is calculated. This time is the larger of the next idle time of the battery swapping station and the estimated arrival time of the vehicle. The estimated end time of the battery swapping (…) is then used… Subtract the first moment Get the total offline time To ensure that the time calculation accurately reflects the time cost of the entire battery swapping process, that is... .
[0086] S204. Based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and energy reward coefficient, determine the reward function value of the candidate battery swapping action.
[0087] In S204, the first state of charge (SOC) value can be the remaining charge of the first vehicle at the first moment, that is, the proportion of the battery's available charge to its nominal capacity.
[0088] Preset average offline time The average reasonable battery swapping time set by enterprises based on their operational experience serves as a benchmark for evaluating battery swapping efficiency; penalties are imposed when the actual vehicle off-line time exceeds this value, while rewards are given when it is less than this value.
[0089] Time reward coefficient To be used for quantifying total offline time Compared with the preset average offline time The weighting parameters for the difference are such that the shorter the offline time, the higher the reward.
[0090] The energy reward coefficient is a weighted parameter used to quantify the impact of the first state of charge (SOC) value on the necessity of battery swapping. The energy reward coefficient includes the first energy reward coefficient. Second power reward coefficient and the third electricity reward coefficient β i First battery reward coefficient Greater than the second battery reward coefficient The third electricity reward coefficient β i The third energy reward coefficient is between the first and second energy reward coefficients, and decreases linearly with the increase of the first state of charge (SOC) value (the closer the first SOC value is to the first preset SOC, the closer the third energy reward coefficient is to the first energy reward coefficient; the closer the first SOC value is to the second preset SOC, the closer the third energy reward coefficient is to the second energy reward coefficient, where the second preset SOC is greater than the first preset SOC). The energy reward coefficient is higher when swapping batteries with low energy, so as to encourage vehicles with low energy to swap batteries first.
[0091] reward function value It is a reward coefficient when the battery is not swapped, encouraging vehicles with high battery levels to continue operating.
[0092] In practice, it can be first determined whether the first state of charge (SOC) value is lower than a preset first SOC. If it is lower than the preset first SOC, a higher first energy reward coefficient is used, combined with the difference between the total offline time and the preset average offline time, and a time reward coefficient, to calculate the initial reward value. If it is higher than or equal to a preset second SOC, a lower second energy reward coefficient is used to calculate the initial reward value. If it is between the preset first SOC and the preset second SOC, a dynamic third energy reward coefficient is used to calculate the initial reward value, and the initial reward value is scaled and smoothed to avoid extreme values affecting the evaluation stability. Finally, the reward function value of the candidate battery swapping action is obtained, providing a quantitative basis for action decision-making.
[0093] By quantifying the total time a vehicle takes to go offline from the current moment until the battery swap is completed, and combining it with factors such as the preset average offline time and vehicle battery level, a highly guiding reward signal is constructed. This enables the intelligent agent to learn how to reduce the time a vehicle is away from operation while rationally arranging the timing of battery swaps and site selection, thereby improving overall scheduling efficiency.
[0094] S205, determine the reward function value for not swapping batteries based on the first state of charge (SOC) value.
[0095] In S205, the reward function value for not swapping batteries Numerical indicators used to quantify the merits of non-battery swapping actions, and to evaluate the overall contribution of this action to operational continuity, safety, and resource utilization.
[0096] In practice, the SOC (State of Charge) of the first vehicle at the decision-making moment can be obtained first to ensure real-time and accurate data. Then, the preset non-battery swapping incentive coefficient can be retrieved. This coefficient is set based on the operational needs of a closed-loop work environment and is used to balance the correlation between the SOC value and the reward. The SOC value is then compared with the non-battery swapping reward coefficient. Multiplying the values yields the reward function value for the non-battery swapping action, providing a quantitative basis for subsequent action decisions. This ensures that vehicles with high battery levels receive high rewards to encourage continued operation, while vehicles with low battery levels receive low rewards to guide the selection of battery swapping.
[0097] For example, in the case where the decision is made not to swap batteries for the vehicle: For the "no battery swapping" action, the reward is determined directly based on the vehicle's SOC value, ensuring both assessment efficiency (no complex calculations required) and alignment with actual operational needs (continuing to operate vehicles with high battery levels is more valuable). By directly linking the SOC value to the reward, vehicles with high battery levels are encouraged to continue operating (receiving higher rewards), while vehicles with low battery levels are prevented from forced operation (receiving lower rewards), thus promoting a reasonable behavior pattern of "operating with high battery levels and swapping batteries with low battery levels." At the same time, there is no need to calculate additional parameters such as driving time and battery swapping time, allowing for rapid reward assessment of the "no battery swapping" action. Combined with the quantitative assessment of the "go to battery swapping" action, this ensures decision-making quality while improving real-time response speed.
[0098] In one embodiment, the power reward coefficient includes a first power reward coefficient, a second power reward coefficient, and a third power reward coefficient. The first power reward coefficient is greater than the second power reward coefficient, and the third power reward coefficient is between the first power reward coefficient and the second power reward coefficient, and decreases linearly with the increase of the first state of charge (SOC) value. Figure 3 The second process for determining the reward function value is shown, as follows: Figure 3 As shown, the process includes the following steps: S301, determine whether the first state of charge (SOC) value is less than the preset first SOC. If the first state of charge (SOC) value is less than the preset first SOC, execute S302. If the first state of charge (SOC) value is greater than or equal to the preset second SOC, execute S303. If the first state of charge (SOC) value is greater than or equal to the preset first SOC and less than the preset second SOC, execute S304.
[0099] In S301, the preset first SOC and preset second SOC are power thresholds set according to the work intensity, vehicle energy consumption rate and battery swapping station distribution in the closed operation scenario, and are used to distinguish the urgency of vehicle battery swapping.
[0100] In practice, the specific value of the first state of charge (SOC) of the first vehicle can be obtained, and a preset SOC threshold parameter can be retrieved. The first SOC value is compared with the preset SOC. If the first SOC value is less than the preset first SOC value, it indicates that the vehicle's battery is low and the battery swapping requirement is urgent, and S302 is executed. If the first SOC value is greater than or equal to the preset second SOC value, it indicates that the vehicle's battery is sufficient and the battery swapping requirement is not urgent, and S303 is executed. If the first SOC value is greater than or equal to the preset first SOC and less than the preset second SOC, the vehicle does not need to be forced to swap batteries (to avoid the safety risk of low battery) nor does it need to prohibit battery swapping (to reserve battery swapping flexibility to cope with subsequent task fluctuations).
[0101] S302, based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and first energy reward coefficient, determine the reward function value of the candidate battery swapping action.
[0102] In practice, if the first state of charge (SOC) value is less than the preset first SOC value, it indicates that the vehicle's battery power is low and the need for battery swapping is urgent. In this case, the first battery power reward coefficient is applied. Specifically, you can retrieve the total offline time. Preset average offline time Time reward coefficient and the first battery reward coefficient parameter Calculate the total offline time. Compared with the preset average offline time The difference, i.e. Multiply the difference by the time reward coefficient. The time dimension reward component is obtained, that is Calculate the difference between 100 and the first state of charge (SOC) value. Multiply the difference by the first power reward coefficient. Receive a reward component based on battery power, i.e. The two reward components are added together to obtain the reward function value of the candidate battery swapping action in this scenario. .
[0103] It should be noted that the difference is calculated using a preset average offline time. Subtract total offline time Total offline time The larger the value, the higher the corresponding reward function value. The smaller the value, the better. In practical applications, a larger reward function value is better, as it corresponds to a shorter total offline time. When the decision-making module in the intelligent agent makes a decision on the battery swapping action of the first vehicle based on multiple reward function values, it will select the battery swapping action with the larger reward function value, that is, select a battery swapping station with a shorter total offline time for the vehicle.
[0104] For example, when a vehicle decides to choose battery swapping and its State of Charge (SOC) is below a set threshold: S303, based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and second energy reward coefficient, determine the reward function value of the candidate battery swapping action.
[0105] In practice, if the first state of charge (SOC) value is greater than or equal to the preset second SOC value, it indicates that the vehicle has sufficient battery power and the battery swapping need is not urgent, and the second battery power reward coefficient is applied. Using the same calculation logic as S302, the reward function value of the candidate battery swapping action in this scenario is obtained. .
[0106] For example, when a vehicle decides to choose battery swapping and its State of Charge (SOC) is not lower than a set threshold: S304. Based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and third energy reward coefficient, determine the reward function value of the candidate battery swapping action.
[0107] In practical implementation, if the State of Charge (SOC) value is between a preset first SOC and a preset second SOC, it means that the vehicle neither needs to be forced to swap batteries (to avoid the safety risks of low battery) nor needs to prohibit battery swapping (to reserve battery swapping flexibility to cope with subsequent task fluctuations), and a third battery reward coefficient is adopted. Using the same calculation logic as S302, the reward function value of the candidate battery swapping action in this scenario is obtained. .
[0108] For example, when the decision-making vehicle neither requires mandatory battery swapping nor prohibits battery swapping, and the first state of charge (SOC) value is between a preset first SOC and a preset second SOC: By setting the first power reward coefficient ( ) greater than the second electricity reward coefficient ( This approach prioritizes battery swapping for low-battery vehicles, ensuring that low-battery vehicles receive higher rewards and preventing task interruptions or safety risks due to depleted battery power. Simultaneously, it reduces the reward weight for battery swapping for high-battery vehicles, minimizing the resource consumption of swapping stations due to ineffective swapping by high-battery vehicles and ensuring that swapping resources are allocated to low-battery vehicles that genuinely need them, thus improving resource utilization. A dynamic third reward coefficient (β) is established, falling between the first and second battery reward coefficients and decreasing linearly with increasing State of Charge (SOC) value. i This approach balances the necessity of battery swapping with the rationality of resource utilization by allowing vehicles with intermediate battery levels to neither be forced to swap batteries (avoiding safety risks associated with low battery levels) nor prohibited from swapping batteries (reserving flexibility for battery swapping to cope with fluctuations in subsequent tasks). By differentiating reward coefficients based on battery level thresholds, unreasonable behaviors such as "not swapping batteries when low" or "frequent swapping when high," are avoided. This ensures that the scheduling strategy conforms to vehicle energy replenishment patterns and adapts to the high-continuity operational needs of closed work scenarios.
[0109] To prevent the reward value range from being too large and affecting training stability, in one embodiment, the first multi-dimensional state vector is input into the pre-trained battery swapping action prediction model of the agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The embodiment also includes a process of scaling and smoothing the reward function values. The scaling and smoothing process of the reward function values can be achieved through... Figure 4 The method shown determines this; specifically, the process includes the following steps: S401, scale and smooth the reward function value to obtain the processed reward function value.
[0110] In S401, scaling can adjust the numerical range of the reward function value using preset coefficients, eliminating dimensional differences caused by parameters of different dimensions and ensuring that the reward value is within the range suitable for model training. Smoothing can use non-linear transformations to reduce the impact of extreme reward values, making reward value changes more continuous and avoiding the impact of abrupt numerical changes on model training.
[0111] In practice, the preset scaling factor can be retrieved first. This coefficient, set based on model training requirements and scene parameter characteristics, is used to map the original reward function value to the target numerical range. With scaling factor Multiplication completes the scaling process. Then, the scaled reward value is smoothed using the hyperbolic tangent function. This function's non-linear properties compress extreme values, ensuring the processed reward value falls within a fixed range. The final processed reward function value is obtained. This approach preserves the differences in quality between different actions while avoiding the interference of extreme values on training stability, thus providing a reliable basis for the agent's accurate decision-making.
[0112] For example, the original reward After scaling and smoothing, the result is : ,in This is the scaling factor.
[0113] For the "no battery swapping" action, the reward is determined directly based on the vehicle's SOC value, ensuring both assessment efficiency (no complex calculations required) and alignment with actual operational needs (continuing to operate vehicles with high battery levels is more valuable). By directly linking the SOC value to the reward, vehicles with high battery levels are encouraged to continue operating (receiving higher rewards), while vehicles with low battery levels are prevented from forced operation (receiving lower rewards), thus promoting a reasonable behavior pattern of "operating with high battery levels and swapping batteries with low battery levels." At the same time, there is no need to calculate additional parameters such as driving time and battery swapping time, allowing for rapid reward assessment of the "no battery swapping" action. Combined with the quantitative assessment of the "go to battery swapping" action, this ensures decision-making quality while improving real-time response speed.
[0114] In one embodiment, the following can be used: Figure 5 The method shown is used to determine the target battery swapping action of the first vehicle, such as Figure 5 As shown, the process includes the following steps S501 to S502: S501 uses an action masking mechanism to filter out legitimate candidate battery swapping actions from multiple candidate battery swapping actions.
[0115] In S501, the action masking mechanism refers to a mechanism that dynamically judges the legitimacy of candidate battery swapping actions through preset rules, filtering out unreasonable or dangerous actions. This is used to narrow down the effective action space and improve decision-making safety and efficiency. The action masking mechanism can directly filter out unreasonable battery swapping actions, preventing the agent from wasting exploration time in the wrong direction. Legitimate candidate battery swapping actions refer to battery swapping actions that comply with vehicle operation safety requirements and battery swapping resource utilization rules, and will not lead to safety risks or resource waste.
[0116] In one embodiment, the following can be used: Figure 6 The method shown filters out legitimate candidate battery swapping actions, such as Figure 6 As shown, the process includes S601-S603: S601, determine whether the first state of charge (SOC) value is less than the first SOC threshold or greater than the second SOC threshold. If the first state of charge (SOC) value is less than the first SOC threshold, execute S602. If the first state of charge (SOC) value is greater than the second SOC threshold, execute S603.
[0117] In S601, the first SOC threshold is a safety critical energy level set based on the vehicle's energy consumption rate, the distance of the work scenario, and the distribution of battery swapping stations. Below this value, it indicates that the vehicle's battery power is insufficient to support the next round of work, and a battery swap must be performed. The second SOC threshold is a battery swapping limit set to avoid resource waste. Above this value, it indicates that the vehicle's battery power is sufficient, and multiple rounds of work can be completed without battery swapping.
[0118] S602, from multiple candidate battery swapping actions, eliminate non-battery swapping actions to obtain the battery swapping action to be swapped.
[0119] In S602, after confirming that the first State of Charge (SOC) value is less than the first SOC threshold, an action masking process is initiated. For all candidate battery swapping actions, actions that do not involve battery swapping are identified and masked, while all actions involving heading to the battery swapping station are retained. This operation forces low-battery vehicles to only select battery swapping-related actions, preventing them from running out of power due to not swapping, thus ensuring operational safety and continuity. For example, for the current state... If the vehicle's battery level is below the safe threshold Then the "no battery swapping" action will be disabled.
[0120] S603 filters out battery swapping actions from multiple candidate battery swapping actions, resulting in a no-battery-swap action.
[0121] In S603, after confirming that the first State of Charge (SOC) value is greater than the second SOC threshold, the action blocking process is initiated. For all candidate battery swapping actions, all actions to proceed to the battery swapping station are identified and blocked, retaining only actions that do not involve battery swapping. For example, if the vehicle's battery level is higher than the prohibited battery swapping threshold... If the battery level is not high, the "go to battery swap" action will be blocked, and the "do not swap battery" action will be retained as a valid candidate for battery swapping. This operation avoids high-battery vehicles occupying battery swapping resources, frees up battery swapping space for low-battery vehicles, optimizes battery swapping resource allocation, and improves overall operational efficiency.
[0122] For example, for state mask function Defined as: By using a first SOC threshold (safety threshold) and a second SOC threshold (battery swapping prohibition threshold), the boundaries of action legality are quantified, avoiding subjective judgments in the action masking mechanism and ensuring the consistency of the screening logic. At the same time, vehicles with low battery levels (below the first SOC threshold) are forced to choose "go for battery swapping," fundamentally eliminating safety issues such as work interruptions and breakdowns caused by vehicles running out of power. Vehicles with high battery levels (above the second SOC threshold) are forced to choose "not to swap batteries," preventing high-battery vehicles from occupying battery swapping station resources, freeing up battery swapping space for low-battery vehicles, and reducing queuing congestion at battery swapping stations.
[0123] S502: Select the battery swapping action with the largest reward function value from the legal candidate battery swapping actions to obtain the target battery swapping action for the first vehicle.
[0124] In S502, the target battery swap is the action with the optimal reward function value selected from the legitimate candidate battery swap actions, and it can achieve global scheduling optimization.
[0125] The agent's final decision will be based on the set of legal actions. Choosing the action with the highest Q-value significantly improves learning efficiency and safety. The Q-value, a term in reinforcement learning, represents the expected value of all cumulative rewards that can be obtained by choosing that action.
[0126] By using an action masking mechanism to block invalid or dangerous candidate battery swapping actions, the agent is prevented from choosing unreasonable behaviors such as "not swapping batteries when the battery is low" or "swapping batteries when the battery is high," thus reducing scheduling risks. At the same time, invalid options other than legal actions can be excluded, reducing the decision-making computation load of the agent and ensuring that the target action is output within seconds / minutes, meeting the real-time scheduling requirements of closed operation scenarios. Furthermore, by selecting the action with the largest reward function value within the range of legal actions, system failures caused by illegal actions are avoided, while ensuring the optimality of the decision result, thus balancing safety and efficiency.
[0127] In one embodiment, after the first vehicle completes its battery swap, training samples can be constructed for training the agent. Figure 7 The process of constructing training samples is shown, such as... Figure 7 As shown, the process includes the following steps: S701, after the first vehicle completes its battery swap, update the state vector of the closed operation scenario at the second moment, and the second multidimensional state vector; the second moment is the moment when the first vehicle completes its battery swap.
[0128] In S701, the second moment is the specific time when the first vehicle completes the battery swapping operation and restores its operational capability. The second multi-dimensional state vector is a standardized data set used to characterize the overall state of the closed operation scenario at the second moment, and is the updated result of the state at the first moment after the battery swapping action.
[0129] In practice, after the first vehicle completes its battery swap, the current time can be recorded as the second moment. The latest data after the battery swap is executed is retrieved: the first State of Charge (SOC) value is updated to full charge, along with its new task completion time; the next idle time of the target battery swapping station is updated as the second moment; the time difference from the first moment to the second moment is calculated, and the SOC values of all vehicles are updated based on the energy consumption rate of each online vehicle; the SOC distribution information of the entire fleet and the predicted battery swapping demand data for the future preset time period are re-calculated. Following the same structure and processing rules as the first multidimensional state vector, the updated data is integrated and normalized to form the second multidimensional state vector.
[0130] S702 constructs training samples for training the agent based on the first multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second multidimensional state vector.
[0131] In S702, the four types of data are structurally combined in a fixed order of "first multidimensional state vector - target battery swapping action - reward function value - second multidimensional state vector" to form a complete training sample. The training sample is stored in a preset training sample set. If the sample set has a hierarchical structure, it is stored in the heuristic experience pool or the exploration experience pool according to the source of the sample (generated by heuristic rules or generated by exploration of real scene) to provide data support for the subsequent training of the agent.
[0132] On the one hand, by recording the state transitions of the entire battery swapping scheduling process (first multi-dimensional state vector → target battery swapping action → reward value → second multi-dimensional state vector), high-quality training samples from real-world scenarios are continuously generated, providing data support for agent optimization. On the other hand, samples from real-world scenarios contain uncertainties in dynamic environments (such as fluctuations in power consumption and changes in tasks), making them more practical than simulation samples. This allows the agent to continuously learn and adapt to the dynamic changes in the scenario. At the same time, the scheduling process also serves as a sample collection process, enabling the agent to continuously iterate and upgrade in practical applications, gradually improving decision-making accuracy and adapting to the long-term operational needs of closed operation scenarios.
[0133] In one embodiment, the agent is trained using machine learning, where... Figure 8 The training process of the agent is shown, such as Figure 8 As shown, the process includes the following steps: S801, sample from the training sample set to obtain the target training sample.
[0134] For example, each training sample in the training sample set includes: the first multidimensional state vector sample, the battery swapping action sample, the reward function value sample, and the second multidimensional state vector sample corresponding to the closed operation scenario sample under multiple different first time-time samples, where the second time-time sample is the time corresponding to the closed operation scenario sample when the battery swapping action sample is completed.
[0135] S802, the first multi-dimensional state vector sample is input into the agent to be trained to obtain the battery swapping prediction action.
[0136] In S802, the agent to be trained can be trained using the Double Dueling Deep Q-Network (D3QN) algorithm.
[0137] In practice, the first multi-dimensional state vector sample from the target training sample can be input into the agent to be trained. The agent extracts and parses the features of the state vector through the network layer, calculates the Q value of each candidate battery swapping action in combination with the current model parameters, filters the legal actions through the action masking mechanism, selects the action with the largest Q value as the battery swapping prediction action, and outputs it to the subsequent steps.
[0138] The Dueling network structure decomposes the Q-value into a state value function. and dominance function This structure can effectively evaluate the value of a state itself and the relative advantage of a specific action, and its output is: (11) The Double Q-learning mechanism is used to calculate the target Q value to mitigate overestimation. (12) in, and These are the parameters of the online network and the target network, respectively. Both "online network" and "target network" are technical terms in reinforcement learning. The online network is responsible for interacting with the environment in real time, executing actions, and collecting experience data. It learns the optimal policy by updating its network parameters. Its parameters are updated immediately after each training step for immediate policy optimization and exploration. The target network periodically replicates the parameters of the online network to calculate the target Q-value, reducing training fluctuations. Its parameters are kept constant for a longer period before updating to ensure the stability of the learning objective and avoid divergence problems caused by frequent changes.
[0139] S803 determines the predicted value of the first reward function based on the battery swapping prediction action and the first multi-dimensional state vector sample.
[0140] In S803, the SOC value of the corresponding vehicle can be extracted from the first multi-dimensional state vector sample. Combined with the battery swap prediction action type (battery swap or no battery swap), the preset reward function calculation rules are called. If it is a battery swap action, it needs to be calculated based on the total offline time implicit in the sample, the preset average offline time, and the corresponding energy reward coefficient. If it is a no-battery swap action, it is calculated based on the SOC value and the no-battery swap reward coefficient, and finally the predicted value of the first reward function is obtained.
[0141] S804, Based on the predicted value of the first reward function and the sample reward function value, determine the loss function value.
[0142] In S804, the loss function value is an indicator that quantifies the difference between the predicted value of the first reward function and the sample reward function value. It reflects the degree of deviation between the model's current decision and the true optimal decision, and the Huber loss function can be used.
[0143] In practice, reward function value samples from the target training samples can be retrieved and input into the Huber loss function along with the predicted value of the first reward function. The degree of deviation between the two is then calculated to obtain the loss function value. The larger this value, the greater the current prediction deviation of the model, requiring optimization through parameter adjustment; conversely, the smaller the value, the closer the model's decision is to the true optimum.
[0144] S805: Determine whether the loss function value meets the iteration stopping condition. If the iteration stopping condition is not met, execute S807. If the iteration stopping condition is met, execute S806.
[0145] S806, thus obtaining the trained agent.
[0146] When the iteration stopping condition is met, the model training is stopped, all network parameters of the agent to be trained are frozen, and it is identified as a trained agent that can be used for actual battery swapping scheduling in subsequent closed operation scenarios.
[0147] S807 adjusts the model parameters of the agent to be trained based on the loss function value.
[0148] In S807, based on the currently calculated loss function value, a gradient descent optimization algorithm is used to backpropagate to each network layer of the agent, adjusting the network parameters according to a preset learning rate. The adjustment process aims to minimize the loss function value. With the goal of improving the accuracy of model prediction and the rationality of decision-making, we optimize the parameters of state feature extraction, Q-value calculation and other processes.
[0149] For example, network parameters are updated by minimizing the Huber loss: Where N is the number of samples; The true reward value (i.e., the reward function value sample) of the i-th training sample is the target value predicted by the model, which reflects the true value of the action; It is the state vector in the i-th training sample (i.e., the first multidimensional state vector sample), which describes the environmental state in the i-th scene; θ represents the action in the i-th training sample (i.e., the battery swapping action sample), and θ represents the actual battery swapping / non-battery swapping action performed in the i-th scenario; θ is the model parameter. The agent inputs the state of the i-th sample based on the current parameter θ. and actions The Q-value (predicted action value) output afterward represents the model's estimate of the value of that state-action combination.
[0150] S808 resamples the training sample set to obtain the newly sampled target training sample; and returns the first multi-dimensional state vector sample to be trained into the agent until the iteration stopping condition is met.
[0151] In S808, the newly sampled target training samples are scene data samples that are re-extracted from the training sample set for the next model iteration.
[0152] In practice, following the same sampling rules as S801, samples covering diverse scenarios are re-extracted from the training sample set to form new target training samples. Then, returning to step S802, the first multi-dimensional state vector sample of the new samples is input into the adjusted agent to be trained, and the subsequent prediction, loss calculation, and parameter adjustment processes are repeated until the iteration stopping condition is met.
[0153] By adjusting model parameters based on the loss function value, the agent continuously learns the mapping relationship between "state-action-reward," gradually approaching the optimal scheduling strategy and ensuring the effectiveness and convergence of the agent's training. Simultaneously, the training sample set covers multiple state vectors and action combinations at different times, enabling the agent to learn scheduling patterns in diverse scenarios, avoiding overfitting to a single scenario, and improving the model's generalization ability and robustness. Furthermore, by controlling the training pace through iteration stopping conditions (such as the loss function value falling below a threshold or the number of iterations reaching a target), overtraining or undertraining is avoided, ensuring the agent's stable performance.
[0154] In one embodiment, it can be achieved by Figure 9 The target training samples obtained by the new sampling are shown in the manner described, where, Figure 9 The process of obtaining the target training samples from the newly sampled data is shown, such as... Figure 9 As shown, the process includes the following steps: S901, counts the current iteration number.
[0155] In S901, the current iteration number It is the number of times a single "sampling-prediction-loss calculation-parameter adjustment" cycle has been completed during the training of the agent, and it is the core indicator for measuring the training progress.
[0156] S902, determine the ratio of the current iteration number to the preset iteration number.
[0157] In S902, the number of iterations is preset. The total number of training iterations, determined based on model training requirements, sample size, and target performance, serves as a benchmark for judging the training phase. For example, It can be a combination of decay steps.
[0158] In practice, the preset number of iterations is retrieved, and the current number of iterations obtained from S901 is divided by the preset number of iterations to obtain the ratio. The closer this ratio is to 0, the earlier the training is in the initial stage; the closer it is to 1, the closer the training is to completion, providing a quantitative basis for the dynamic adjustment of the sampling ratio.
[0159] For example, the ratio is .
[0160] S903, based on the ratio and the minimum retention ratio of the first training sample set, determine the mixed sampling ratio of the first training sample set and the second training sample set.
[0161] In S903, the training sample set includes a first training sample set and a second training sample set. The training samples in the first training sample set are constructed based on expert experience and are also called a heuristic experience pool. It stores high-quality decision trajectories generated by prior heuristic rules.
[0162] The training samples in the second training sample set are samples generated during simulated battery swapping in simulated closed operation scenarios, also known as the exploration experience pool. , storing new samples generated by the online interaction between the intelligent agent and the environment.
[0163] Minimum retention ratio To ensure training stability, the first training sample set is set to have the lowest sampling percentage throughout the entire training process, thus avoiding a complete departure from expert experience.
[0164] It should be noted that, as the number of iterations increases, the training process of the agent can be divided into three stages: Initial stage (when the number of iterations is relatively small): Mainly from Sampling allows the agent to quickly learn a better initial policy.
[0165] Transition phase (iteration count gradually increases): gradually decrease Introducing more from The exploration samples enable the agent to adapt to the uncertainty and diversity of the environment.
[0166] Mature stage (when there are many iterations): The intelligent agent mainly learns through its own exploration, gradually breaking free from the limitations of heuristic rules and learning to cope with diverse and unknown scenarios.
[0167] In one embodiment, it can be achieved by Figure 10 The method shown determines the mixed sampling ratio, such as Figure 10 As shown, the process includes the following steps: S1001, calculate the difference between 1 and the ratio, i.e. .
[0168] S1002, the larger of the minimum retention ratio and the difference is determined as the mixed sampling ratio of the first training sample set and the second training sample set.
[0169] In S1002, the mixed sampling ratio is the proportion of samples drawn from the first training sample set and the second training sample set in a single training iteration. The proportion of the first training sample set is the result determined in this step, and the proportion of the second training sample set is 1 minus this result.
[0170] During training, from the heuristic experience pool and explore the experience pool Sampling is performed proportionally across the two experience pools. Mixing ratio. With training iteration rounds Dynamic decay: in It is the number of mixed decay steps. This represents the minimum retention ratio of pre-filled samples, ensuring that high-quality samples are still occasionally encountered towards the end of training, preventing complete forgetting. n is the current training iteration round (or total number of steps), and the mixing ratio. Indicates the current round from Sampling ratio, number of mixed decay steps A "schedule" is defined, which specifies how many training iterations the agent needs to go through to shift the primary source of experience sampling from the heuristic experience pool. Basically switched to the exploration experience pool .
[0171] For example, if the difference is 0.2 and the minimum retention ratio is 0.1, then 0.2 is selected as the sampling ratio of the first training sample set; if the difference is 0.05 and the minimum retention ratio is 0.1, then 0.1 is selected as the sampling ratio of the first training sample set, ensuring that the sampling ratio of the first training sample set is not lower than the preset lower limit.
[0172] By calculating the difference between "1 - number of iterations / preset number of iterations" and combining it with the minimum retention ratio, the sampling ratio of the first training sample set is gradually reduced to avoid model training oscillations caused by sudden changes in the sampling ratio. The minimum retention ratio (p_min) ensures that a certain proportion of expert experience samples are retained throughout the training process, preventing the agent from completely deviating from reasonable scheduling rules and ensuring the bottom-line safety of the strategy. Furthermore, the dynamically adjusted mixed sampling ratio makes full use of the guiding value of the initial expert samples and fully leverages the exploratory value of the simulation samples in the later stages, so that the sample resources at different stages are optimally utilized.
[0173] S1004, Collect target training samples corresponding to the mixed sampling ratio from the first training sample set and the second training sample set.
[0174] In S1004, the target training samples are the set of samples used for model training extracted from the two training sample sets according to the mixed sampling ratio in this iteration, which need to cover the corresponding ratio of experience samples and exploration samples.
[0175] In practice, the number of samples to be drawn from the first and second training sample sets can be calculated based on the mixed sampling ratio. A corresponding number of high-quality experience samples are randomly drawn from the first training sample set, and a corresponding number of exploration samples are randomly drawn from the second training sample set. The two types of samples are combined to form the target training samples for this iteration, providing data support for the parameter adjustment of the agent.
[0176] The first training sample set (expert experience samples) provides high-quality initial learning materials for the agent, enabling it to quickly master basic scheduling rules and avoid inefficiencies caused by random exploration in the early stages of training. By dynamically adjusting the mixed sampling ratio, the initial stage mainly uses expert experience samples (for quick learning) and the later stage mainly uses simulation samples (to overcome the limitations of experience), so that the agent has both basic scheduling capabilities and can learn better strategies for complex scenarios. At the same time, by combining the "determinism" of expert experience with the "uncertainty" of simulation scenarios, the training sample set covers both normal and extreme scenarios (such as high-concurrency battery swapping), improving the agent's coping ability.
[0177] In one embodiment, it can be achieved by Figure 11 The second training sample set is constructed as shown, such as Figure 11 As shown, the process includes the following steps: S1101, construct a simulation environment for a closed operation sample scenario, and initialize the environmental state simulation data in the simulation environment.
[0178] For example, environmental status data includes vehicle data and battery swapping station data in closed operation scenarios.
[0179] In S1101, the simulation environment of the closed operation sample scenario can be a virtual environment that accurately simulates the processes of vehicle operation, power consumption, battery swapping queuing, and dynamic task assignment in closed scenarios such as ports and mines, with discrete event-driven operation as the core.
[0180] In practice, the core parameters of the simulation environment, including the total number of vehicles, can be set according to the scale and operational characteristics of the real closed operation scenario. Number of battery swapping stations (Taking 3 as an example), the range of operation time distribution, the initial remaining power distribution interval, etc. Construct a discrete event simulation framework to clarify the core logic such as vehicle task generation, power consumption calculation, and battery swapping station service scheduling.
[0181] Initialize vehicle data: Assign an initial position to each simulated vehicle, randomly generate initial remaining battery power and tasks that have started but not yet been completed, and determine the task completion time.
[0182] For example, all battery swapping stations The "next time when battery swapping is available" Initialized to 0, indicating an initial idle state. For each vehicle... Randomly generate initial power Given an initial position, generate an initial job task that has started but not yet completed: randomly generate job length. (Following a specified distribution), and a start time is randomly generated. ,satisfy and This allows us to calculate the initial task end time (i.e., the offline time). .
[0183] Initialize battery swapping station data: Set the next idle time of all battery swapping stations as the initial time, define fixed parameters such as the duration of a single battery swap, and complete the initialization of environmental state simulation data.
[0184] For example, the simulation main loop starts from the current time. Beginning, until (generally The time (i.e., 24 hours) ends.
[0185] S1102, when the simulation clock advances to the first simulation moment when the first simulation vehicle completes its task, the first simulation multidimensional state vector of the simulation environment at the first simulation moment is obtained from the environmental state simulation data; the first simulation vehicle is any one of the multiple simulation vehicles.
[0186] In S1102, the simulation clock is a timing tool used to record the running time of the simulation environment. Its scale is synchronized with real time to ensure the temporal continuity of the simulation process. The first simulation vehicle is any vehicle among multiple simulation vehicles that is the first to complete its current task and trigger the battery swapping decision event. The first simulation moment is the specific simulation time point at which the first simulation vehicle completes its current task and is ready to accept new scheduling instructions. The first simulation multidimensional state vector is a standardized data set characterizing the overall state of the simulation environment at the first simulation moment. Its structure is consistent with the multidimensional state vector of the real scene and is used for agent input.
[0187] In practice, after the simulation clock is started, the simulation environment runs according to preset logic, monitoring the work progress of all simulated vehicles in real time. When a simulated vehicle completes its current task, that moment is recorded as the first simulation moment, and that vehicle is designated as the first simulated vehicle.
[0188] For example, the simulator finds the minimum task completion time among all vehicles. And advance the simulation clock to that moment, i.e. This means that at least one vehicle has completed its mission and is ready to receive new dispatch instructions.
[0189] Key information for the first simulation moment is extracted from the environmental state simulation data: the remaining battery power of the first simulation vehicle, the travel time to each battery swapping station, the next idle time of each battery swapping station, the statistical distribution of the remaining battery power of the entire fleet, and the prediction of battery swapping demand within a preset future time period. After normalizing and logarithmizing this information, it is combined in a fixed order to form the first simulation multidimensional state vector.
[0190] S1103, the first simulated multidimensional state vector is input into the battery swapping action prediction model in the agent to be trained, and the simulation reward function values corresponding to different candidate battery swapping actions of the first simulated vehicle are obtained.
[0191] In practice, the first simulated multidimensional state vector is input into the agent to be trained. The agent extracts and parses the features of the state vector through the network layer. Based on the type of battery swapping action (battery swapping or no battery swapping) in the simulation environment, a preset reward function calculation rule is invoked: for battery swapping, the reward function is calculated by combining the remaining battery power of the first simulated vehicle, the total offline time, the preset average offline time, and the corresponding battery power reward coefficient; for no battery swapping, the reward function is calculated based on the remaining battery power and the no-battery-swap reward coefficient. The calculated initial reward values are scaled and smoothed to obtain the simulation reward function values corresponding to each candidate battery swapping action.
[0192] S1104, through the decision-making module in the intelligent agent, the battery swapping action of the first simulated vehicle is decided based on multiple simulation reward function values, so as to obtain the battery swapping prediction action of the first simulated vehicle.
[0193] In practice, an action masking mechanism can be invoked to compare the remaining battery power of the first simulated vehicle with a preset threshold: if the remaining battery power is below the safety threshold, the "no battery swap" action is blocked; if the remaining battery power is above the prohibition threshold, all battery swap actions are blocked, and legal candidate battery swap actions are selected. The simulation reward function values corresponding to the legal actions are compared, and the action with the largest value is selected as the predicted battery swap action for the first simulated vehicle, determining whether the vehicle will subsequently perform a battery swap or continue operation.
[0194] For example, all in Vehicles whose missions are about to end are selected in sequence and used as the current decision-making vehicle. For each decision-making vehicle, the intelligent agent needs to consider its action space. Choose one action ,in This indicates "no battery replacement". Indicates "to go to the "Electricity swapping at individual battery swapping stations."
[0195] S1105, after the first simulated vehicle completes the battery swap prediction action, update the state variables in the simulation environment to obtain the second simulation multidimensional state vector.
[0196] In practice, if the predicted action for battery swapping is no battery swapping, a new task is generated for the first simulated vehicle, the task duration is randomly determined, and the task end time is updated.
[0197] For example, if the action (Without battery swapping): This refers to the vehicle. Generate a new job task with a job length of The new task is randomly generated from a given distribution. The start time of the new task is the current time. The end time has been updated to .
[0198] If the predicted action for battery swapping is to go to the designated battery swapping station, calculate the time it takes for the vehicle to travel to the target battery swapping station, the estimated arrival time, the start and end times of battery swapping, update the next idle time of the battery swapping station, reset the vehicle's remaining battery power to full and update its operation end time.
[0199] For example, if the action (Heading to the battery swapping station) ), calculate vehicles Drive to the battery swapping station Time required Calculate vehicles Estimated arrival time Calculate the battery swap start time ,in It is a battery swapping station The next available time will be the time to provide battery swapping services, and the battery swapping end time will be calculated. ,in Minutes is a fixed battery swapping time; the battery swapping station is being updated. Status: Record vehicle Queue time Updating vehicles The task end time is and its After the battery swap is complete, the battery will be reset to full charge.
[0200] Then, the time difference from the first simulation moment to the current moment is calculated. Based on the energy consumption rate of each online simulation vehicle, the remaining power of all vehicles is updated, and the distribution of remaining power of the entire fleet and the prediction of future battery swapping needs are recalculated.
[0201] For example, from the previous event point (the first simulation time) to the current time... Time progressed According to each vehicle energy consumption rate Update the battery levels of all vehicles operating online: .
[0202] Finally, based on the structure and processing rules of the first simulation multidimensional state vector, all the updated data are integrated to form the second simulation multidimensional state vector.
[0203] S1106, based on the first simulated multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second simulated multidimensional state vector, a second training sample set is constructed for training the agent.
[0204] In practice, the four types of data are structurally combined according to a fixed order: "first simulation multidimensional state vector - target battery swapping action - simulation reward function value - second simulation multidimensional state vector" to form a single simulation training sample. This sample is then stored in the second training sample set. The simulation environment operation, state acquisition, action decision-making, and state update process are repeated to continuously generate diverse simulation samples, enriching the second training sample set and providing sufficient data support for the iterative training of the agent.
[0205] By constructing a high-fidelity simulation environment, the battery swapping scheduling process under a large number of different scenarios (number of vehicles, task distribution, battery swapping demand) can be quickly simulated, and a massive amount of training samples can be efficiently generated, solving the problems of long sample collection cycles and high costs in real scenarios. On the one hand, the simulation environment can flexibly adjust parameters (such as the number of vehicles, battery swapping station capacity, and task frequency) to generate diverse samples such as normal scenarios, high-concurrency scenarios, and extreme power scenarios, enabling the agent to learn comprehensive scheduling rules. On the other hand, the simulation environment can safely test the effects of different battery swapping actions, avoiding operational losses caused by incorrect scheduling in real scenarios. At the same time, through iterative training of simulation samples, the performance of the agent can be optimized in advance, improving the reliability after actual deployment.
[0206] In this embodiment, by responding to vehicle operation completion requests in real time, vehicle data and battery swapping station data of all operating vehicles in the closed operation scenario are dynamically collected. A multi-dimensional state vector is constructed and input into the battery swapping action prediction model in the trained agent to obtain multiple reward function values. Based on these reward function values, the decision module in the agent selects target battery swapping actions, ensuring that the decision results take into account both individual vehicle needs and global optimization of battery swapping resources. This achieves real-time and adaptive scheduling of battery swapping demand. This method overcomes the shortcomings of traditional rule-based scheduling, which is rigid and has poor real-time performance in operations optimization. It can make intelligent scheduling decisions based on the global state in dynamic and uncertain closed operation scenarios (such as ports and mines), thereby reducing queuing time and improving vehicle attendance and operational continuity.
[0207] To better understand the implementation of the above battery swapping scheduling method, the complete implementation of the above battery swapping scheduling method is given below in combination with a specific scenario.
[0208] The application scenario is a large-scale automated container terminal. This terminal deploys 150 electric heavy-duty trucks for horizontal container transport and has three battery swapping stations, each equipped with two swapping bays. The system operates continuously 24 hours a day. During the initialization phase, the simulation environment is configured with the following parameters based on historical operational data: the initial battery level of the vehicles follows a uniform distribution, ranging from 30% to 80%; the task duration for each vehicle follows an exponential distribution with a mean of 45 minutes; the vehicle's speed within the terminal is fixed, and the travel time between stations is between 3 and 8 minutes; the duration of a single battery swap operation is 5 minutes.
[0209] The initial idle time for all battery swapping stations is set to 0 minutes. Each vehicle is randomly assigned an initial task, the start time of which is set before the simulation begins, ensuring that the vehicle is in operation when the simulation starts.
[0210] Suppose that at the 125th minute of the simulation, the electric heavy truck with the serial number V057 completes a container transfer task and reports the "task completed" status to the central dispatch system through the on-board terminal.
[0211] The central dispatch system responds to the reporting request from the first vehicle in the closed operation scenario upon completion of its task, and acquires vehicle data for all vehicles in the closed operation scenario and battery swapping station data for all battery swapping stations. In other words, upon receiving the reporting request, the dispatch system immediately responds and collects global real-time data at that moment. Vehicle data: Obtain the real-time SOC, current location, and current task status (whether it is in operation and the estimated completion time) of all 150 vehicles.
[0212] Battery swapping station data: Obtain the next available service time, the number of vehicles currently in the queue, and the status of each workstation for each of the three battery swapping stations.
[0213] Subsequently, based on vehicle data and battery swapping station data, a first multi-dimensional state vector for the closed operation scenario at the first moment is constructed. That is, based on the above data, a 17-dimensional state vector is constructed for the decision-making vehicle V057, the details of which are as follows: V057's SOC value: its current battery level is 42%, normalized to 0.42.
[0214] The travel time of V057 to each battery swapping station was calculated to be 5 minutes, 7 minutes, and 4 minutes respectively. After normalization by the global maximum and minimum times (3 minutes and 8 minutes), three characteristic values were obtained.
[0215] The next available time for each battery swapping station is as follows: Station 1 will be available in the 128th minute, Station 2 in the 135th minute, and Station 3 in the 131st minute. Calculate the difference between this time and the current time of 125 minutes, and normalize it after logarithmic transformation.
[0216] Full fleet SOC distribution statistics: This involves counting the number of vehicles in the entire fleet (150 vehicles) across five battery levels (1-20%, 21-40%, 41-60%, 61-80%, 81-100%) and calculating the percentages. For example, currently, 35% of the vehicles are in the 41-60% range.
[0217] Short-term battery swapping demand forecast: Count the number of vehicles that will complete the task within the next 10 minutes (125 to 135 minutes), and calculate the proportion according to the five battery capacity ranges mentioned above.
[0218] Next, the first multi-dimensional state vector is input into the pre-trained battery swapping action prediction model of the agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. Specifically, the constructed state vector can be input into the trained D3QN agent model to obtain different candidate battery swapping actions of the first vehicle. When a candidate battery swapping action is selected as the vehicle to proceed with the swapping, it is determined whether the first state of charge (SOC) value is less than a preset first SOC. If the first SOC value is less than the preset first SOC, the reward function is determined based on the first SOC value and the total offline time. Preset average offline time Time reward coefficient First Battery Reward Coefficient Determine the reward function value of the candidate battery swapping action. . .
[0219] When the first state of charge (SOC) value is greater than or equal to a preset second SOC, based on the first SOC value and the total offline time... Preset average offline time Time reward coefficient Second power reward coefficient Determine the reward function value of the candidate battery swapping action. . .
[0220] When the first state of charge (SOC) value is between a preset first SOC and a preset second SOC, based on the first SOC value and the total offline time... Preset average offline time Time reward coefficient Third electricity reward coefficient Determine the reward function value of the candidate battery swapping action. . .
[0221] If the candidate for battery swapping is selected as not swapping, the reward coefficient for not swapping will be applied. And the first state of charge (SOC) value, to determine the reward function value for not swapping batteries. .
[0222] To prevent the reward value from being too large and affecting training stability, the reward function value can be scaled and smoothed to obtain a processed reward function value.
[0223] For example, the model outputs the Q-value (long-term expected return) for each optional action in V057: Action 0 (no battery replacement): Q value = 1.23 Action 1 (Go to Station 1): Q value = 2.15 Action 2 (Go to Station 2): Q value = 1.87 Action 3 (Go to Station 3): Q value = 2.41 Subsequently, using an action masking mechanism, legitimate candidate battery swapping actions are selected from multiple candidate actions. From these legitimate candidate actions, the battery swapping action with the highest reward function value is selected to obtain the target battery swapping action for the first vehicle. If the first State of Charge (SOC) value is less than a first SOC threshold, non-battery swapping actions are masked from multiple candidate actions, resulting in the battery swapping action to proceed with. If the first SOC value is greater than a second SOC threshold, battery swapping actions are masked from multiple candidate actions, resulting in the non-battery swapping action.
[0224] Since V057's SOC is 42%, which is 20% higher than the set safety threshold (second SOC threshold), the "no battery swap" action is legal. Simultaneously, its battery level is below the prohibition threshold for battery swapping (first SOC threshold) by 80%, so the "go to battery swapping station" action is also legal. None of the actions were blocked. Ultimately, the decision logic selects action 3, which has the highest Q value, instructing V057 to go to battery swapping station number 3 for battery swapping.
[0225] The system issues a dispatch command to the V057 vehicle's onboard terminal: "Proceed to Battery Swap Station No. 3 for battery swapping." Simultaneously, the system internally updates the simulation status. Calculate the travel time of V057 (4 minutes), and the estimated arrival time at station 3 is 129 minutes.
[0226] Since the next available time at station 3 is in the 131st minute, V057 will need to queue for 2 minutes. The battery swap will begin in the 131st minute and end in the 136th minute.
[0227] The next free time for station #3 is now set for the 136th minute.
[0228] The queuing time for V057 was recorded as 2 minutes, and the total offline time was 11 minutes (4 minutes of driving + 2 minutes of queuing + 5 minutes of battery swapping).
[0229] Update the mission end time of V057 to the 136th minute and schedule its SOC to be reset to 100% at that time.
[0230] While V057 was en route to the battery swapping station, the system continued to advance the simulation clock, process task completion events for other vehicles, and make a new round of scheduling decisions based on the latest global state.
[0231] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0232] In one embodiment, such as Figure 12 As shown, this application provides a battery swapping dispatching device, the battery swapping dispatching device 1200 including: The acquisition module 1201 is used to respond to the reporting request of the first vehicle in the closed operation scenario to complete the operation task at the first moment, and to acquire vehicle data of all operation vehicles and battery swapping station data of all battery swapping stations in the closed operation scenario. Module 1202 is used to construct the first multi-dimensional state vector of the closed operation scenario at the first moment based on vehicle data and battery swapping station data. The input module 1203 is used to input the first multidimensional state vector into the pre-trained agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The decision module 1204 is used to make a decision on the battery swapping action of the first vehicle based on multiple reward function values, and to obtain the target battery swapping action of the first vehicle. The scheduling module 1205 is used to schedule the battery swapping of the first vehicle based on the target battery swapping action.
[0233] In one embodiment, such as Figure 13As shown, this application provides an electronic device, which includes: a processor 1301 and a memory 1302 storing computer program instructions; When the processor 1301 executes computer program instructions, it implements the battery swapping scheduling method described above.
[0234] In one embodiment, this application provides a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the battery swapping scheduling method described above.
[0235] In one embodiment, this application provides a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the battery swapping scheduling method described above.
[0236] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0237] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0238] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0239] The foregoing flowcharts and / or block diagrams describing the method for determining the open-circuit voltage of a battery, the battery management system, and the power-consuming device according to embodiments of this application have described various aspects of this application. It should be understood that each block in the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to create a machine such that these instructions, executable by the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowcharts and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0240] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and they should all be covered within the scope of the claims and specification of this application. In particular, as long as there is no structural conflict, the various technical features mentioned in the embodiments can be combined in any way. This application is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. A battery swapping dispatch method, characterized in that, include: In response to the first vehicle in the closed operation scenario reporting the completion of the operation task at the first moment, the vehicle data of all operation vehicles and the battery swapping station data of all battery swapping stations in the closed operation scenario are obtained. Based on the vehicle data and the battery swapping station data, a first multi-dimensional state vector of the closed operation scenario at the first moment is constructed. The first multidimensional state vector is input into the battery swapping action prediction model in the pre-trained agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The decision-making module in the intelligent agent makes a decision on the battery swapping action of the first vehicle based on multiple reward function values, and obtains the target battery swapping action of the first vehicle. Based on the target battery swapping action, battery swapping scheduling is performed on the first vehicle; The first multidimensional state vector includes the multidimensional state vector of the first vehicle and the global state vector in the closed operation scenario. The multidimensional state vector of the first vehicle includes: the first state of charge (SOC) value of the first vehicle and the time required for the first vehicle to travel to each of the battery swapping stations; The global state vector includes: the next idle time of each battery swapping station, the duration of a single battery swap, the SOC distribution statistics of all operating vehicles, and the predicted demand information within a preset time period; The step of making a decision on the battery swapping action of the first vehicle based on multiple reward function values to obtain the target battery swapping action of the first vehicle includes: When the first state of charge (SOC) value is less than the first SOC threshold, the non-swap action is filtered out from the multiple candidate swapping actions to obtain the swapping action that does not swap. If the first state of charge (SOC) value is greater than the second SOC threshold, the battery swapping action is filtered out from the multiple candidate battery swapping actions to obtain the non-battery swapping action. The battery swapping action with the largest reward function value is selected from the legal candidate battery swapping actions to obtain the target battery swapping action for the first vehicle. The first SOC threshold is a safety critical power level set based on vehicle energy consumption rate, operation scene distance, and battery swapping station distribution. If the power level is below the first SOC threshold, it indicates that the vehicle's power is insufficient to support the next round of operation and battery swapping is necessary. The second SOC threshold is a battery swapping limit set to avoid resource waste. If the power level is above the second SOC threshold, it indicates that the vehicle's power is sufficient and multiple rounds of operation can be completed without battery swapping.
2. The method according to claim 1, characterized in that, The step of inputting the first multidimensional state vector into the pre-trained battery swapping action prediction model of the agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle includes: The first multidimensional state vector is input into the battery swapping action prediction model in the pre-trained agent to obtain different candidate battery swapping actions of the first vehicle. In the case where the candidate battery swapping option is selected as the battery swapping option, for each candidate battery swapping action: Based on the first moment, the time required for the first vehicle to travel to the corresponding battery swapping station, and the battery swapping time at the battery swapping station, the total offline time of the first vehicle from the first moment to the completion of the battery swapping is determined when the first vehicle is swapping its battery at the battery swapping station. The reward function value of the candidate battery swapping action is determined based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and energy reward coefficient.
3. The method according to claim 2, characterized in that, The power reward coefficient includes a first power reward coefficient, a second power reward coefficient, and a third power reward coefficient. The first power reward coefficient is greater than the second power reward coefficient, and the third power reward coefficient is between the first power reward coefficient and the second power reward coefficient, and decreases linearly with the increase of the first state of charge (SOC) value. The step of determining the reward function value of the candidate battery swapping action based on the first state of charge (SOC) value, total offline time, preset average offline time, time reward coefficient, and energy reward coefficient includes: When the first state of charge (SOC) value is less than the preset first SOC, the reward function value of the candidate battery swapping action is determined based on the first SOC value, the total offline time, the preset average offline time, the time reward coefficient, and the first energy reward coefficient. When the first state of charge (SOC) value is greater than or equal to a preset second SOC, the reward function value of the candidate battery swapping action is determined based on the first SOC value, the total offline time, the preset average offline time, the time reward coefficient, and the second energy reward coefficient. When the first state of charge (SOC) value is greater than or equal to a preset first SOC and less than a preset second SOC, the reward function value of the candidate battery swapping action is determined based on the first SOC value, the total offline time, the preset average offline time, the time reward coefficient, and the third energy reward coefficient.
4. The method according to claim 2, characterized in that, The step of inputting the first multidimensional state vector into a pre-trained intelligent agent's battery swapping action prediction model to obtain reward function values corresponding to different candidate battery swapping actions of the first vehicle further includes: If the candidate for battery swapping is not selected, the reward function value for the non-battery swapping is determined based on the first State of Charge (SOC) value.
5. The method according to claim 3 or 4, characterized in that, The step of inputting the first multidimensional state vector into a pre-trained intelligent agent's battery swapping action prediction model to obtain reward function values corresponding to different candidate battery swapping actions of the first vehicle further includes: The reward function value is scaled and smoothed to obtain the processed reward function value.
6. The method according to claim 1, characterized in that, The method further includes: After the first vehicle completes its battery swap, the state vector of the closed operation scenario at the second time point, the second multidimensional state vector, is updated; the second time point is the moment when the first vehicle completes its battery swap. Based on the first multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second multidimensional state vector, training samples for training the agent are constructed.
7. The method according to claim 1, characterized in that, The method further includes: The target training sample is obtained by sampling from the training sample set; each training sample in the training sample set includes: the first multidimensional state vector sample, the battery swapping action sample, the reward function value sample, and the second multidimensional state vector sample corresponding to the closed operation scenario sample under multiple different first time samples, wherein the second time sample is the time corresponding to the closed operation scenario sample when the battery swapping action sample is completed. The first multidimensional state vector sample is input into the agent to be trained to obtain the battery swapping prediction action. Based on the battery swapping prediction action and the first multidimensional state vector sample, determine the predicted value of the first reward function; Based on the predicted value of the first reward function and the sample reward function value, the loss function value is determined; If the iteration stopping condition is not met, the model parameters of the agent to be trained are adjusted based on the loss function value. The training sample set is resampled to obtain a newly sampled target training sample; and the first multidimensional state vector sample is input into the agent to be trained until the iteration stopping condition is met.
8. The method according to claim 7, characterized in that, The training sample set includes a first training sample set and a second training sample set. The training samples in the first training sample set are samples constructed based on expert experience, while the training samples in the second training sample set are samples generated during simulated battery swapping in a simulated closed operation scenario. The step of resampling from the training sample set to obtain newly sampled target training samples includes: Count the current iteration number; Determine the ratio of the current iteration number to the preset iteration number; Based on the ratio and the minimum retention ratio of the first training sample set, the mixed sampling ratio of the first training sample set and the second training sample set is determined; Target training samples corresponding to the mixed sampling ratio are collected from the first training sample set and the second training sample set.
9. The method according to claim 8, characterized in that, Determining the mixed sampling ratio of the first and second training sample sets based on the ratio and the minimum retention ratio of the first training sample set includes: Calculate the difference between 1 and the ratio; The larger of the minimum retention ratio and the difference is determined as the mixed sampling ratio of the first training sample set and the second training sample set.
10. The method according to claim 8 or 9, characterized in that, The process of constructing the second training sample set includes: A simulation environment for a closed-loop operation sample scenario is constructed, and the environmental state simulation data in the simulation environment is initialized; the environmental state simulation data includes vehicle data and battery swapping station data in the closed-loop operation scenario; When the simulation clock advances to the first simulation moment when the first simulation vehicle completes its task, the first simulation multidimensional state vector of the simulation environment at the first simulation moment is obtained from the environmental state simulation data; the first simulation vehicle is any one of multiple simulation vehicles; The first simulated multidimensional state vector is input into the battery swapping action prediction model in the agent to be trained to obtain the simulation reward function values corresponding to different candidate battery swapping actions of the first simulated vehicle. The decision-making module in the intelligent agent makes a decision on the battery swapping action of the first simulated vehicle based on multiple simulation reward function values, and obtains the predicted battery swapping action of the first simulated vehicle. After the first simulated vehicle completes the battery swapping prediction action, the state variables in the simulation environment are updated to obtain the second simulation multidimensional state vector. Based on the first simulated multidimensional state vector, the target battery swapping action, the reward function value corresponding to the target battery swapping action, and the second simulated multidimensional state vector, a second training sample set is constructed for training the agent.
11. A battery swapping dispatching device, characterized in that, include: The acquisition module is used to respond to the reporting request from the first vehicle in the closed operation scenario at the first moment to complete the operation task, and to acquire vehicle data of all operation vehicles and battery swapping station data of all battery swapping stations in the closed operation scenario. The construction module is used to construct the first multi-dimensional state vector of the closed operation scenario at the first moment based on the vehicle data and the battery swapping station data; The input module is used to input the first multi-dimensional state vector into the battery swapping action prediction model in the pre-trained intelligent agent to obtain the reward function values corresponding to different candidate battery swapping actions of the first vehicle. The decision module is used to make a decision on the battery swapping action of the first vehicle based on multiple reward function values through the decision module in the intelligent agent, so as to obtain the target battery swapping action of the first vehicle. The scheduling module is used to schedule the battery swapping of the first vehicle based on the target battery swapping action. The first multidimensional state vector includes the multidimensional state vector of the first vehicle and the global state vector in the closed operation scenario. The multidimensional state vector of the first vehicle includes: the first state of charge (SOC) value of the first vehicle and the time required for the first vehicle to travel to each of the battery swapping stations; The global state vector includes: the next idle time of each battery swapping station, the duration of a single battery swap, the SOC distribution statistics of all operating vehicles, and the predicted demand information within a preset time period; The decision-making module is specifically used for: When the first state of charge (SOC) value is less than the first SOC threshold, the non-swap action is filtered out from the multiple candidate swapping actions to obtain the swapping action that does not swap. If the first state of charge (SOC) value is greater than the second SOC threshold, the battery swapping action is filtered out from the multiple candidate battery swapping actions to obtain the non-battery swapping action. The battery swapping action with the largest reward function value is selected from the legal candidate battery swapping actions to obtain the target battery swapping action for the first vehicle. The first SOC threshold is a safety critical power level set based on vehicle energy consumption rate, operation scene distance, and battery swapping station distribution. If the power level is below the first SOC threshold, it indicates that the vehicle's power is insufficient to support the next round of operation and battery swapping is necessary. The second SOC threshold is a battery swapping limit set to avoid resource waste. If the power level is above the second SOC threshold, it indicates that the vehicle's power is sufficient and multiple rounds of operation can be completed without battery swapping.
12. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, it implements the battery swapping scheduling method as described in any one of claims 1-10.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the battery swapping scheduling method as described in any one of claims 1-10.
14. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the battery swapping scheduling method as described in any one of claims 1-10.
Citation Information
Patent Citations
Electric vehicle scheduling method, device and equipment and storage medium
CN120875282A
Battery replacement decision-making method for electric unmanned mine car
CN121073168A