Charging station operation and maintenance path planning method combining clustering station chain decomposition and reinforcement learning

By dividing the charging piles into multiple operation and maintenance areas and using reinforcement learning algorithms to generate the optimal operation and maintenance path, the problem of low efficiency in traditional operation and maintenance path planning is solved, rapid response and efficient fault handling are achieved, and the efficiency and economic benefits of the operation and maintenance of the charging piles are significantly improved.

CN120218375APending Publication Date: 2025-06-27NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510306609.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

Smart Images

  • Figure CN120218375A_ABST
    Figure CN120218375A_ABST
Patent Text Reader

Abstract

The invention provides a charging station operation and maintenance path planning method combining clustering station chain decomposition and reinforcement learning. The method mainly comprises the following steps that longitudes and latitudes of all target charging stations are collected, all stations are divided into a plurality of smaller operation and maintenance areas with the number of operation and maintenance personnel of the stations in the current area as the limit, and the Euclidean distance from the center of each operation and maintenance area to the charging stations in the area is the shortest; then, a reinforcement learning Q-learning algorithm is adopted to carry out path planning on sites in the operation and maintenance area, path working time after planning is calculated and iteration is carried out until the proportion of the operation and maintenance area with all path working time smaller than the working time limit is larger than a limit threshold value; and outputting path planning routes and route working time of all the target operation and maintenance sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A method for planning the operation and maintenance path of charging stations by combining clustering site chain decomposition and reinforcement learning belongs to the technical field of path planning. Technical Background

[0002] With the rapid popularization of electric vehicles, charging piles, as an important part of electric vehicle infrastructure, their operation and maintenance efficiency directly affects the charging experience of users and the operating costs of operators. However, charging piles are widely distributed, have various fault types, and limited operation and maintenance resources. Traditional operation and maintenance scheduling methods often rely on manual experience, making it difficult to achieve efficient path planning and resource allocation, resulting in too long response times for operation and maintenance personnel and untimely fault handling, seriously affecting the availability of charging piles and user satisfaction.

[0003] To solve this problem, this patent proposes a method for planning the operation and maintenance path of charging stations by combining clustering site chain decomposition and reinforcement learning. This method divides charging piles into multiple operation and maintenance areas and combines the Q-learning algorithm to dynamically generate the optimal operation and maintenance path for each area. This algorithm can comprehensively consider multi-dimensional factors such as the geographical distribution of charging piles, fault priorities, and operation and maintenance resource distributions, automatically learn and optimize the operation and maintenance path, thereby significantly shortening the response time of operation and maintenance personnel and improving the timeliness and efficiency of fault handling. At the same time, this algorithm can dynamically adjust the path strategy according to real-time operation and maintenance requirements, more reasonably allocate operation and maintenance resources, avoid resource waste and repeated scheduling, and significantly improve the overall efficiency and economic benefits of charging pile operation and maintenance.

[0004] The technical solution of this patent is not only applicable to the operation and maintenance scenario of charging piles, but also can be extended to the operation and maintenance fields of other similar distributed facilities, with broad application prospects and important commercial value. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to solve the defect that the selection of the current operation and maintenance route depends on manual experience judgment. By using a clustering algorithm, the operation and maintenance sites are divided into several operation and maintenance areas, and then the reinforcement learning algorithm is used to iteratively find an optimal path for maintaining each site in the operation and maintenance area, so as to achieve efficient optimization of the maintenance efficiency of all operation and maintenance sites. The present invention not only solves the problem that traditional path optimization algorithms have low optimization efficiency and are not easy to converge when facing a large number of path points, but also provides a practical basis for the operation and maintenance department to increase or decrease operation and maintenance staff and improve the company's efficiency.

[0006] The technical solution adopted by the present invention to solve its technical problems is: a method for planning the operation and maintenance path of charging stations by combining clustering site chain decomposition and reinforcement learning, including the following steps:

[0007] Step 1: Collect and obtain the longitude and latitude information of all target charging stations, process it and mark it on the map;

[0008] Step 2: Initialize the number of regional operation and maintenance centers, and divide them into several operation and maintenance regions based on K-means clustering, so that the Euclidean distance from the regional operation and maintenance center to the stations within the region is the shortest;

[0009] Step 3: Use the reinforcement learning Q-learning algorithm for the stations within the region to find the shortest operation and maintenance route;

[0010] Step 4: Calculate the maintenance time corresponding to the shortest operation and maintenance route in each regional operation and maintenance center. If it exceeds the restricted working time, re-divide the regional operation and maintenance center based on Step 2;

[0011] Step 5: Output the operation and maintenance path planning lines and the total cost required for all charging stations.

[0012] Preferably, in Step 1, the longitude and latitude information of all target charging stations is collected through the Baidu Map Open Platform. After being processed, these information are marked on the map. The specific implementation process is as follows:

[0013] Step 1-1: Obtain the longitude and latitude information of all target charging stations by calling the Baidu Map API interface;

[0014] Step 1-2: Go to the Baidu Map Open Platform to register and apply to become an individual developer, and then apply for an APIKey;

[0015] Step 1-3: For obtaining the longitude and latitude information of charging stations, the "geocoding" API interface needs to be used;

[0016] Step 1-4: After the Baidu Map server calls the interface, it will return a JSON-formatted response data, and the information of each charging station will contain its longitude and latitude coordinates;

[0017] Step 1-5: Extract the longitude and latitude information of all target charging stations and store it in a suitable data structure for subsequent map marking and operation and maintenance path planning;

[0018] Step 1-6: Mark the corresponding coordinates on the map to obtain a visual representation of the location information of all charging stations.

[0019] Preferably, in Step 2, the initial number of regional operation and maintenance centers is set according to the operation conditions of the operation and maintenance centers, and the K-means clustering algorithm is used to divide these stations into several operation and maintenance regions to ensure that the Euclidean distance from the stations within each region to their corresponding operation and maintenance centers is the shortest. The specific implementation process is as follows:

[0020] Step 2-1: Initialize by randomly selecting K points from the site set as the initial regional centers, where K is the number of operation and maintenance personnel arranged initially. These regional centers can be actual points in the site set or randomly generated points;

[0021] Step 2-2: Assign regions to each site, that is, calculate the Euclidean distance from each data point to all centroids. For data point x i and the first centroid μ Ⅰ , the Euclidean distance calculation formula is shown in the following figure:

[0022]

[0023] where M refers to the level of data for which the clustering algorithm needs to be used. In two-dimensional plane data, it is divided into longitude and latitude. d(x i , μ Ⅰ ) refers to the distance between data point x i and the first centroid μ Ⅰ . It represents the square root of the sum of the squared differences between the i-th data point x im and the first centroid μ Ⅰm at the current level m;

[0024] Assign each site to the region numbered j corresponding to the regional center closest to it, that is:

[0025]

[0026] where ||x i -μ j || and ||x i -μ Ⅰ || respectively refer to the Euclidean distances from data point x i to the centroid μ j and the centroid μ Ⅰ . μ j refers to the centroid numbered j in the centroid set μ k from 1 to k, that is, when the distance from data point x i to other centroids μ j is less than the distance from data point x i to the original centroid μ Ⅰ , it is assigned to the j-th new region C j ;

[0027]

[0028] Step 2-3: Update the regional center for each region C j Calculate the average of the longitudes and latitudes of all sites within the region and use this average as the new regional center. The formula for the new regional center is:

[0029]

[0030] Among them, C j refers to the category numbered j, and |C j | refers to the number of data points in category C j . Calculate the average value of all data points in region j as the new centroid μ of category j j ;

[0031] Step 2-4: Repeat the steps of allocating regions and updating the region centers until the positions of the region centers no longer change or the maximum number of iterations is reached.

[0032] Preferably, in step 3, the Q-learning algorithm of reinforcement learning is used for each site inside the divided operation and maintenance region to explore the optimal operation and maintenance route. To ensure the practical operation feasibility of the operation and maintenance work for the maintenance time corresponding to each shortest operation and maintenance route, the specific implementation process is as follows:

[0033] Step 3-1: Configure the Q-learning algorithm environment and initialize and define the following concepts respectively:

[0034] State definition: Define the current site as the state, such as site S1, S2, etc. Each state represents the position where the operation and maintenance personnel may be currently;

[0035] Action set: The action set contains all possible operations that can move from the current site to adjacent sites. These actions constitute the movement options that the operation and maintenance personnel can take;

[0036] Reward function design: Among them, the movement reward is set according to the movement cost. For example, if the distance is 10 kilometers, the reward is -10 to reflect the economic cost of movement. Completion reward: When all sites are visited and the starting point (or the end point) is reached, a positive reward Reward is given to encourage the algorithm to find a complete operation and maintenance path. Illegal movement penalty: If an attempt is made to move to an illegal site (such as a non-existent path), a high penalty Penalty is given to avoid the algorithm generating invalid or incorrect movement strategies. The setting of the reward function is as follows:

[0037]

[0038] Among them, cost(s,a) represents the movement cost of executing action a from state s, and the termination condition is reaching the target site or exceeding the maximum number of steps (to avoid infinite loops);

[0039] Termination condition: Set reaching the target site or exceeding the maximum number of steps as the termination condition to ensure that the algorithm finds the optimal solution or reaches a stable state within a reasonable time;

[0040] Step 3-2: Initialize Q-learning training:

[0041] Create a Q-table according to the number of states and actions, with a size equal to the number of states multiplied by the number of actions, and initialize its values to 0. The Q-table is used to store the expected return values for each action in each state;

[0042]

[0043] Among them, Q(s,a) is the return value obtained by the agent when completing action a in state s;

[0044] The hyperparameters of the Q-table are set as the learning rate α, which is used to control the update speed of new information over old information, the discount factor γ, which represents the importance of future returns relative to current returns, and the exploration rate ε, which is used to balance between exploring new strategies and exploiting known optimal strategies;

[0045] Step 3-3: Execute the Q-learning training loop. At the start of each training episode, randomly select a starting site as the initial state s, then perform action selection. According to the exploration rate ε, randomly select an action for exploration, or select the action with the largest Q-value in the current state in the Q-table for greedy selection. Execute the selected action and move to the new site s';

[0046]

[0047] Among them, ε is the exploration rate set initially, argmax a' Q(s',a') represents the maximum Q-value among all possible actions a' in the next state s', which is an estimate of the maximum future return that can be obtained;

[0048] And calculate the reward r according to whether the move is legal, whether the path is completed, and whether an illegal move is attempted.

[0049] r = R(s,a,s')

[0050] (Equation 8)

[0051] Among them, R(s,a,s') is the value calculated according to the reward function R when the agent is in state s' after executing action a in state s, and this value is assigned to the parameter r and recorded in the Q-table;

[0052] Update the Q-value using the Bellman equation, that is:

[0053] Q(s,a) ← Q(s,a) + α[r + γ maxQ(s',a') - Q(s,a)]

[0054] (Equation 9)

[0055] Among them, Q(s,a) is the reward value obtained by the agent for completing action a in state s. α is the learning rate, a positive number between 0 and 1, which determines the speed at which new information overrides old information. r is the reward, the immediate reward obtained by the agent from the environment after executing an action. γ is the discount factor, which determines the current value of future rewards. maxQ(s',a') represents the maximum Q value among all possible actions a' in the next state s', and it is an estimate of the maximum reward that may be obtained in the future;

[0056] Step 3-4: Perform convergence judgment, that is, monitor the change of the Q-table. When the change of the Q value in consecutive multiple trainings is less than the set threshold, it is considered that the algorithm has converged to the optimal solution or reached a stable state;

[0057] Step 3-5: After the training is completed, set the exploration rate ε to 0, and only use the known optimal strategy for path selection. Starting from the starting point of the area, select the action with the maximum Q value at each step according to the trained Q-table, and record the path sequence and the total cost. Finally, output the optimal path sequence and its corresponding total cost for each operation and maintenance area;

[0058] C = M + W

[0059] (Equation 10)

[0060] Among them, C refers to the total cost of each operation and maintenance area, which is composed of the displacement time cost M between stations and the station maintenance time cost W.

[0061] Preferably, as described in step 4, once the proportion of the routes with exceeded operation and maintenance time in a certain area exceeds the threshold, the K-means clustering step will be re-executed to adjust the division of the area operation and maintenance center. The specific implementation process is as follows:

[0062] Step 4-1: First, calculate the total duration of the optimal routes in each operation and maintenance area, and count the areas where the total duration exceeds the maximum working duration. If the proportion exceeds the threshold, it is determined that the number of allocated operation and maintenance areas is too small, and then step 4-2 is carried out. If it does not exceed the threshold, directly output the path planning routes and the total cost of all target operation and maintenance sites;

[0063] Step 4-2: After increasing the number of operation and maintenance areas, return to step 2 and use the K-means clustering step again for re-division.

[0064] Preferably, as described in step 5, output the operation and maintenance path planning lines and the total cost required for all charging stations. The specific implementation process is as follows:

[0065] Step 5-1: Until the proportion of the reasonable working duration areas after re-division in step 4 exceeds the threshold or exceeds the maximum number of iterations, finally output the path planning routes and the total cost of all target operation and maintenance sites.

[0066] Preferably, a charging station operation and maintenance path planning system combining clustering site chain decomposition and reinforcement learning, the system comprising:

[0067] A site clustering module for screening and classifying sites at different geographical locations, and calculating the site chain with the shortest Euclidean distance between each other as the alternative pool of operation and maintenance route sites according to the pre-set number of operation and maintenance personnel;

[0068] A reinforcement learning module for planning a path with the shortest relative distance and the best comprehensive conditions selected by the operation and maintenance personnel from the alternative pool of operation and maintenance route sites;

[0069] A path optimization module for feedback on whether to increase or decrease the site chain based on the working time and the current number of operation and maintenance personnel, and adjusting the alternative pool of operation and maintenance route sites.

[0070] An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that: when the processor executes the program, the charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning is implemented.

[0071] A computer-readable storage medium, on which a computer instruction is stored, characterized in that: when the computer instruction is executed by a processor, the charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning is implemented.

[0072] Compared with the prior art, the advantages of the present invention are as follows: Through the charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning, in view of the situation that when the traditional heuristic path optimization solves the charging station operation and maintenance problem, if there are too many sites to be traversed, the calculation time required by the random mutation method will increase geometrically. If the data volume is large, it will cause multiple sites with large Euclidean distances to be adjacent in the path planning result, thus wasting a large amount of computing power. It is proposed that on the premise of clustering idea, combined with the K-Means algorithm, cities with close geographical intervals are used as smaller route optimization pools, and then the shortest path planning in each divided route optimization pool is obtained through reinforcement learning iteration. The exploration rate ε of the action selection ε-greedy strategy of the reinforcement learning algorithm and the reward function are improved to solve the problems faced in the actual operation and maintenance environment, maintain the balance between the optimization time efficiency and the solution of operation and maintenance problems. At the same time, taking advantage of the characteristics of the operation and maintenance process including repair and movement processes, the operation and maintenance time of the path is statistically calculated and used as another measurement index in addition to the shortest path, to solve the drawback that the allocation of urban agglomerations in individual route optimization pools is not reasonable enough and does not conform to the actual operation and maintenance environment, improve the feasibility of path optimization, improve the efficiency of the operation and maintenance process, and consider the situation of multiple causes of failures and uneven distribution of failure levels existing in the actual operation and maintenance environment. It is proposed to use the Q-learning algorithm of reinforcement learning to design a reward function that conforms to the actual operation and maintenance environment, so that the algorithm has the characteristics of learning from previous routes and experiences. Compared with the result of the traditional heuristic algorithm being a fixed path, this operation and maintenance path planning method is more in line with the actual operation and maintenance conditions of each region, more in line with the actual working ability of operation and maintenance employees, and can intelligently change the working path to improve the operation and maintenance satisfaction while ensuring the benefits of employees under various special working conditions, such as extreme weather conditions like freezing rain, and dynamically adjust the working hours. In addition, by comparing the operation and maintenance time of each optimized path with the working time of operation and maintenance employees, the number of clustering path divisions is feedback-adjusted, providing practical data support for the human resources department of the operation and maintenance main company to adjust the number of employees and regions, reducing the labor cost of the operation and maintenance main company while improving the working efficiency of operation and maintenance employees, and achieving a win-win situation for the charging station operation and maintenance main company, operation and maintenance employees, and customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 is a flowchart of the charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning;

[0074] Figure 2 is a diagram of the division of operation and maintenance sites after the clustering algorithm;

[0075] Figure 3 is a module division and interaction diagram corresponding to the charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning in an embodiment of the present invention based on the clustering idea and the reinforcement learning method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0076] The present invention will be further described in detail below in conjunction with the accompanying drawings and implementation processes.

[0077] Example: As Figure 1 shown is a flowchart of a charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning, which specifically includes the following steps:

[0078] Step 1: Collect the longitude and latitude information of all target charging stations through the Baidu Map Open Platform

[0079] Step 1-1: Obtain the longitude and latitude information of all target charging stations by calling the Baidu Map API interface;

[0080] Step 1-2: Go to the Baidu Map Open Platform to register and apply to become an individual developer, and then apply for an APIKey;

[0081] Step 1-3: For obtaining the longitude and latitude information of the charging station, the "geocoding" API interface needs to be used;

[0082] Step 1-4: After the Baidu Map server calls the interface, it will return a JSON-formatted response data, and the information of each charging station will include its longitude and latitude coordinates;

[0083] Step 1-5: Extract the longitude and latitude information of all target charging stations and store it in a suitable data structure for subsequent map annotation and operation and maintenance path planning, as shown in Table 1;

[0084] Table 1: Table of longitude and latitude values obtained by each target charging station in the example through calling the interface;

[0085]

[0086] Step 1-6: Mark the corresponding coordinates on the map to obtain a visual representation of the location information of all charging stations. The specific steps are to use a compilation tool (such as vscode) to mark the longitude and latitude information of each station obtained in Step 1-1, where the longitude and latitude information is stored in a specified Excel document, and then by calling a toolkit, such as matplotlib.pyplot in the Python language, convert the longitude and latitude fields into points on the map and output a PNG-format picture, as Figure 2 shown.

[0087] Step 2: With the number of operation and maintenance personnel as the limit, cluster and divide all stations into several operation and maintenance areas, and the Euclidean distance from the area center to the stations within the area is the shortest. The specific steps include the following:

[0088] Step 2-1: Initialize. Randomly select K points from the set of sites as the initial regional centers. K is the number of operation and maintenance personnel arranged initially, which is set to 3 in this embodiment. These regional centers can be actual points in the site set or randomly generated points;

[0089] Step 2-2: Assign regions to each site, that is, calculate the Euclidean distance from each data point to all centroids. For the data point x i and the first centroid μ Ⅰ , the Euclidean distance calculation formula is as shown in the following figure:

[0090]

[0091] where M refers to the level of the data for which the clustering algorithm needs to be used. In two-dimensional plane data, it is divided into longitude and latitude. d(x i , μ Ⅰ ) refers to the distance between the data point x i and the first centroid μ Ⅰ . It represents the square root of the sum of the squared differences between the i-th data point x im and the first centroid μ Ⅰm at the current level m;

[0092] Assign each site to the region numbered j corresponding to the nearest regional center, that is:

[0093]

[0094] where ||x i - μ j || and ||x i - μ Ⅰ || respectively refer to the Euclidean distances from the data point x i to the centroid μ j and the centroid μ Ⅰ . μ j refers to the centroid numbered j in the centroid set μ k numbered from 1 to k, that is, when the distance from the data point x i to other centroids μ j is less than the distance from the data point x i to the original centroid μ Ⅰ , it is assigned to the j-th new region C j ;

[0095]

[0096] Step 2-3: Update the regional centers for each region C j Calculate the average of the longitudes and latitudes of all sites within the region and use this average as the new regional center. The formula for the new regional center is:

[0097]

[0098] Among them, C j refers to the category numbered j, |C j | refers to the number of data points in category C j For all data points in region j, calculate the average value as the new centroid μ of category j j ;

[0099] Step 2-4: Repeat the steps of allocating regions and updating the region centers until the positions of the region centers no longer change or the maximum number of iterations is reached. The final result is as shown in Figure 2 shown.

[0100] Step 3: Use the reinforcement learning Q-learmimg algorithm to find the shortest operation and maintenance route for the sites within the region, which specifically includes the following steps:

[0101] Step 3-1: Configure the Q-learning algorithm environment and initialize and define the following concepts respectively:

[0102] State definition: Define the current site as the state, such as site S1, S2, etc. Each state represents the possible position where the operation and maintenance personnel may be currently;

[0103] Action set: The action set contains all possible operations that can move from the current site to adjacent sites. These actions constitute the movement options that the operation and maintenance personnel can take;

[0104] Reward function design: Among them, the movement reward is set according to the movement cost. For example, if the distance is 10 kilometers, the reward is -10 to reflect the economic cost of movement. Completion reward: When all sites are visited and the starting point (or the end point) is returned, a positive reward of +100 is given to encourage the algorithm to find a complete operation and maintenance path. Illegal movement penalty: If an attempt is made to move to an illegal site (such as a non-existent path), a high penalty of -50 is given to avoid the algorithm generating invalid or incorrect movement strategies. In this embodiment, the reward function is set as follows:

[0105]

[0106] Among them, cost(s,a) represents the movement cost of executing action a from state s, and the termination condition is reaching the target site or exceeding the maximum number of steps;

[0107] Termination condition: Set reaching the target site or exceeding the maximum number of steps as the termination condition to ensure that the algorithm finds the optimal solution or reaches a stable state within a reasonable time;

[0108] Step 3-2: Initialize Q-learning training

[0109] Create a Q - table according to the number of states and actions. Its size is the number of states multiplied by the number of actions, and the initial value is set to 0. The Q - table is used to store the expected return values for each action in each state.

[0110]

[0111] Among them, Q(s,a) is the return value obtained by the agent when completing action a in state s;

[0112] The hyperparameters of the Q - table are set as the learning rate α, which is used to control the update speed of new information over old information. The discount factor γ, which represents the importance of future returns relative to current returns. The exploration rate ε, which is used to balance between exploring new strategies and exploiting known optimal strategies.

[0113]

[0114] In this embodiment, let the learning rate α = 0.1, the discount factor γ = 0.9, and the exploration rate ε = 0.3;

[0115] Step 3 - 3: Execute the Q - learning training loop. At the beginning of each training episode, randomly select a starting site as the initial state s. Then perform action selection. According to the exploration rate ε, randomly select an action for exploration, or select the action with the largest Q - value in the current state in the Q - table for greedy selection. Execute the selected action and move to the new site s';

[0116]

[0117] Among them, ε is the exploration rate set initially, argmax a' Q(s',a') represents the maximum Q - value among all possible actions a' in the next state s'. It is an estimate of the maximum return that may be obtained in the future;

[0118] And calculate the reward r according to whether the move is legal, whether the path is completed, and whether an illegal move is attempted.

[0119] r = R(s,a,s')

[0120] (Equation 9)

[0121] Among them, R(s,a,s') is the value calculated according to the reward function R when the agent is in state s and executes action a to make the agent in state s', and this value is assigned to the parameter r and recorded in the Q - table;

[0122] Update the Q - value using the Bellman equation, that is:

[0123] Q(s,a) ← Q(s,a) + α[r + γ max_{a'}Q(s',a') - Q(s,a)]

[0124] (Equation 10)

[0125] Among them, Q(s,a) is the reward value obtained by the agent for completing action a in state s. α is the learning rate, a positive number between 0 and 1. It determines the speed at which new information overrides old information. r is the reward, the immediate reward obtained by the agent from the environment after executing the action. γ is the discount factor, which determines the current value of future rewards. max_{a'}Q(s',a') represents the maximum Q value among all possible actions a' in the next state s'. It is an estimate of the maximum reward that may be obtained in the future;

[0126] Step 3 - 4: Perform convergence judgment, that is, monitor the change of the Q - table. When the change of the Q - value in consecutive multiple trainings is less than the set threshold (such as 0.001), it is considered that the algorithm has converged to the optimal solution or reached a stable state;

[0127] Step 3 - 5: After the training is completed, set the exploration rate ε to 0, and only use the known optimal policy for path selection. Starting from the origin of the area, select the action with the maximum Q - value at each step according to the trained Q - table, and record the path sequence and the total cost. Finally, output the optimal path sequence and its corresponding total cost for each operation and maintenance area:

[0128] C = M + W

[0129] (Equation 11)

[0130] Among them, C refers to the total cost of each operation and maintenance area, which is composed of the displacement time cost M between stations and the station maintenance time cost W.

[0131] Step 4: Calculate the maintenance time of the shortest operation and maintenance route in the area, and judge whether the proportion of the reasonable working - hour area exceeds the threshold. Specifically, it includes the following steps:

[0132] Step 4 - 1: First, calculate the total duration of the optimal route in each operation and maintenance area, and count whether the proportion of the areas with a total duration exceeding the maximum working duration exceeds the threshold;

[0133] Step 4 - 2: If it exceeds the threshold, it is determined that the number of allocated operation and maintenance areas is too small. After increasing the number of operation and maintenance areas, return to Step 2 to re - divide again.

[0134] Step 5: Output the operation and maintenance path planning routes and the total cost required for all charging stations. Specifically, it includes the following steps:

[0135] Step 5-1: Until the proportion of the reasonable working duration area after re-division in Step 4 exceeds the threshold or exceeds the maximum number of iterations, finally output the path planning routes and the total cost of all target operation and maintenance sites. The optimal paths of one batch of sites are as follows:

[0136] Optimal path of site chain 1 = [S1, S2, S3, S4, S5, S1], total duration 7.1h < threshold 8h

[0137] Optimal path of site chain 2 = [S6, S7, S8, S9, S10], total duration 7.5h < threshold 8h

[0138] Optimal path of site chain 3 = [S11], total duration 5.8h < threshold 8h

[0139] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. A charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning, characterized in that: The steps include: (1) Collect and obtain the longitude and latitude information of all target charging stations, process them and mark them on the map; (2) Initialize the number of regional operation and maintenance centers and divide them into several operation and maintenance areas based on K-means clustering. Make the European distance from the regional operation and maintenance center to the sites in the region the shortest; (3) Use the reinforcement learning Q-learning algorithm to find the shortest operation and maintenance route for sites in the region; (4) Calculate the maintenance time corresponding to the shortest operation and maintenance route in each regional operation and maintenance. If the maintenance time exceeds the limited working time, re-divide the regional operation and maintenance center based on step (2); (5) Output the operation and maintenance path planning routes and route working time required for all charging stations.

2. The charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning according to claim 1 is characterized in that: Collect the longitude and latitude of all target charging stations and mark them on the map. The specific implementation process is as follows: Step 1: Obtain the latitude and longitude information of all target charging stations by calling the Baidu Map API interface; Step 2: Mark the corresponding coordinates on the map to obtain the location information of all charging stations.

3. The charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning according to claim 1 is characterized in that: Using the K-means algorithm in the clustering algorithm, all charging stations are divided into several areas through iterative optimization, so that the sum of the squares of the distances from the center of each area to other stations is minimized. The specific implementation process is as follows: Step 1: Initialize and randomly select K points from the site set as the initial regional centers. K is the number of operation and maintenance personnel initially arranged. These regional centers can be actual points in the site set or randomly generated points. Step 2: Assign a region to each site, that is, calculate the Euclidean distance from each data point to all centroids; for data point x i and the first centroid μ Ⅰ , the Euclidean distance calculation formula is shown in the following figure: Among them, M refers to the level of data that needs to use the clustering algorithm. In the plane, the two-dimensional data is divided into longitude and latitude, d(x i ,μ Ⅰ ) refers to the data point x i and the first centroid μ Ⅰ The distance, which means that in the current level m, the i-th data point x im and the first centroid μ Ⅰm The square root of the sum of squares of the differences; Assign each site to the region numbered j corresponding to the regional center closest to it, that is: Among them, ||x i -μ j || and ||x i -μ Ⅰ || refers to the data point x i To the center of mass μ j and the center of mass μ Ⅰ The Euclidean distance, μ j refers to the set of centroids numbered from 1 to k μ k The centroid numbered j in the i To other centroids μ j The distance is less than the data point x i To the original center of mass μ Ⅰ , then assign it to the jth new area C j middle; Step 3: For each region C i Update the regional center, calculate the average longitude and latitude of all sites in the region, and use the average as the new regional center. The calculation formula for the new regional center is: Among them, C j refers to the category number j, |C j |Refers to category C j The number of data points in the region j is calculated, and the average value of all data points in region j is used as the new centroid μ of category j. j ; Step 4: Repeat the steps of allocating regions and updating region centers until the position of the region center no longer changes or the maximum number of iterations is reached.

4. The charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning according to claim 1 is characterized in that: The reinforcement learning Q-learning algorithm is used to output the optimal path for several operation and maintenance areas s. The specific implementation process is as follows: Step 1: Configure the Q-learning algorithm: When configuring the Q-learning algorithm to be applied to several operation and maintenance areas, it is first necessary to model the environment of a single operation and maintenance area, which includes defining the state as the current site, the action as the operation that can be moved from the current site to the adjacent site; and the reward function, which is set according to the movement cost, path completion, and whether illegal movement is attempted. Specifically, each time you move from one site to another, the reward value is the negative value of the movement cost to encourage low-cost movement. When all sites are visited and return to the starting point (or reach the end point), a positive reward is given to encourage the completion of the path. If you try to move to an illegal site, a high penalty is given. The termination condition is set to reach the target site or exceed the maximum number of steps to avoid the algorithm from falling into an infinite loop. The reward function is set as follows: Among them, cost(s,a) represents the movement cost of executing action a from state s, the termination condition is reaching the target site or exceeding the maximum number of steps, Reward is the reward value for completing the path, and Penalty is the penalty value for performing illegal moves; Next, Q-learning training is performed for each operation and maintenance area. The training process starts with initialization, which includes creating a Q table with a size equal to the number of states multiplied by the number of actions, and setting its initial value to 0. Among them, Q(s,a) is the reward value obtained by the agent when completing action a in state s; At the same time, set the hyperparameters, including learning rate α, discount factor γ, and exploration rate ε; The training cycle includes multiple training episodes. Each episode starts from a randomly selected starting point. In the path exploration phase, the action a is randomly selected according to the exploration rate table, i.e., the following formula, or the action with the largest Q value in the current state in the Q table is selected. Then the action is executed and the corresponding reward is calculated. Among them, ε is the exploration rate set at the beginning, argmax a 'Q(s',a') represents the maximum Q value of all possible actions a' in the next state s', which is an estimate of the maximum possible reward in the future; The reward r is calculated based on whether the move is legal, whether the path is completed, and whether an illegal move is attempted. Subsequently, the Q table is updated using the Bellman equation, and the next action is selected based on the updated Q table; r=R(s,a,s') (Formula 8) Among them, R(s,a,s') is the value calculated by the reward function R when the action a is executed in state s to put the agent in state s', and this value is assigned to the parameter r and recorded in the Q table; This process continues until the path is completed or the maximum number of steps is reached. The basis for convergence is to monitor the changes in the Q table in the following formula. When the change in the Q value of multiple consecutive trainings is less than the set threshold, the algorithm is considered to have converged; Q(s,a)←Q(s,a)+α[r+γa'maxQ(s',a')-Q(s,a)] (Equation 9) Where Q(s,a) is the reward value obtained by the agent when completing action a in state s, α is the learning rate, a positive number between 0 and 1, which determines the speed at which new information covers old information, r is the reward, which is the immediate reward obtained by the agent from the environment after performing an action, γ is the discount factor, which determines the current value of future rewards, max_{a'}Q(s',a') represents the maximum Q value among all possible actions a' in the next state s', which is an estimate of the maximum reward that may be obtained in the future; Step 2: Output the optimal path for each operation and maintenance area: Finally, the optimal path of each operation and maintenance area is output. This step requires that the exploration rate is fixed to 0 in advance. Only the known optimal strategy is used. Starting from the starting point of the area, the maximum Q value action of each step is selected according to the trained Q table, and the path sequence and total cost are recorded. This process continues until the end point is reached or all sites are visited and returned to the starting point. Finally, the optimal path sequence and the corresponding total cost of each operation and maintenance area are output; C=M+W (Equation 10) Where C refers to the total cost of each operation and maintenance area, which consists of the inter-site displacement time cost M and the site maintenance time cost W.

5. The charging station operation and maintenance path planning method combining clustering site chain decomposition and reinforcement learning according to claim 1 is characterized in that: Calculate the total duration of the optimal route in each operation and maintenance area. If the proportion of areas where the total duration exceeds the maximum working duration exceeds the threshold, it is determined that the number of operation and maintenance areas is too small. After increasing the number of operation and maintenance areas, return to step (2) and divide again until the proportion of areas with reasonable working duration exceeds the threshold or exceeds the maximum number of iterations.

6. A charging station operation and maintenance path planning system combining clustering site chain decomposition and reinforcement learning, characterized in that: A charging station operation and maintenance path planning method combining clustered site chain decomposition and reinforcement learning for implementing any one of claims 1 to 5, the system comprising: The site clustering module is used to screen and classify sites in different geographical locations, and calculate the site chain with the shortest Euclidean distance between them according to the pre-set number of operation and maintenance personnel as the site candidate pool for the operation and maintenance route; The reinforcement learning module is used to plan a route with the shortest relative distance and the best comprehensive conditions, which is selected by the operation and maintenance personnel from the pool of candidate operation and maintenance route sites; The path optimization module is used to provide feedback on whether to increase or decrease the site chain based on working hours and the current number of operation and maintenance personnel, and adjust the site candidate pool for the operation and maintenance route.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the charging station operation and maintenance path planning method combining cluster site chain decomposition and reinforcement learning as described in any one of claims 1 to 5 above is implemented.

8. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the computer instructions are executed by the processor, the charging station operation and maintenance path planning method combining cluster site chain decomposition and reinforcement learning as described in any one of claims 1 to 5 is implemented.