A battery allocation method for battery swap cabinets based on deep learning
Through deep learning-based methods, predicting the battery requirements of battery swap cabinets and combining reinforcement learning generation and provisioning strategies, the problems of uneven resource allocation and lagging response in traditional provisioning methods are solved, and more efficient and lower-cost battery provisioning is achieved.
Patent Information
- Application Number
- CN202510209948.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-25
AI Technical Summary
When facing dynamic demand changes in traditional battery swap cabinet battery distribution methods, there are problems of uneven resource allocation, lagging response and high operating costs, making it difficult to effectively deal with fluctuations in users' battery swap demand.
Using a deep learning method, the real-time operation data of the battery swap cabinet is collected through IoT devices, a multi-dimensional feature data set is constructed, and a LSTM-GNN model integrating spatiotemporal features is used to predict battery demand. Combining factors such as battery health status and geographical location, a reinforcement learning model is constructed to generate a global allocation strategy, and the strategies are monitored and adjusted in real time.
It improves the accuracy and efficiency of battery deployment, reduces operating costs, improves user experience, and can better cope with the volatility of battery demand.
Smart Images

Figure CN119692735B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of battery allocation technology, and in particular to a battery allocation method for a battery swap cabinet based on deep learning. Background Art
[0002] The current battery allocation method for battery swap cabinets is to allocate resources based on historical data and preset rules. For example, by analyzing the long-term battery swap demand trend in a certain area, a certain number of batteries are allocated in advance to meet the predicted demand.
[0003] However, in actual applications, users' demand for battery swapping does not always follow a fixed pattern. For example, morning and evening commuting and high-frequency takeout time periods will cause sudden increases in demand. When demand surges, the allocation strategy cannot follow up in time, which will cause the battery inventory of some battery swap cabinets to be quickly depleted, and the user experience will decline accordingly. On the other hand, in order to avoid battery shortages, operators tend to over-replenish when demand is stable, resulting in long-term idleness of battery resources in some battery swap cabinets and low overall utilization.
[0004] In order to cope with the volatility of demand, some operators have increased manual inspections and multiple dispatches to make up for the shortcomings of the allocation strategy. However, this method is inefficient and costly, and it is difficult to provide effective protection during the real peak of demand. It can be seen that the traditional battery allocation method has significant problems of uneven resource allocation, delayed response and high operating costs when facing dynamic demand changes. A battery allocation method for battery swap cabinets based on deep learning is urgently needed to solve such problems. Summary of the invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a battery allocation method for battery swap cabinets based on deep learning to solve the problem of surge in battery swap demand during peak hours and battery shortage in some battery swap cabinets.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] An embodiment of the present invention provides a battery allocation method for a battery swap cabinet based on deep learning, which includes:
[0009] Step S1, collecting real-time operation data of the power exchange cabinet through the Internet of Things (IoT) device to build a multi-dimensional feature data set;
[0010] Step S2, based on the multi-dimensional feature data set constructed in step S1, using the deep learning LSTM-GNN model integrating spatiotemporal features, predicting the battery demand of each battery swap cabinet in a specific time period in the future, and generating demand distribution data;
[0011] Step S3, combining the demand distribution data generated in step S2, obtain the real-time health status data of the batteries in the battery swap cabinet, and analyze the battery health status: evaluate the availability of the current inventory of the battery swap cabinet, and calculate the number of batteries that meet the demand;
[0012] And determine the number and priority of batteries to be deployed based on their health status and urgency of demand;
[0013] Step S4: Based on the number and priority of batteries to be deployed determined in step S3, a reinforcement learning model is constructed. The model takes the battery deployment path and the number of allocations as optimization goals, and combines the geographical location, deployment cost, and battery health status between the battery swap cabinets to generate a global deployment strategy.
[0014] Step S5, apply the deployment strategy generated in step S4, and monitor the battery usage and demand fluctuations of the battery swap cabinet after deployment in real time; based on the real-time monitoring data, dynamically adjust the deployment strategy generated in step S4.
[0015] As a preferred solution of the deep learning-based battery allocation method for battery swap cabinets described in the present invention, the real-time operation data includes battery inventory, historical battery swap records, user behavior, surrounding traffic flow and weather conditions.
[0016] As a preferred solution of the deep learning-based battery allocation method for battery swap cabinets described in the present invention, the demand distribution data includes the expected number of battery swaps and the corresponding required number of batteries for each battery swap cabinet in the target time period.
[0017] As a preferred solution of the battery allocation method of the battery swap cabinet based on deep learning described in the present invention, the step of predicting the battery demand of each battery swap cabinet in a specific time period in the future and generating demand distribution data is as follows:
[0018] Constructing the network topology of the battery swap cabinet , defined as:
[0019] ,
[0020] in, It represents the network diagram of the power exchange cabinet. Represents a battery swap cabinet set, each node represents a battery swap cabinet, represents the edge set between nodes, indicating the possible deployment path between the battery swap cabinets. Represents the weight set of edges, which is used to represent the deployment cost, including distance and time.
[0021] Define node feature matrix and the adjacency matrix :
[0022] ,
[0023] in, and Respectively represent a A matrix and a The matrix of Indicates the number of power exchange cabinets. Represents the characteristic dimension of each node (battery inventory, battery replacement frequency, surrounding traffic flow, etc.), is the adjacency matrix;
[0024] Use a multi-layer graph convolutional network GCN to extract spatial features. The extraction process is expressed as:
[0025] ,
[0026] in, Indicates The node feature matrix of the layer, initially represents the normalized adjacency matrix with self-loops added, represents the identity matrix, represents the degree matrix, Indicates The weight matrix of the layer, represents the ReLU activation function,
[0027] Extracted spatial features And historical battery replacement records Input LSTM for time dynamic modeling, the modeling formula is:
[0028] ,
[0029] in, Indicates time The hidden state of Indicates time The spatial feature matrix of Indicates time The hidden state of
[0030] Output the LSTM Mapped to the battery demand of each battery swap cabinet, the mapping formula is:
[0031] ,
[0032] in, Indicates time Demand distribution of each battery swap cabinet, and is a trainable parameter, Used to normalize the output into a probability distribution.
[0033] As a preferred solution of the battery allocation method for battery swap cabinets based on deep learning described in the present invention, the health status data includes the remaining capacity, cycle life and internal resistance change of the battery.
[0034] As a preferred solution of the battery allocation method for a battery swap cabinet based on deep learning described in the present invention, wherein: the step of combining the demand distribution data generated in step S2 to obtain the real-time health status data of the battery in the battery swap cabinet and analyzing the battery health status is as follows:
[0035] Get the battery health status data of the battery swap cabinet and define the battery health status as a matrix :
[0036] ,
[0037] in, Represents the health status matrix of each battery in the battery swap cabinet, Indicates the number of power exchange cabinets. Indicates the maximum number of batteries in each battery swap cabinet. Indicates In the power exchange cabinet The health status value of each battery;
[0038] Defining health status is a weighted combination index, expressed as:
[0039] ,
[0040] in, Indicates the remaining battery capacity. Indicates the battery cycle life, Indicates the change in the internal resistance of the battery. Represents weight, satisfying ;
[0041] According to health status and forecast demand distribution , calculate the battery swap cabinet Number of available batteries , the calculation formula is:
[0042] ,
[0043] in, is the indicator function, when hour, , otherwise , is the threshold of health status,
[0044] Calculate the gap in the number of batteries required to meet demand using the following formula:
[0045] ,
[0046] in, Indicates power exchange cabinet The gap between battery demand and current inventory, Indicates power exchange cabinet In time Battery requirements,
[0047] According to the demand gap and health status , define the deployment priority score :
[0048] ,
[0049] in, For battery swap cabinet The first item represents the demand ratio, and the second item represents the normalized weight of the health status ratio.
[0050] As an optimal solution of the battery allocation method for battery swap cabinets based on deep learning described in the present invention, the allocation strategy includes specific allocation paths and time nodes.
[0051] As a preferred solution of the battery allocation method for battery swap cabinets based on deep learning described in the present invention, the steps of constructing a reinforcement learning model, which takes the battery allocation path and allocation quantity as optimization goals, combines the geographical location, allocation cost and battery health status between the battery swap cabinets, and generates a global allocation strategy are as follows:
[0052] Model battery deployment as a Markov decision process MDP and define the state space , Action Space , reward function and state transition probability :
[0053] ,
[0054] in, Indicates the battery inventory and health status of the current battery swap cabinet. Indicates executable deployment actions, including deployment path and quantity. The reward value after executing an action measures the efficiency and cost of deployment. Represents the state transition probability, and the next state is determined by the current state and action;
[0055] Defining states for:
[0056] ,
[0057] in, Indicates time The battery health status matrix of each battery swap cabinet, Indicates time The battery demand gap matrix, Represents the geographical location matrix between the battery swap cabinets,
[0058] Defining actions for:
[0059] ,
[0060] in, Represents the deployment path matrix, the elements on the path Indicates power exchange cabinet arrive The deployment priority of represents the deployment quantity matrix, in which the elements For battery swap cabinet Deploy to The number of batteries,
[0061] Reward Function The comprehensive deployment cost and demand satisfaction are expressed as:
[0062] ,
[0063] in, Indicates time The total reward value, Indicates the number of power exchange cabinets. is the weight of demand satisfaction, To adjust the cost weight, Indicates time No. The demand gap of each battery swap cabinet is Indicates that all power-swap cabinets are at time The total demand gap Indicates power exchange cabinet arrive The deployment cost, Indicates power exchange cabinet Deploy to The number of batteries;
[0064] The policy network is optimized using the policy gradient-based reinforcement learning algorithm PPO, with the goal of maximizing the reward:
[0065] ,
[0066] in, Representation strategy The gradient of Indicated in strategy The expected value under represents the gradient of the policy network parameters, Indicates that the policy network is in state Next select action The probability of Indicates time The advantage function of
[0067] After training is completed, the policy network generates the optimal deployment strategy, and the generation formula is:
[0068] ,
[0069] in, For the optimal deployment strategy, Indicates the time frame for deployment decisions, Represents the expected function.
[0070] As a preferred solution of the battery allocation method for a battery swap cabinet based on deep learning described in the present invention, wherein: the demand fluctuation is determined by combining the predicted data of step S2 with the difference between the actual battery swap demand of the battery swap cabinet;
[0071] The battery usage is dynamically collected by IoT devices, including real-time battery replacement operations of the battery replacement cabinet, battery inventory changes and abnormal conditions, among which abnormal conditions include equipment failures and abnormal battery aging.
[0072] As a preferred solution of the battery allocation method of the power swap cabinet based on deep learning described in the present invention, wherein: the real-time monitoring of the battery usage and demand fluctuation of the power swap cabinet after allocation; based on the real-time monitoring data, the step of dynamically adjusting the allocation strategy generated in step S4 is,
[0073] Dynamically collect the operation data of the power exchange cabinet through the IoT device to build a real-time monitoring data set :
[0074] ,
[0075] in, Indicates time Real-time monitoring data collection, Indicates time Real-time battery swap operation data, including swap frequency and battery consumption, Indicates time The battery inventory change data of the battery swap cabinet, Indicates time abnormal data, including equipment failure and abnormal battery aging;
[0076] Based on real-time data and the demand predicted in step S2 , calculate the actual demand matrix of the battery swap cabinet And the demand difference matrix, the calculation formula is:
[0077] ,
[0078] in, Indicates time The demand difference matrix,
[0079] Indicates time The actual demand matrix, from extract, Indicates time Forecasted demand matrix,
[0080] Based on the demand difference matrix and battery health status , dynamically adjust the deployment priority matrix , the adjustment formula is:
[0081] ,
[0082] in, Indicates time Updated deployment priority matrix, Indicates time The original priority matrix, To update the weights, Indicates time The maximum value of the demand difference of all battery swap cabinets;
[0083] Using the updated priority matrix Adjust the deployment path and quantity. The adjustment formula is:
[0084] ,
[0085] in, Indicates time The adjusted deployment path matrix, Indicates time The adjusted allocation quantity matrix, Represents the function of replanning the path and quantity, combined with and deployment cost optimization results,
[0086] According to the adjusted deployment strategy and real-time data , calculate the deployment effect , the calculation formula is:
[0087] ,
[0088] in, Indicates time The deployment effect, Indicates the total number of power-swap cabinets. For battery swap cabinet The demand satisfaction weight, For battery swap cabinet The actual demand satisfaction rate, Represents the total number of all deployment paths, For path The allocation cost weight, For path The actual deployment cost,
[0089] Based on the deployment effect , optimize the policy network and generate an updated deployment strategy :
[0090] ,
[0091] in, Indicates time The optimized deployment strategy Indicates time Original Strategy Network, Represents the policy optimization function in reinforcement learning.
[0092] The beneficial effects of the present invention are as follows: the present invention constructs a data set by combining battery inventory, user behavior and traffic flow, and adopts a deep learning LSTM-GNN model that integrates spatiotemporal features to predict the battery demand distribution in a specific time period. The model constructs the topological structure of the battery swap cabinet network, captures spatial dependencies, and realizes time dynamic modeling in combination with historical battery swap records to improve the accuracy of the prediction; by analyzing the real-time health status of the batteries in the battery swap cabinet, the availability of the current inventory is evaluated in combination with the demand forecast results, the number of batteries that meet the demand is calculated, and the priority of battery allocation is determined according to the urgency of the demand and the health status; in terms of allocation strategy generation, the battery allocation problem is modeled as a Markov decision process, and an optimization model is constructed through the reinforcement learning algorithm PPO. The allocation path, quantity and cost are comprehensively integrated to generate a global optimal allocation strategy. In addition, the battery usage and demand fluctuations of the battery swap cabinet are monitored in real time, and the allocation priority is dynamically adjusted by using the difference between the demand forecast and the actual demand, and the allocation strategy is updated in real time in combination with the optimized path and quantity to improve the adaptability of the allocation plan.
[0093] In summary, the present invention aims to solve the problems of insufficient demand forecasting, uneven resource allocation and delayed response in traditional methods by combining demand forecasting, real-time response and deployment optimization to reduce operating costs while improving the utilization efficiency of battery deployment and user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0094] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0095] Figure 1 It is a flow chart of the battery allocation method of the battery swap cabinet based on deep learning of the present invention. DETAILED DESCRIPTION
[0096] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0097] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0098] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0099] Example 1, reference Figure 1 This embodiment provides a battery allocation method for a power swap cabinet based on deep learning, comprising the following steps:
[0100] Step S1, collecting real-time operation data of the power exchange cabinet through the Internet of Things (IoT) device to build a multi-dimensional feature data set;
[0101] Real-time operational data includes battery inventory, historical battery replacement records, user behavior, surrounding traffic flow and weather conditions.
[0102] Step S2, based on the multidimensional feature data set constructed in step S1, using the deep learning LSTM-GNN model integrating spatiotemporal features, predicting the battery demand of each battery swap cabinet in a specific time period in the future, and generating demand distribution data;
[0103] The demand distribution data includes the expected number of battery swaps for each battery swap cabinet in the target time period and the corresponding number of batteries required;
[0104] The steps to predict the battery demand of each battery swap cabinet in a specific time period in the future and generate demand distribution data are as follows:
[0105] Constructing the network topology of the battery swap cabinet , defined as:
[0106] ,
[0107] in, It represents the network diagram of the power exchange cabinet. Represents a battery swap cabinet set, each node represents a battery swap cabinet, represents the edge set between nodes, indicating the possible deployment path between the battery swap cabinets. Represents the weight set of edges, which is used to represent the deployment cost, including distance and time.
[0108] Define node feature matrix and the adjacency matrix :
[0109] ,
[0110] in, and Respectively represent a A matrix and a The matrix of Indicates the number of power exchange cabinets. Represents the characteristic dimension of each node (battery inventory, battery replacement frequency, surrounding traffic flow, etc.), is the adjacency matrix;
[0111] A multi-layer graph convolutional network GCN is used to extract spatial features. The extraction process is expressed as:
[0112] ,
[0113] in, Indicates The node feature matrix of the layer, initially represents the normalized adjacency matrix with self-loops added, represents the identity matrix, represents the degree matrix, Indicates The weight matrix of the layer, represents the ReLU activation function,
[0114] Extracted spatial features And historical battery replacement records Input LSTM for time dynamic modeling, the modeling formula is:
[0115] ,
[0116] in, Indicates time The hidden state of Indicates time The spatial feature matrix of Indicates time The hidden state of
[0117] Output the LSTM Mapped to the battery demand of each battery swap cabinet, the mapping formula is:
[0118] ,
[0119] in, Indicates time Demand distribution of each battery swap cabinet, and is a trainable parameter, Used to normalize the output into a probability distribution;
[0120] Specifically, the power cabinet network diagram is represented as an adjacency matrix and a node feature matrix. GNN is then used to capture the topological relationship and extract the spatial dependencies between battery swap cabinets. LSTM is used to perform temporal dynamic modeling of historical battery swap data to achieve multi-dimensional prediction of battery swap demand. Then, through the fully connected layer and Softmax mapping, the battery demand distribution of each battery swap cabinet in a specific time period is obtained, which can more accurately reflect the changing trend of battery swap demand.
[0121] Step S3, combining the demand distribution data generated in step S2, obtain the real-time health status data of the batteries in the battery swap cabinet, and analyze the battery health status: evaluate the availability of the current inventory of the battery swap cabinet, and calculate the number of batteries that meet the demand;
[0122] And determine the number and priority of batteries to be deployed based on their health status and urgency of demand;
[0123] Health status data includes the battery’s remaining capacity, cycle life, and internal resistance change;
[0124] Combined with the demand distribution data generated in step S2, the real-time health status data of the battery in the battery swap cabinet is obtained, and the steps for analyzing the battery health status are as follows:
[0125] Get the battery health status data of the battery swap cabinet and define the battery health status as a matrix :
[0126] ,
[0127] in, Represents the health status matrix of each battery in the battery swap cabinet, Indicates the number of power exchange cabinets. Indicates the maximum number of batteries in each battery swap cabinet. Indicates In the power exchange cabinet The health status value of each battery;
[0128] Defining health status is a weighted combination index, expressed as:
[0129] ,
[0130] in, Indicates the remaining battery capacity. Indicates the battery cycle life, Indicates the change in the internal resistance of the battery. Represents weight, satisfying ;
[0131] According to health status and forecast demand distribution , calculate the battery swap cabinet Number of available batteries , the calculation formula is:
[0132] ,
[0133] in, is the indicator function, when hour, , otherwise , is the threshold of health status,
[0134] Calculate the gap in the number of batteries required to meet demand using the following formula:
[0135] ,
[0136] in, Indicates power exchange cabinet The gap between battery demand and current inventory, Indicates power exchange cabinet In time Battery requirements,
[0137] According to the demand gap and health status , define the deployment priority score :
[0138] ,
[0139] in, For battery swap cabinet The first item represents the demand ratio, and the second item represents the normalized weight of the health status ratio;
[0140] Specifically, based on the real-time battery health status data of the battery swap cabinet, a health status matrix is constructed, and the health status is evaluated based on the remaining capacity, cycle life and internal resistance change. Combined with the battery demand distribution predicted in step S2, the battery demand gap of the battery swap cabinet is calculated, and the number of batteries to be deployed is determined; then, the deployment priority score is calculated comprehensively based on the demand urgency and health status to ensure the accuracy of the battery deployment quantity.
[0141] Step S4: Based on the number and priority of batteries to be deployed determined in step S3, a reinforcement learning model is constructed. The model takes the battery deployment path and the number of allocations as optimization goals, and combines the geographical location, deployment cost, and battery health status between the battery swap cabinets to generate a global deployment strategy.
[0142] The deployment strategy includes specific deployment paths and time nodes;
[0143] Construct a reinforcement learning model, which takes the battery deployment path and allocation quantity as the optimization goal, combines the geographical location, deployment cost and battery health status between the battery swap cabinets, and generates the steps of the global deployment strategy as follows:
[0144] Model battery deployment as a Markov decision process MDP and define the state space , Action Space , reward function and state transition probability :
[0145] ,
[0146] in, Indicates the battery inventory and health status of the current battery swap cabinet. Indicates executable deployment actions, including deployment path and quantity. The reward value after executing an action measures the efficiency and cost of deployment. Represents the state transition probability, and the next state is determined by the current state and action;
[0147] Defining states for:
[0148] ,
[0149] in, Indicates time The battery health status matrix of each battery swap cabinet, Indicates time The battery demand gap matrix, Represents the geographical location matrix between the battery swap cabinets,
[0150] Defining actions for:
[0151] ,
[0152] in, Represents the deployment path matrix, the elements on the path Indicates power exchange cabinet arrive The deployment priority of represents the deployment quantity matrix, in which the elements For battery swap cabinet Deploy to The number of batteries,
[0153] Reward Function The comprehensive deployment cost and demand satisfaction are expressed as:
[0154] ,
[0155] in, Indicates time The total reward value, Indicates the number of power exchange cabinets. is the weight of demand satisfaction, To adjust the cost weight, Indicates time No. The demand gap of each battery swap cabinet is Indicates that all power-swap cabinets are at time The total demand gap Indicates power exchange cabinet arrive The deployment cost, Indicates power exchange cabinet Deploy to The number of batteries;
[0156] The policy network is optimized using the policy gradient-based reinforcement learning algorithm PPO, with the goal of maximizing the reward:
[0157] ,
[0158] in, Representation strategy The gradient of Indicated in strategy The expected value under represents the gradient of the policy network parameters, Indicates that the policy network is in state Next select action The probability of Indicates time The advantage function of
[0159] After training is completed, the policy network generates the optimal deployment strategy, and the generation formula is:
[0160] ,
[0161] in, For the optimal deployment strategy, Indicates the time frame for deployment decisions, represents the expected function;
[0162] Specifically, through reinforcement learning modeling, the battery allocation problem is abstracted into a Markov decision process, and the state, action and reward functions are defined. In particular, the demand satisfaction and allocation cost are quantified as optimization objectives through the reward function, and the policy network is optimized in combination with the policy gradient algorithm. Finally, the policy network generates an allocation strategy with global benefits as the goal, including specific allocation paths and time nodes, which effectively ensures the globality and dynamic adaptability of the allocation decision.
[0163] Step S5, applying the deployment strategy generated in step S4, and monitoring the battery usage and demand fluctuation of the deployed battery swap cabinet in real time; dynamically adjusting the deployment strategy generated in step S4 based on the real-time monitoring data;
[0164] The demand fluctuation is determined by combining the predicted data of step S2 with the difference between the actual battery swapping demand of the battery swapping cabinet;
[0165] Battery usage is dynamically collected by IoT devices, including real-time battery swapping operations of battery swap cabinets, battery inventory changes, and abnormal conditions, including equipment failures and abnormal battery aging.
[0166] Real-time monitoring of battery usage and demand fluctuations of the battery swap cabinet after deployment; based on the real-time monitoring data, the steps of dynamically adjusting the deployment strategy generated in step S4 are as follows:
[0167] Dynamically collect the operation data of the power exchange cabinet through the IoT device to build a real-time monitoring data set :
[0168] ,
[0169] in, Indicates time Real-time monitoring data collection, Indicates time Real-time battery swap operation data, including swap frequency and battery consumption, Indicates time The battery inventory change data of the battery swap cabinet, Indicates time abnormal data, including equipment failure and abnormal battery aging;
[0170] Based on real-time data and the demand predicted in step S2 , calculate the actual demand matrix of the battery swap cabinet And the demand difference matrix, the calculation formula is:
[0171] ,
[0172] in, Indicates time The demand difference matrix,
[0173] Indicates time The actual demand matrix, from extract, Indicates time Forecasted demand matrix,
[0174] Based on the demand difference matrix and battery health status , dynamically adjust the deployment priority matrix , the adjustment formula is:
[0175] ,
[0176] in, Indicates time Updated deployment priority matrix, Indicates time The original priority matrix, To update the weights, Indicates time The maximum value of the demand difference of all battery swap cabinets;
[0177] Using the updated priority matrix Adjust the deployment path and quantity. The adjustment formula is:
[0178] ,
[0179] in, Indicates time The adjusted deployment path matrix, Indicates time The adjusted allocation quantity matrix, Represents the function of replanning the path and quantity, combined with and deployment cost optimization results,
[0180] According to the adjusted deployment strategy and real-time data , calculate the deployment effect , the calculation formula is:
[0181] ,
[0182] in, Indicates time The deployment effect, Indicates the total number of power-swap cabinets. For battery swap cabinet The demand satisfaction weight, For battery swap cabinet The actual demand satisfaction rate, Represents the total number of all deployment paths, For path The allocation cost weight, For path The actual deployment cost,
[0183] Based on the deployment effect , optimize the policy network and generate an updated deployment strategy :
[0184] ,
[0185] in, Indicates time The optimized deployment strategy Indicates time Original Strategy Network, Represents the policy optimization function in reinforcement learning;
[0186] Specifically, in step S5, combined with the dynamic characteristics of real-time data collection, the demand difference matrix and the battery health status matrix are used to dynamically adjust the deployment priority. Through real-time optimization of the deployment path and quantity, the adjusted deployment effect is verified and used as a basis to optimize the strategy network, ensure the adjustment ability of the strategy, and improve the intelligence of the overall deployment plan.
[0187] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A battery allocation method for a power swap cabinet based on deep learning, characterized in that: include, Step S1, collecting real-time operation data of the power exchange cabinet through the Internet of Things (IoT) device to construct a multi-dimensional feature data set; Step S2, based on the multi-dimensional feature data set constructed in step S1, using the deep learning LSTM-GNN model integrating spatiotemporal features, predicting the battery demand of each battery swap cabinet in a specific time period in the future, and generating demand distribution data; Step S3, combining the demand distribution data generated in step S2, obtain the real-time health status data of the batteries in the battery swap cabinet, and analyze the battery health status: evaluate the availability of the current inventory of the battery swap cabinet, and calculate the number of batteries that meet the demand; And determine the number and priority of batteries to be deployed based on their health status and urgency of demand; Step S4: Based on the number and priority of batteries to be deployed determined in step S3, a reinforcement learning model is constructed. The model takes the battery deployment path and the number of allocations as optimization goals, and combines the geographical location, deployment cost, and battery health status between the battery swap cabinets to generate a global deployment strategy. The reinforcement learning model is constructed, which takes the battery deployment path and allocation quantity as the optimization target, combines the geographical location, deployment cost and battery health status between the battery swap cabinets, and generates the global deployment strategy in the following steps: Model battery deployment as a Markov decision process MDP and define the state space , Action Space , reward function and state transition probability : , in, Indicates the battery inventory and health status of the current battery swap cabinet. Indicates executable deployment actions, including deployment path and quantity. The reward value after executing an action measures the efficiency and cost of deployment. Represents the state transition probability, and the next state is determined by the current state and action; Defining states for: , in, Indicates time The battery health status matrix of each battery swap cabinet, Indicates time The battery demand gap matrix, Represents the geographical location matrix between the battery swap cabinets, Defining actions for: , in, Represents the deployment path matrix, the elements on the path Indicates power exchange cabinet arrive The deployment priority of represents the deployment quantity matrix, in which the elements For battery swap cabinet Deploy to The number of batteries, Reward Function The comprehensive allocation cost and demand satisfaction are expressed as: , in, Indicates time The total reward value, Indicates the number of power exchange cabinets. is the weight of demand satisfaction, To adjust the cost weight, Indicates time No. The demand gap of each battery swap cabinet is Indicates that all power-swap cabinets are at time The total demand gap Indicates power exchange cabinet arrive The deployment cost, Indicates power exchange cabinet Deploy to The number of batteries; The policy network is optimized using the policy gradient-based reinforcement learning algorithm PPO, with the goal of maximizing the reward: , in, Representation strategy The gradient of Indicated in strategy The expected value under represents the gradient of the policy network parameters, Indicates that the policy network is in state Next select action The probability of Indicates time The advantage function of After training is completed, the policy network generates the optimal deployment strategy, and the generation formula is: , in, For the optimal deployment strategy, Indicates the time frame for deployment decisions, represents the expected function; Step S5, apply the deployment strategy generated in step S4, and monitor the battery usage and demand fluctuations of the battery swap cabinet after deployment in real time; based on the real-time monitoring data, dynamically adjust the deployment strategy generated in step S4.
2. A battery allocation method for a power swap cabinet based on deep learning as claimed in claim 1, characterized in that: The real-time operational data includes battery inventory, historical battery replacement records, user behavior, surrounding traffic flow and weather conditions.
3. A battery allocation method for a power swap cabinet based on deep learning as described in claim 2, characterized in that: The demand distribution data includes the expected number of battery swaps for each battery swap cabinet in the target time period and the corresponding number of batteries required.
4. A battery allocation method for a power swap cabinet based on deep learning as described in claim 3, characterized in that: The step of predicting the battery demand of each battery swap cabinet in a specific time period in the future and generating demand distribution data is as follows: Constructing the network topology of the battery swap cabinet , defined as: , in, It represents the network diagram of the power exchange cabinet. Represents a battery swap cabinet set, each node represents a battery swap cabinet, represents the edge set between nodes, indicating the possible deployment path between the battery swap cabinets. Represents the weight set of edges, which is used to represent the deployment cost, including distance and time. Define node feature matrix and the adjacency matrix : , in, and Respectively represent a A matrix and a The matrix of Indicates the number of power exchange cabinets. represents the feature dimension of each node, is the adjacency matrix; A multi-layer graph convolutional network GCN is used to extract spatial features. The extraction process is expressed as: , in, Indicates The node feature matrix of the layer, initially represents the normalized adjacency matrix with self-loops added, represents the identity matrix, represents the degree matrix, Indicates The weight matrix of the layer, represents the ReLU activation function, Extracted spatial features And historical battery replacement records Input LSTM for time dynamic modeling, the modeling formula is: , in, Indicates time The hidden state of Indicates time The spatial feature matrix of Indicates time The hidden state of Output the LSTM Mapped to the battery demand of each battery swap cabinet, the mapping formula is: , in, Indicates time Demand distribution of each battery swap cabinet, and is a trainable parameter, Used to normalize the output into a probability distribution.
5. A battery allocation method for a power swap cabinet based on deep learning as described in claim 4, characterized in that: Health status data includes the battery's remaining capacity, cycle life, and internal resistance change.
6. A battery allocation method for a power swap cabinet based on deep learning as claimed in claim 5, characterized in that: The step of combining the demand distribution data generated in step S2 to obtain the real-time health status data of the battery in the battery swap cabinet and analyzing the battery health status is as follows: Get the battery health status data of the battery swap cabinet and define the battery health status as a matrix : , in, Represents the health status matrix of each battery in the battery swap cabinet, Indicates the number of power exchange cabinets. Indicates the maximum number of batteries in each battery swap cabinet. Indicates In the power exchange cabinet The health status value of each battery; Defining health status is a weighted combination index, expressed as: , in, Indicates the remaining battery capacity. Indicates the battery cycle life, Indicates the change in the internal resistance of the battery. Represents weight, satisfying ; According to health status and forecast demand distribution , calculate the battery swap cabinet Number of available batteries , the calculation formula is: , in, is the indicator function, when hour, , otherwise , is the threshold of health status, Calculate the gap in the number of batteries required to meet demand using the following formula: , in, Indicates power exchange cabinet The gap between battery demand and current inventory, Indicates power exchange cabinet In time Battery requirements, According to the demand gap and health status , define the deployment priority score : , in, For battery swap cabinet The first item represents the demand ratio, and the second item represents the normalized weight of the health status ratio.
7. A battery allocation method for a power swap cabinet based on deep learning as claimed in claim 6, characterized in that: The deployment strategy includes specific deployment paths and time nodes.
8. A battery allocation method for a power swap cabinet based on deep learning as claimed in claim 7, characterized in that: The demand fluctuation is determined by combining the predicted data of step S2 with the difference in the actual battery swapping demand of the battery swapping cabinet; The battery usage is dynamically collected by IoT devices, including real-time battery replacement operations of the battery replacement cabinet, battery inventory changes and abnormal conditions, among which abnormal conditions include equipment failures and abnormal battery aging.
9. A battery allocation method for a power swap cabinet based on deep learning as claimed in claim 8, characterized in that: The step of real-time monitoring of battery usage and demand fluctuations of the battery swap cabinet after deployment; dynamically adjusting the deployment strategy generated in step S4 based on the real-time data of monitoring is as follows: Dynamically collect the operation data of the power exchange cabinet through the IoT device to build a real-time monitoring data set : , in, Indicates time Real-time monitoring data collection, Indicates time Real-time battery swap operation data, including swap frequency and battery consumption, Indicates time The battery inventory change data of the battery swap cabinet, Indicates time abnormal data, including equipment failure and abnormal battery aging; Based on real-time data and the demand predicted in step S2 , calculate the actual demand matrix of the battery swap cabinet And the demand difference matrix, the calculation formula is: , in, Indicates time The demand difference matrix, Indicates time The actual demand matrix, from extract, Indicates time Forecasted demand matrix, Based on the demand difference matrix and battery health status , dynamically adjust the deployment priority matrix , the adjustment formula is: , in, Indicates time Updated deployment priority matrix, Indicates time The original priority matrix, To update the weights, Indicates time The maximum value of the demand difference of all battery swap cabinets; Using the updated priority matrix Adjust the deployment path and quantity. The adjustment formula is: , in, Indicates time The adjusted deployment path matrix, Indicates time The adjusted allocation quantity matrix, Represents the function of replanning the path and quantity, combined with and deployment cost optimization results, According to the adjusted deployment strategy and real-time data , calculate the deployment effect , the calculation formula is: , in, Indicates time The deployment effect, Indicates the total number of power-swap cabinets. For battery swap cabinet The demand satisfaction weight, For battery swap cabinet The actual demand satisfaction rate, Represents the total number of all deployment paths, For path The allocation cost weight, For path The actual deployment cost, Based on the deployment effect , optimize the policy network and generate an updated deployment strategy : , in, Indicates time The optimized deployment strategy Indicates time Original Strategy Network, Represents the policy optimization function in reinforcement learning.
Citation Information
Patent Citations
Battery scheduling method, system and device based on deep reinforcement learning, and medium
CN116542498A
Intelligent charging method of battery charging and replacing cabinet based on reinforcement learning algorithm
CN118182238A