Vehicle management method and system based on deep reinforcement learning model

Through the vehicle management method of deep reinforcement learning model, the problem of unbalanced vehicle scheduling in the shared vehicle system is solved, precise scheduling and energy conversion processing are achieved, and user experience and operational efficiency are improved.

CN120672004APending Publication Date: 2025-09-19YOUON TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410310314.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-19
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Unbalanced vehicle scheduling in the shared vehicle system leads to insufficient or excessive electricity/hydrogen in the station, making it difficult for users to borrow vehicles, and low efficiency in battery/hydrogen replacement, which affects user experience and causes economic losses to operators.

Method used

A vehicle management method based on a deep reinforcement learning model is adopted. By collecting historical dynamic data and vehicle demand data, training sample sets and prediction sample sets are generated, and a vehicle demand predictor is established to achieve accurate scheduling and energy conversion processing of shared vehicles.

Benefits of technology

It improves the efficiency of shared vehicle dispatching, meets user needs, reduces the probability of vehicles with insufficient energy storage, and improves user experience and operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672004A_ABST
    Figure CN120672004A_ABST
Patent Text Reader

Abstract

The invention discloses a vehicle management method and system based on a deep reinforcement learning model. The vehicle management method comprises the following steps: collecting historical dynamic data and vehicle actual demand data corresponding to each parking station of a vehicle in a region in each time period, and generating a training sample set and a prediction sample set; training the neural network model based on deep learning to obtain a trained neural network model; generating a vehicle demand predictor according to the trained neural network model and the prediction sample set; and managing the vehicles at the parking stations according to the vehicle prediction demand data generated by the vehicle demand predictor. According to the vehicle management method, the predicted demand data of the vehicle is predicted through the vehicle demand predictor, accurate scheduling of the shared vehicle is realized, transduction processing of the arrival vehicle is also realized, the demand of a user for the shared vehicle is greatly met, the efficiency of scheduling the shared vehicle among the parking stations is improved, and the user experience is improved. And the energy conversion of the vehicle is facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle management method and system based on a deep reinforcement learning model, and belongs to the field of artificial intelligence technology. Background Art

[0002] At present, with the increasing popularity of shared vehicle systems and the continued growth of the two-wheeled vehicle market, citizens' enthusiasm for green and low-carbon travel has not diminished, the number of rides continues to increase, and the requirements for shared travel platforms or system service providers are becoming higher and higher, such as using machine learning algorithms to achieve refined operations based on time, location, traffic conditions, etc., to improve user experience.

[0003] In a shared vehicle system, users' borrowing and returning of vehicles is random and affected by dynamic factors such as weather and time, resulting in unbalanced scheduling of shared vehicles. Manual scheduled maintenance is usually used, but this method is not only time-consuming and labor-intensive, but also fails to accurately schedule shared vehicles based on actual user needs. In addition, the usage of vehicles at each station varies. Some stations have fewer vehicles with low battery / hydrogen levels, but manual labor still has to go to the station to replace them, which is inefficient. In addition, citizens' own vehicles occasionally require rapid battery / hydrogen replacement, and how to promptly respond to users' range requirements is also an urgent problem that needs to be solved.

[0004] Therefore, if conventional vehicle management methods are adopted, it is easy for some stations to be short of vehicles. Users may not be able to borrow a vehicle at the station where there is a shortage of vehicles, or they may encounter the embarrassing problem of being able to pick up a vehicle but lacking electricity / hydrogen. The above situations will cause huge economic losses to operators and seriously affect user experience. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide a vehicle management method and system based on a deep reinforcement learning model to achieve accurate scheduling of shared vehicles in each parking station, while meeting the vehicle's energy conversion requirements, meeting the needs of different users, and improving user service quality. In order to solve the above technical problems, the technical solution of the present invention is:

[0006] On the one hand, the present invention provides a vehicle management method based on a deep reinforcement learning model, comprising the following steps:

[0007] Collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area in each time period to generate a training sample set and a prediction sample set; train a neural network model based on deep learning based on the training sample set to obtain a trained neural network model; generate a vehicle demand predictor based on deep reinforcement learning based on the trained neural network model and the prediction sample set; manage vehicles at parking stations in at least one area based on the vehicle predicted demand data generated by the vehicle demand predictor corresponding to each parking station.

[0008] Furthermore, historical dynamic data and actual vehicle demand data corresponding to each parking stop in the area in each time period are collected to generate a training sample set and a prediction sample set, including the following steps: collecting historical dynamic data and actual vehicle demand data corresponding to each parking stop in each time period; dividing the data sample set consisting of the actual vehicle demand data and historical dynamic data corresponding to each time period into a training sample set and a prediction sample set; wherein any data sample in the training sample set and the prediction sample set includes the historical dynamic data corresponding to the previous time period of the parking stop and the actual vehicle demand data corresponding to the next time period.

[0009] Furthermore, according to the training sample set, the neural network model based on deep learning is trained to obtain a trained neural network model, including the following steps: calling a preset neural network model, taking the historical dynamic data of the parking site corresponding to the previous time period of any data sample in the data sample set as the input of the neural network model, and taking the actual vehicle demand data corresponding to the next time period as the output, repeatedly training the neural network model to obtain a trained neural network model.

[0010] Furthermore, a vehicle demand predictor based on deep reinforcement learning is generated based on the trained neural network model and the prediction sample set, including the following steps: establishing a reinforcement learning model based on the prediction sample set and the neural network model; generating a vehicle demand predictor based on the neural network model under the optimal strategy of the reinforcement learning model;

[0011] Furthermore, the vehicle demand forecast data generated by the vehicle demand forecaster corresponding to each parking site includes: vehicle demand forecast data corresponding to the time period after the parking site calculated by the vehicle demand forecaster based on the dynamic data collected corresponding to the time period before the parking site; the management of vehicles at parking sites in at least one area includes scheduling shared vehicles of multiple parking sites in the same area, scheduling shared vehicles between adjacent areas in multiple areas, and performing energy conversion processing on vehicles in any parking site.

[0012] Furthermore, scheduling shared vehicles at multiple parking sites in the same area includes the following steps:

[0013] Based on the collected historical dynamic data of the previous time period corresponding to any parking station in the same area, the vehicle demand predictor corresponding to each parking station calculates the vehicle forecast demand data corresponding to the shared vehicles at the parking station in the next time period; the difference between the vehicle forecast demand data for the next time period corresponding to each parking station and the actual number of shared vehicles displayed at the end of the previous time period is calculated as the shared vehicle demand difference; based on the shared vehicle demand difference corresponding to each parking station and a preset first scheduling threshold, a vehicle scheduling request corresponding to each parking station is generated; wherein the vehicle scheduling request includes a vehicle transfer-in request and a vehicle transfer-out request; based on the vehicle scheduling request corresponding to each parking station, shared vehicles at multiple parking stations in the same area are scheduled.

[0014] Furthermore, based on the shared vehicle demand difference corresponding to each parking stop and a preset first scheduling threshold, a vehicle dispatch request corresponding to each parking stop is generated, including the following steps: when the shared vehicle demand difference is a positive number and its absolute value is not less than the first scheduling threshold, a vehicle transfer-in request corresponding to the parking stop is generated; when the shared vehicle demand difference is a negative number and its absolute value is not less than the first scheduling threshold, a vehicle transfer-out request corresponding to the parking stop is generated; based on the vehicle dispatch request corresponding to each parking stop, shared vehicles at multiple parking stops in the same area are dispatched, including the following steps: when transfer-in requests and transfer-out requests occur simultaneously at multiple parking stops in the same area, the dispatch vehicle is controlled to perform the dispatch task according to the preset vehicle dispatch strategy, so as to transfer the shared vehicle from at least one parking stop that initiated the vehicle transfer-out request, and then transfer the transferred shared vehicle into the parking stop that initiated the vehicle transfer-in request.

[0015] Furthermore, the multiple areas include a first area and a second area, and the second area is any area adjacent to the first area; scheduling shared vehicles between adjacent areas in the multiple areas includes the following steps: when the shared vehicle demand difference corresponding to any parking station in the first area is negative and the absolute value is not less than a preset second scheduling threshold, and the shared vehicle demand difference corresponding to any parking station in the second area is positive and the absolute value is not less than the second scheduling threshold, the scheduling vehicle is controlled to perform the scheduling task according to the vehicle scheduling strategy to dispatch shared vehicles from the first area and then dispatch the dispatched shared vehicles to each parking station in the second area; wherein the second scheduling threshold is not less than the first scheduling threshold.

[0016] Furthermore, in the process of dispatching shared vehicles from parking sites within at least one area, dispatching a shared vehicle from a parking site that initiated a vehicle dispatch request includes the following steps: calculating a total difference M in shared vehicle demand corresponding to all parking sites that initiated the vehicle dispatch request and a total incoming demand N for all parking sites that initiated the vehicle dispatch request; performing a first round of energy storage sorting on multiple shared vehicles from highest to lowest based on the energy storage corresponding to each shared vehicle at all parking sites that initiated the vehicle dispatch request, thereby generating a first energy storage list; when M is greater than N, selecting the shared vehicles ranked last M in the first energy storage list to generate a second energy storage list; selecting the shared vehicles ranked last N in the second energy storage list, switching the energy of the shared vehicles ranked last N and dispatching them, and then dispatching them to the multiple parking sites that initiated the vehicle dispatch request according to the vehicle dispatch strategy; when N is greater than M, switching the energy of the shared vehicles ranked last M in the first energy storage list and dispatching them, and then dispatching them to the multiple parking sites that initiated the vehicle dispatch request according to the vehicle dispatch strategy.

[0017] On the other hand, the present invention provides a vehicle management system, including: a data acquisition module, which is used to collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area in each time period, and generate a training sample set and a prediction sample set; a model training module, which is used to train a neural network model based on deep learning according to the training sample set to obtain a trained neural network model; a model generation module, which is used to generate a vehicle demand predictor based on deep reinforcement learning according to the trained neural network model and the prediction sample set; a vehicle scheduling module, which is used to manage vehicles at parking stations in at least one area according to the vehicle prediction demand data generated by the vehicle demand predictor corresponding to each parking station.

[0018] The vehicle management method based on a deep reinforcement learning model provided by this application uses a vehicle demand predictor based on a deep reinforcement learning model to predict the expected vehicle demand data. Based on the predicted vehicle demand data, the method accurately dispatches shared vehicles between multiple parking sites in the same area and between adjacent areas. It also realizes the energy conversion processing of vehicles arriving at the station, greatly meeting the user's demand for shared vehicles, improving the efficiency of shared vehicle scheduling between parking sites, and facilitating the energy conversion of vehicles at the site. This application can also select shared vehicles with low energy storage at the site and convert their energy, reducing the probability of users picking shared vehicles with low energy storage when they go to the parking site, and ensuring that the shared vehicles transferred to the parking sites that initiate vehicle transfer requests are all shared vehicles with sufficient energy storage, greatly improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the invention and, together with the description, serve to explain the principles of the invention.

[0020] Figure 1 is a flow chart of a vehicle management method based on a deep reinforcement learning model according to one embodiment;

[0021] Figure 2 is a schematic diagram of vehicle scheduling between adjacent areas according to yet another embodiment;

[0022] Figure 3 FIG. 4 is a schematic diagram of a vehicle management system that can be used to implement another embodiment. DETAILED DESCRIPTION

[0023] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments in conjunction with the accompanying drawings.

[0024] <Method Example>

[0025] like Figure 1 As shown, this embodiment provides a vehicle management method 1000 based on a deep reinforcement learning model, including steps S1000 to S4000:

[0026] Step S1000: Collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area in each time period to generate a training sample set and a prediction sample set.

[0027] In this embodiment, historical dynamic data and vehicle demand data for each parking station in each region, corresponding to each time period, are collected to generate training sample sets and prediction sample sets. The historical dynamic data includes dynamic parameters such as light intensity, humidity, temperature, wind speed, visibility, weekend, daytime, nighttime, and holiday information for each time period (specifically, down to the year, month, day, and time of day). Actual vehicle demand data can be the number of vehicles actually borrowed from a parking station by users during each time period. The training sample sets and prediction sample sets are generated based on the historical dynamic data and actual vehicle demand data for each parking station during each time period.

[0028] In one embodiment, step S1000: collecting historical dynamic data and actual vehicle demand data corresponding to each parking station in the area at each time period to generate a training sample set and a prediction sample set, including steps S1010 to S1020:

[0029] Step S1010: Collect historical dynamic data and actual vehicle demand data for each parking station corresponding to each time period.

[0030] In this embodiment, historical dynamic data and actual vehicle demand data for each parking station corresponding to each time period are first collected. The length of each time period can be set to one hour or other time lengths, which are manually set according to actual conditions and are not limited here.

[0031] Collect historical dynamic data for each parking station corresponding to each time period. This historical dynamic data includes a series of dynamic parameters that can be measured and collected in real time using a multi-dimensional data measurement device installed at each parking station. As some parameters (such as temperature, humidity, wind speed, and other variable parameters) change every minute and every second over time, the collection frequency of each parameter can be set to three times per time period, and then the average value of each parameter at each collection time point is taken as the corresponding parameter value for that time period. The collection frequency of each parameter here can be manually set according to actual conditions and is not limited here.

[0032] Step S1020: Divide the data sample set consisting of the actual vehicle demand data and historical dynamic data corresponding to each time period into a training sample set and a prediction sample set; wherein any data sample in the training sample set and the prediction sample set includes the historical dynamic data corresponding to the previous time period of the parking station and the actual vehicle demand data corresponding to the next time period.

[0033] In this embodiment, a data sample set is formed based on data samples consisting of actual vehicle demand data and historical dynamic data corresponding to parking stations in various time periods. It should be noted that each data sample in the data sample set includes historical dynamic data corresponding to the previous time period of the parking station and actual vehicle demand data corresponding to the next time period. For example, a data sample in the data sample set includes historical dynamic data for the time period from 10:00 AM to 11:00 AM on March 6, 2015, and actual vehicle demand data for the time period from 11:00 AM to 12:00 AM on March 6, 2015. On this basis, all data samples in the data sample set are pre-processed by labeling and then divided into a training sample set and a prediction sample set. The ratio of the number of data samples included in the training set and the prediction set can be set to 5:5 or 6:4. The specific ratio can be set manually according to actual conditions and is not limited here.

[0034] Step S2000: Train the deep learning-based neural network model according to the training sample set to obtain a trained neural network model.

[0035] In this embodiment, the neural network model is trained using data samples from a training sample set to repeatedly train the deep learning-based neural network model to obtain a trained neural network model. The trained deep neural network model can predict vehicle demand data for each parking station corresponding to a subsequent period based on dynamic data from the previous period.

[0036] In one embodiment, step S2000: training a neural network model based on deep learning according to a training sample set to obtain a trained neural network model includes step S2010:

[0037] Step S2010: Call the preset neural network model, use the historical dynamic data of the parking station corresponding to the previous time period of any data sample in the data sample set as the input of the neural network model, and use the actual vehicle demand data corresponding to the next time period as the output, repeatedly train the neural network model to obtain a trained neural network model.

[0038] In this embodiment, a preset neural network model is called for training. The historical dynamic data of the parking station in the previous period for each data sample in the training sample set is used as input, and the actual vehicle demand data of the parking station in the subsequent period is used as output for repeated training. Specifically, the training is repeated for each data sample in the data sample set until the optimal state-action-value function is obtained. The trained neural network model is then stopped and saved. This is well understood by those skilled in the art and will not be further elaborated here.

[0039] For example, a data sample in the training sample set includes historical dynamic data from 10:00 AM to 11:00 AM on March 6, 2015, and actual vehicle demand data from 11:00 AM to 12:00 AM on March 6, 2015. During the neural network training process, this embodiment uses the historical dynamic data from 10:00 AM to 11:00 AM on March 6, 2015, as input and the actual vehicle demand data from 11:00 AM to 12:00 AM on March 6, 2015, as output. In other words, this process uses a deep learning supervised learning approach to train the preset neural network model, ultimately producing a trained neural network model.

[0040] Step S3000: Generate a vehicle demand predictor based on deep reinforcement learning based on the trained neural network model and prediction sample set.

[0041] In this embodiment, a reinforcement learning model is established based on the trained neural network model and the data samples included in the prediction sample set. Based on the neural network model under the optimal strategy of the reinforcement learning model, a vehicle demand predictor based on deep reinforcement learning is generated.

[0042] In one embodiment, step S3000: generating a vehicle demand forecaster based on deep reinforcement learning according to the trained neural network model and the prediction sample set, further includes steps S3010 to S3020:

[0043] Step S3010: Establish a reinforcement learning model based on the prediction sample set and the neural network model.

[0044] In this embodiment, a reinforcement learning model is established based on the prediction sample set and the neural network model trained in the previous stage. Then, the trained neural network model is used as the agent of the reinforcement learning model in this stage, and each prediction action of the neural network model is used as the action of the reinforcement learning model. Based on the prediction sample set, the MSE (mean square error) of the vehicle prediction demand data of the next period predicted by the neural network model based on the historical dynamic data of the previous period in each data sample in the prediction sample set and the vehicle actual demand data corresponding to the next period in the data sample is used as the environment of the reinforcement learning model, and the size of the MSE of the vehicle prediction demand data and the vehicle actual demand data is used as the basis for setting the reward of the reinforcement learning model (reward mechanism of reinforcement learning).

[0045] In this embodiment, the preferred algorithm used in the above process is the Q-Learning algorithm, a value-based reinforcement learning algorithm. As the agent's reinforcement learning continues, its actions will increasingly favor the strategy that maximizes the reward (evaluation metric), resulting in a smaller MSE value and more accurate predictions. The goal is to find the strategy that maximizes the reward. The problem can be abstracted into a Markov decision process involving the agent, environment, reward, and action.

[0046] When setting the reward of the reinforcement learning model, the following rules can be followed: when the MSE is in the range [2.5, +∞), the evaluation index is -30; when the MSE is in the range [1.5, 2.5), the evaluation index is -5; when the MSE is in the range [1, 1.5), the evaluation index is -1; when the MSE is in the range [0, 1), the evaluation index is +30. It should be noted that the specific values ​​of the MSE range and evaluation index in this embodiment can be set manually according to actual conditions and are not limited here.

[0047] Step S3020: Generate a vehicle demand predictor based on the neural network model under the optimal strategy of the reinforcement learning model. In this embodiment, a new predictor is generated by obtaining the neural network model under the optimal strategy of the reinforcement learning model as the vehicle demand predictor.

[0048] Step S4000: managing vehicles at parking sites within at least one area based on vehicle demand forecast data generated by a vehicle demand forecaster corresponding to each parking site.

[0049] In this embodiment, a vehicle demand predictor corresponding to each parking station predicts vehicle demand data corresponding to a subsequent time period based on collected historical dynamic data from a previous time period. Vehicles at parking stations within at least one area are managed based on the vehicle demand data corresponding to the subsequent time period for each parking station. Each area may include multiple parking stations.

[0050] In one embodiment, step S4000: based on the vehicle demand forecast data generated by the vehicle demand forecaster corresponding to each parking station, may include step S4010:

[0051] Step S4010: Based on the collected dynamic data corresponding to the time period before the parking stop, the vehicle demand forecaster calculates the vehicle demand data corresponding to the time period after the parking stop.

[0052] The vehicles include shared vehicles and non-shared vehicles, and the management of vehicles at parking sites within at least one area includes scheduling shared vehicles at multiple parking sites within the same area, scheduling shared vehicles between adjacent areas within multiple areas, and performing energy conversion processing on vehicles at any parking site.

[0053] The vehicles in this embodiment may include shared vehicles and non-shared vehicles. Parking sites can park shared vehicles and non-shared vehicles. Managing vehicles in parking sites within at least one area includes scheduling shared vehicles in multiple parking sites, scheduling shared vehicles in adjacent areas of multiple areas, and also includes providing manual or self-service energy conversion services for non-shared vehicles in any parking site. The energy conversion process here may include replacing the vehicle's on-board lithium battery, on-board hydrogen storage device or other type of energy storage device, which is not limited here.

[0054] The advantage of doing this is that the vehicle demand forecaster can predict the expected vehicle demand data, and based on the vehicle demand forecast data, accurate scheduling of shared vehicles between multiple parking sites in the same area and between adjacent areas can be achieved. It also realizes the energy conversion processing of vehicles arriving at the station, which greatly meets the user's demand for shared vehicles, improves the efficiency of shared vehicle scheduling between parking sites, and makes it easier for vehicles to go to the site for energy conversion, greatly improving the user experience.

[0055] In another embodiment, the scheduling of shared vehicles at multiple parking sites in the same area in step S4010 may further include steps S4011 to S4014, which are specifically as follows:

[0056] Step S4011: Based on the collected historical dynamic data of the previous time period corresponding to any parking station in the same area, the vehicle demand predictor corresponding to each parking station calculates the vehicle demand data corresponding to the shared vehicles of the parking station in the next time period.

[0057] In this embodiment, shared vehicles are dispatched for multiple parking stations within the same area. Specifically, based on the collected historical dynamic data for the previous time period corresponding to any parking station within the same area, the vehicle demand forecaster corresponding to each parking station performs a forecast to obtain the predicted vehicle demand data for each parking station in the next time period.

[0058] Step S4012: Calculate the difference between the predicted vehicle demand data for the next period corresponding to each parking stop and the actual number of shared vehicles displayed at the end of the previous period as the shared vehicle demand difference.

[0059] In this embodiment, after calculating the predicted vehicle demand data for each parking station in the next period, the difference between the predicted vehicle demand data and the actual number of shared vehicles corresponding to the parking station at the end of the previous period is calculated as the shared vehicle demand difference. The actual number of shared vehicles is the number of shared vehicles parked at the parking station at the end of the previous period.

[0060] Step S4013: Generate a vehicle dispatch request corresponding to each parking stop based on the shared vehicle demand difference corresponding to each parking stop and a preset first dispatch threshold; wherein the vehicle dispatch request includes a vehicle transfer-in request and a vehicle transfer-out request.

[0061] In this embodiment, a vehicle dispatch request corresponding to each parking stop is generated based on the numerical relationship between the difference in shared vehicle demand corresponding to each parking stop and the first dispatch threshold. It should be noted that the specific value of the first dispatch threshold can be set manually based on actual conditions and is not limited here.

[0062] In some other embodiments, step S4013: generating a vehicle dispatch request corresponding to each parking stop based on the shared vehicle demand difference corresponding to each parking stop and a preset first dispatch threshold value, may further include step S4013A, specifically as follows:

[0063] Step S4013A: When the difference in shared vehicle demand is a positive number and its absolute value is not less than the first scheduling threshold, a vehicle entry request corresponding to the parking site is generated; when the difference in shared vehicle demand is a negative number and its absolute value is not less than the first scheduling threshold, a vehicle exit request corresponding to the parking site is generated.

[0064] In this embodiment, when the shared vehicle demand difference corresponding to a parking stop is a positive number (i.e., the vehicle forecast demand data is greater than the first scheduling threshold, indicating that the vehicle forecast demand in the subsequent time period exceeds the actual number of shared vehicles in the parking stop displayed at the end of the previous time period) and the absolute value is not less than the first scheduling threshold, a vehicle transfer request corresponding to the parking stop is generated.

[0065] When the difference in shared vehicle demand corresponding to a parking stop is zero (i.e., the vehicle forecast demand data is equal to the first scheduling threshold, indicating that the forecast vehicle demand in the subsequent time period is equal to the actual number of shared vehicles in the parking stop displayed at the end of the previous time period), no vehicle scheduling request corresponding to the parking stop is generated.

[0066] When the difference in shared vehicle demand corresponding to a parking stop is negative (i.e., the vehicle forecast demand data is less than the first scheduling threshold, indicating that the vehicle forecast demand in the subsequent time period is less than the actual number of shared vehicles in the parking stop displayed at the end of the previous time period) and the absolute value is not less than the first scheduling threshold, a vehicle dispatch request corresponding to the parking stop is generated.

[0067] In addition, in the two cases where the shared vehicle demand difference corresponding to a parking stop is a negative number but the absolute value is less than the first scheduling threshold, or the shared vehicle demand difference corresponding to a parking stop is a positive number but the absolute value is less than the first scheduling threshold, no vehicle scheduling request corresponding to the parking stop is generated.

[0068] Step S4014: Dispatch shared vehicles at multiple parking sites in the same area according to the vehicle dispatch request corresponding to each parking site.

[0069] In this embodiment, a shared vehicle demand predictor corresponding to each parking station within the same area is used to calculate the shared vehicle demand difference corresponding to each parking station. Based on the relationship between the shared vehicle demand difference corresponding to each parking station and the first scheduling threshold, a vehicle dispatch request corresponding to each parking station is generated. Based on the vehicle dispatch requests corresponding to each parking station within the same area, shared vehicles are dispatched to multiple parking stations within the same area.

[0070] In another embodiment of the present application, another vehicle demand forecaster is provided for performing energy conversion processing on vehicles in any parking station. The energy conversion processing on vehicles in any parking station in step S4010 may further include the following steps:

[0071] In this embodiment, the historical dynamic data and actual vehicle demand data corresponding to each parking station in the area are collected in each time period to generate a training sample set and a prediction sample set corresponding to the vehicle. Among them, the historical dynamic data includes dynamic data such as light intensity, humidity, temperature, wind speed, visibility, whether it is a weekend, daytime or nighttime, and whether it is a holiday in each time period. The actual vehicle demand data is the number of vehicles that went to each parking station for energy conversion in each time period in the past. Based on the training sample set and prediction sample set of the vehicle, the same deep reinforcement learning method as described above is used to generate a vehicle demand predictor based on deep reinforcement learning corresponding to the energy conversion vehicle. Since the deep reinforcement learning method for the energy conversion vehicle here is consistent with the above method, it will not be repeated here. According to the vehicle demand forecast data of the next time period generated by the vehicle demand predictor corresponding to each parking station, the vehicles in any parking station are converted.

[0072] The vehicle demand forecaster at each parking station generates the vehicle demand forecast data for the next period based on the historical dynamic data of the parking station in the previous period. It should be noted that the vehicle demand forecast data in this embodiment refers to the number of vehicles expected to go to each parking station for energy conversion in the next period.

[0073] In this embodiment, based on the vehicle demand forecast data for the next period generated by the vehicle demand forecaster corresponding to each parking stop, the solution for implementing energy conversion processing for vehicles going to each parking stop is: a self-service energy conversion exchange cabinet is set up at each parking stop, and lithium batteries, hydrogen storage devices or other energy storage devices are provided, etc., for users to self-service energy conversion. In addition, the back-end can also assign operation and maintenance personnel to carry full-capacity energy storage devices to the corresponding parking stop to provide manual energy conversion services for users, or assign operation and maintenance personnel to maintain the self-service energy conversion exchange cabinet in the station in advance, manually take out the low-capacity energy storage device replaced by the user in the cabinet, and replace it with a full-capacity energy storage device for the user to come for energy conversion.

[0074] Specifically, if a parking lot's vehicle demand forecaster, based on historical dynamic data from the previous period, predicts that the predicted vehicle demand for the next period exceeds the number of energy storage devices available for energy conversion within the parking lot's energy conversion self-service exchange cabinet, this indicates that a large number of vehicles are expected to visit the parking lot for energy conversion in the next period, and the energy storage devices available for energy conversion within the energy conversion self-service exchange cabinet are expected to be in short supply. In this case, the difference between the predicted vehicle demand for the next period and the number of energy storage devices available for energy conversion within the energy conversion self-service exchange cabinet can be calculated to enable operations and maintenance personnel to bring the same number of energy storage devices as the difference to the parking lot in advance, and manually convert the energy for the vehicles visiting the parking lot to meet the vehicle energy conversion needs in the next period.

[0075] In some other embodiments, step S4014 dispatches shared vehicles to multiple parking sites in the same area according to the vehicle dispatch request corresponding to each parking site, including step S4014A:

[0076] Step S4014A: When transfer-in requests and transfer-out requests occur simultaneously at multiple parking stops in the same area, the dispatch vehicle is controlled to perform the dispatch task according to the preset vehicle dispatch strategy, so as to dispatch a shared vehicle from at least one parking stop that initiated the vehicle transfer-out request, and then dispatch the dispatched shared vehicle to the parking stop that initiated the vehicle transfer-in request.

[0077] In this embodiment, vehicle dispatch is only executed if at least one parking station in the same area issues a vehicle dispatch request and at least one parking station in the same area issues a vehicle dispatch request. This allows a vehicle to be dispatched from the parking station that initiated the dispatch request and then dispatched to the parking station that initiated the dispatch request. In other words, if at least one parking station in the same area issues a vehicle dispatch request but no parking station in the same area issues a vehicle dispatch request, vehicle dispatch is not executed.

[0078] For example, assume the first scheduling threshold is 5. Figure 2 The first area shown is a circle with a diameter of 2 kilometers, and includes five parking sites, namely parking sites A, B, C, D and E.

[0079] If the actual number of shared vehicles displayed at parking station A at the end of the previous time period is 50, and the vehicle demand predictor corresponding to parking station A predicts 30 vehicles for the next time period based on the historical dynamic data of the previous time period, then the calculated shared vehicle demand difference of parking station A is negative 20, that is, the number of vehicles that can be dispatched from parking station A in the next time period is 20, and the absolute value of the shared vehicle demand difference of parking station A is greater than 5, and a vehicle dispatch request corresponding to parking station A is generated and issued.

[0080] If the actual number of shared vehicles displayed at parking station B at the end of the previous time period is 20, and the vehicle demand predictor corresponding to parking station B predicts that the vehicle demand data for the next time period is 26 based on the historical dynamic data of the previous time period, then the calculated shared vehicle demand difference of parking station B is positive 6, that is, the number of vehicles that need to be transferred into parking station B in the next time period is 6, and the absolute value of the shared vehicle demand difference of parking station B is greater than 5, and a vehicle transfer request corresponding to the parking station is generated and issued.

[0081] Example 1: If parking lot A issues a vehicle dispatch request and has 20 vehicles available for dispatch in the next period; parking lot B issues a vehicle dispatch request and needs to dispatch 6 vehicles in the next period; parking lot C issues a vehicle dispatch request and needs to dispatch 6 vehicles in the next period; parking lot D issues a vehicle dispatch request and needs to dispatch 6 vehicles in the next period; and parking lot E does not issue any vehicle dispatch requests. Based on this, the vehicle dispatch policy in this example is to dispatch 18 vehicles from parking lot A, then dispatch 6 of the dispatched vehicles to parking lot B, 6 to parking lot C, and 6 to parking lot D.

[0082] Example 2: If parking lot A issues a vehicle dispatch request and the number of vehicles available for dispatch in the next period is 5; if parking lot A issues a vehicle dispatch request and the number of vehicles available for dispatch in the next period is 7; parking lot C issues a vehicle dispatch request and the number of vehicles required for dispatch in the next period is 6; parking lot D issues a vehicle dispatch request and the number of vehicles required for dispatch in the next period is 12; and parking lot E issues a vehicle dispatch request and the number of vehicles required for dispatch in the next period is 6. Clearly, the total number of vehicles dispatchable by parking lots A and B cannot fully meet the vehicle dispatch needs of the other three parking lots. Based on this, the vehicle dispatch plan can be: Since the number of vehicles required for dispatch in the next period corresponding to parking lots C, D, and E is in a ratio of 1:2:1, 5 vehicles can be dispatched from parking lot A and 7 vehicles from parking lot B. Then, 3 of the dispatched vehicles can be dispatched to parking lot C, 6 to parking lot D, and 3 to parking lot E. That is, the vehicle dispatching strategy in this example may be to allocate the number of dispatched vehicles according to the ratio of the number of vehicles to be dispatched at each parking station waiting for dispatched vehicles, thereby coordinating the dispatching of vehicles at each station.

[0083] The step S4010 of scheduling shared vehicles between adjacent areas in the plurality of areas further includes step S4015:

[0084] Step S4015: When the difference in shared vehicle demand corresponding to any parking station in the first area is negative and the absolute value is not less than the preset second scheduling threshold, and the difference in shared vehicle demand corresponding to any parking station in the second area is positive and the absolute value is not less than the second scheduling threshold, the scheduling vehicle is controlled to perform the scheduling task according to the vehicle scheduling strategy to dispatch shared vehicles from the first area and then dispatch the dispatched shared vehicles to each parking station in the second area; wherein the second scheduling threshold is not less than the first scheduling threshold.

[0085] In this embodiment, the multiple areas include a first area and a second area, and the second area can be any area adjacent to the first area. The adjacent area in this embodiment is defined as a distance between the center lines of the two areas that is no greater than a preset distance. The preset distance here can be set manually according to actual conditions and is not limited here. Figure 2 As shown in the figure, assuming the preset distance is 4 kilometers, the first and second areas are circles with a diameter of 2 kilometers. The center of the second area is 3 kilometers away from the center of the first area. Therefore, the first and second areas are adjacent areas, and vehicle scheduling can be performed between the two areas. The first area includes parking lots A, B, C, D, and E, and the second area includes F, G, H, J, and K.

[0086] Based on this, when the difference in shared vehicle demand corresponding to any parking station in the first area is negative and its absolute value is not less than the preset second scheduling threshold, and the difference in shared vehicle demand corresponding to any parking station in the second area is positive and its absolute value is not less than the second scheduling threshold, the dispatch vehicle is controlled to perform the vehicle scheduling task according to the above-mentioned vehicle scheduling strategy to dispatch shared vehicles from each parking station in the first area and then dispatch the dispatched shared vehicles to each parking station in the second area. Specifically, the allocation method of shared vehicles dispatched to each parking station in the second area in this embodiment can refer to the vehicle scheduling strategy of "distributing dispatched vehicles based on the ratio of the number of vehicles to be dispatched to each parking station waiting for dispatched vehicles" in step S4014A above, and will not be repeated here. In this embodiment, the second scheduling threshold is not less than the first scheduling threshold. The specific value of the second scheduling threshold can be manually set according to actual conditions and is not limited here.

[0087] In some other embodiments of the present application, managing vehicles at parking sites within at least one area in step S4010 includes scheduling shared vehicles at multiple parking sites within the same area and / or scheduling shared vehicles between adjacent areas within multiple areas. In this embodiment, regardless of whether shared vehicles are scheduled at multiple parking sites within the same area or between adjacent areas within multiple areas, in both scheduling processes, the step of retrieving a shared vehicle from the parking site that initiated the vehicle retrieving request may specifically include the following steps:

[0088] The total difference in shared vehicle demand M corresponding to all parking sites that initiated vehicle transfer-out requests and the total incoming demand N of all parking sites that initiated vehicle transfer-in requests are calculated. Here, M and N are integers not less than zero.

[0089] In this embodiment, in the process of managing vehicles at parking sites within at least one area, including scheduling shared vehicles at multiple parking sites within the same area and / or scheduling shared vehicles between adjacent areas within multiple areas, the sum M of the shared vehicle demand differences corresponding to all parking sites that initiate vehicle transfer-out requests and the total amount N of vehicle transfer-in demands of all parking sites that initiate vehicle transfer-in requests (i.e., the sum of the shared vehicle demand differences of all parking sites that initiate vehicle transfer-in requests) are first calculated.

[0090] According to the energy storage capacity corresponding to each shared vehicle in all parking sites that initiate vehicle dispatch requests, multiple shared vehicles are sorted from high to low in the first round of energy storage to generate a first energy storage list.

[0091] In this embodiment, the internal energy storage capacity of each shared vehicle in all parking sites that initiate the vehicle dispatch request is different. According to the corresponding energy storage capacity of each shared vehicle in all parking sites, the shared vehicles are sorted from high to low to generate a first energy storage list.

[0092] When M is greater than N, the shared vehicles ranked last M in the first energy storage list are selected to generate a second energy storage list. The shared vehicles ranked last N in the second energy storage list are selected, and the shared vehicles in the last N positions are transferred out after energy conversion. Then, according to the vehicle scheduling strategy, they are transferred to the multiple parking sites that initiated the vehicle transfer request.

[0093] In this embodiment, when M is greater than N, the shared vehicles ranked at the last M in the first energy storage list are first preliminarily screened out to generate a second energy storage list. During this preliminary screening process, the background system can also abandon abnormal vehicles in the first energy storage list (such as vehicles that have frequently reported fault codes recently) based on the stored vehicle fault records, and instead select vehicles in other positions in the first energy storage list to enter the second energy storage list to fill the position. For example, assuming that M is 10, the shared vehicle ranked 10th from the bottom in the first energy storage list was originally screened out to enter the second energy storage list, but the shared vehicle ranked 10th from the bottom has recently frequently reported fault codes. Although it does not affect riding at present, it does not rule out the possibility of riding failures in the future. In this case, the vehicle ranked 11th from the bottom can be selected to replace it and enter the second energy storage list. This can ensure the effectiveness of this vehicle dispatch and avoid the dispatched vehicle from failing later and affecting the user experience.

[0094] Then, the last N shared vehicles in the second energy storage list are screened out and manually recharged. The recharged shared vehicles are then dispatched out according to the vehicle dispatch strategy described in Example 1 in step S4014A above and then recharged into the multiple parking lots that initiated the vehicle recharge request. Manual recharge can be performed by an operation and maintenance personnel carrying an energy storage device to the parking lots corresponding to the last N shared vehicles in the screen out and recharging the corresponding shared vehicles.

[0095] When N is greater than M, the shared vehicles ranked last M in the first energy storage list are switched out after energy conversion, and then transferred to the multiple parking sites that initiated the vehicle transfer request according to the vehicle scheduling strategy.

[0096] In this embodiment, when N is greater than M, the total number of available shared vehicles is clearly insufficient to fully meet the vehicle transfer requirements of all parking stations. Therefore, this embodiment can directly transfer the energy of the last M shared vehicles in the first energy storage list and then transfer them to the multiple parking stations that initiated the vehicle transfer request according to the vehicle scheduling strategy described in Example 2 in step S4014A above.

[0097] The advantage of doing so in the embodiment of the present application is that the shared bicycles with lower energy storage capacity in the original parking station are transferred out from the parking station that initiated the vehicle transfer request. Not only can the shared vehicles with lower energy storage capacity be selected and converted in time, reducing the probability of users picking up shared vehicles with lower energy storage capacity when going to the parking station, but after the conversion, it can also be ensured that the shared vehicles transferred into the parking station that initiated the vehicle transfer request are all vehicles with sufficient energy storage capacity, thereby ensuring the subsequent user experience.

[0098] <System Example>

[0099] In one embodiment of the present application, a vehicle management system 2000 is also provided. Figure 3 As shown, the vehicle management system 2000 includes a data acquisition module 2100, a model training module 2200, a model generation module 2300 and a vehicle scheduling module 2400. Among them:

[0100] The data collection module 2100 is used to collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area at each time period, and generate training sample sets and prediction sample sets;

[0101] The model training module 2200 is used to train the neural network model based on deep learning according to the training sample set to obtain a trained neural network model;

[0102] A model generation module 2300 is used to generate a vehicle demand predictor based on deep reinforcement learning based on the trained neural network model and the prediction sample set;

[0103] The vehicle dispatch module 2400 is configured to manage vehicles at parking sites within at least one area based on vehicle demand forecast data generated by a vehicle demand forecaster corresponding to each parking site.

[0104] According to the vehicle management system provided in the embodiment of the present application, a vehicle demand predictor based on a deep reinforcement learning model is established, and the vehicle expected demand data is predicted by the vehicle demand predictor. Based on the vehicle expected demand data, accurate scheduling of shared vehicles between multiple parking sites in the same area and between adjacent areas is achieved, and energy conversion processing of vehicles arriving at the station is also achieved, which greatly meets the user's demand for shared vehicles, improves the efficiency of shared vehicle scheduling between parking sites, and facilitates private vehicles to go to the site for energy conversion, greatly improving the user experience.

[0105] The above specific embodiments further illustrate the technical problems, technical solutions and beneficial effects solved by the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A vehicle management method based on a deep reinforcement learning model, characterized in that: include: Collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area at each time period to generate training sample sets and prediction sample sets; Training a deep learning-based neural network model according to the training sample set to obtain a trained neural network model; Generating a vehicle demand predictor based on deep reinforcement learning according to the trained neural network model and the prediction sample set; The vehicles at the parking sites in at least one area are managed according to the vehicle demand forecast data generated by the vehicle demand forecaster corresponding to each parking site.

2. The vehicle management method according to claim 1, characterized in that: The process of collecting historical dynamic data and actual vehicle demand data corresponding to each parking station in the area at each time period to generate a training sample set and a prediction sample set includes the following steps: Collect historical dynamic data and actual vehicle demand data for each parking station corresponding to each time period; The data sample set consisting of actual vehicle demand data and historical dynamic data corresponding to each time period is divided into a training sample set and a prediction sample set; wherein any data sample in the training sample set and the prediction sample set includes the historical dynamic data corresponding to the previous time period of the parking station and the actual vehicle demand data corresponding to the next time period.

3. The vehicle management method according to claim 2, characterized in that: The step of training a neural network model based on deep learning according to the training sample set to obtain a trained neural network model comprises the following steps: The preset neural network model is called, and the historical dynamic data of the parking site corresponding to the previous time period of any data sample in the data sample set is used as the input of the neural network model, and the actual vehicle demand data corresponding to the next time period is used as the output. The neural network model is repeatedly trained to obtain the trained neural network model.

4. The vehicle management method according to claim 3, characterized in that: Generating a vehicle demand predictor based on deep reinforcement learning according to the trained neural network model and the prediction sample set includes the following steps: Establishing a reinforcement learning model based on the prediction sample set and the neural network model; The vehicle demand predictor is generated according to the neural network model under the optimal strategy of the reinforcement learning model.

5. The vehicle management method according to claim 4, characterized in that: The vehicle demand forecast data generated by the vehicle demand forecaster corresponding to each parking site includes: The vehicle demand forecaster calculates the vehicle demand data corresponding to the next period after the parking station based on the collected dynamic data corresponding to the previous period of the parking station; The management of the vehicles at the parking sites in at least one area includes scheduling shared vehicles at multiple parking sites in the same area, scheduling shared vehicles between adjacent areas in multiple areas, and performing energy conversion processing on vehicles at any parking site.

6. The vehicle management method according to claim 5, characterized in that: The method of dispatching shared vehicles at multiple parking sites in the same area includes the following steps: The vehicle demand forecaster corresponding to each parking station calculates the predicted vehicle demand data corresponding to the shared vehicles at the parking station in a subsequent time period based on the collected historical dynamic data of the previous time period corresponding to any parking station in the same area; Calculating the difference between the predicted vehicle demand data for the next period corresponding to each parking stop and the actual number of shared vehicles displayed at the end of the previous period as the shared vehicle demand difference; Generating a vehicle dispatch request corresponding to each parking stop based on the shared vehicle demand difference corresponding to each parking stop and a preset first dispatch threshold; wherein the vehicle dispatch request includes a vehicle transfer-in request and a vehicle transfer-out request; According to the vehicle dispatching request corresponding to each parking station, shared vehicles at multiple parking stations in the same area are dispatched.

7. The vehicle management method based on deep reinforcement learning model according to claim 6 is characterized in that: Generating a vehicle dispatch request corresponding to each parking stop based on the shared vehicle demand difference corresponding to each parking stop and a preset first dispatch threshold comprises the following steps: generating the vehicle transfer request corresponding to the parking site when the shared vehicle demand difference is a positive number and the absolute value is not less than the first scheduling threshold; generating the vehicle dispatch request corresponding to the parking station when the shared vehicle demand difference is a negative number and the absolute value thereof is not less than the first scheduling threshold; The method of dispatching shared vehicles at multiple parking sites in the same area according to the vehicle dispatch request corresponding to each parking site includes the following steps: When transfer-in requests and transfer-out requests occur simultaneously at multiple parking stops in the same area, the dispatch vehicle is controlled to perform the dispatch task according to the preset vehicle dispatch strategy, so as to dispatch a shared vehicle from at least one parking stop that initiated the vehicle transfer-out request, and then dispatch the dispatched shared vehicle to the parking stop that initiated the vehicle transfer-in request.

8. The vehicle management method based on deep reinforcement learning model according to claim 7 is characterized in that: The plurality of regions include a first region and a second region, wherein the second region is any region adjacent to the first region; The method of scheduling shared vehicles between adjacent areas within a plurality of areas includes the following steps: If the shared vehicle demand differences corresponding to any parking station in the first area are all negative numbers and the absolute value is not less than a preset second scheduling threshold, and the shared vehicle demand differences corresponding to any parking station in the second area are all positive numbers and the absolute value is not less than the second scheduling threshold, the scheduling vehicle is controlled to perform the scheduling task according to the vehicle scheduling strategy to dispatch shared vehicles from the first area and then dispatch the dispatched shared vehicles to each parking station in the second area; wherein the second scheduling threshold is not less than the first scheduling threshold.

9. The vehicle management method based on deep reinforcement learning model according to claim 8 is characterized in that: In the process of dispatching the shared vehicles at the parking sites in at least one area, dispatching the shared vehicle from the parking site that initiated the vehicle dispatch request comprises the following steps: Calculate the total difference M of shared vehicle demand corresponding to all parking sites that initiate the vehicle transfer-out request and the total incoming demand N of all parking sites that initiate the vehicle transfer-in request; According to the energy storage corresponding to each shared vehicle in all parking spots that initiated the vehicle dispatch request, a first round of energy storage sorting is performed on the multiple shared vehicles from high to low, to generate a first energy storage list; When M is greater than N, the shared vehicles ranked last M in the first energy storage list are screened out to generate a second energy storage list; the shared vehicles ranked last N in the second energy storage list are screened out, the shared vehicles ranked last N are switched for energy and then transferred out, and then transferred into the multiple parking sites that initiated the vehicle transfer request according to the vehicle scheduling strategy; When N is greater than M, the shared vehicles ranked last M in the first energy storage list are switched out after energy conversion, and then transferred to the multiple parking sites that initiated the vehicle transfer request according to the vehicle scheduling strategy.

10. A vehicle management system, characterized in that: include: The data collection module is used to collect historical dynamic data and actual vehicle demand data corresponding to each parking station in the area at each time period, and generate training sample sets and prediction sample sets; A model training module is used to train a neural network model based on deep learning according to the training sample set to obtain a trained neural network model; A model generation module, configured to generate a vehicle demand predictor based on deep reinforcement learning according to the trained neural network model and the prediction sample set; The vehicle dispatching module is used to manage the vehicles at the parking sites in at least one area according to the vehicle demand forecast data generated by the vehicle demand forecaster corresponding to each parking site.