Path planning method and device for transport vehicle, vehicle and storage medium
By constructing and optimizing the path planning problem model in the transportation vehicle path planning, the problem of insufficient consideration of static geographic information and dynamic cost in the existing technology is solved, and the efficiency and cost balance of transportation paths are achieved, which is convenient for promotion and application.
Patent Information
- Application Number
- CN202510305042.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
In the prior art, only static geographical information is considered to lead to limited transportation efficiency and lack dynamic considerations on transportation costs, resulting in high overall transportation costs and inconvenient for promotion and application.
By pre-constructing a path planning problem model and optimizing and adjusting the model every time the target transport vehicle performs a transportation task, the optimized path planning problem model outputs a planning path that meets the preset operating cost conditions, and comprehensively considers the task objectives and dynamic costs.
It realizes path planning based on dynamic data under the current transportation operating conditions, balancing efficiency and cost, reducing overall transportation costs, and making it easier to promote and apply.
Smart Images

Figure CN120146352A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of electronic digital data processing, and particularly relates to a path planning method, device, vehicle and storage medium for a transport vehicle. Background Art
[0002] In the related art, the path planning for transport vehicles mainly focuses on transportation tasks. Aiming to ensure sufficient energy, based on the static geographical information between transportation tasks, such as distance, traffic conditions, etc., an efficient planned path is generated.
[0003] However, in the related art, the improvement of transportation efficiency by only considering static geographical information is limited, and due to the lack of dynamic consideration of transportation costs, the overall transportation cost is relatively high, and there are high limitations in actual popularization and use, which needs to be improved. Summary of the Invention
[0004] This application provides a path planning method, device, vehicle and storage medium for a transport vehicle, so as to solve the technical problems in the related art that the improvement of transportation efficiency by only considering static geographical information is limited, and there is a lack of dynamic consideration of transportation costs, resulting in a relatively high overall transportation cost and inconvenience in popularization and application.
[0005] In a first aspect embodiment of this application, a path planning method for a transport vehicle is provided, including the following steps: obtaining the current transportation task of the target transport vehicle; obtaining the current position information, current transportation itinerary and current state data of the target transport vehicle under the current transportation working condition based on the current transportation task; inputting the current position information, current transportation itinerary and current state data into a path planning problem model updated by the previous transportation task to output a planned path that meets the preset operating cost condition, where the path planning problem model is accumulated by a reward function corresponding to the historical transportation tasks of the target transport vehicle.
[0006] Optionally, in an embodiment of this application, the step of inputting the current position information, current transportation itinerary and current state data into a path planning problem model updated by the previous transportation task to output a planned path that meets the preset operating cost condition includes: constructing a state space of the path planning problem model by using at least one of the current position information, power data in the current state data, node data in the current transportation itinerary and current electricity price data; constructing an action space of the path planning problem model by using at least one of the data of the next target node in the current transportation itinerary and the itinerary target; and cumulatively updating the reward function of the path planning problem model by using at least one of the estimated transportation cost, estimated charging cost and estimated power consumption data for traveling to the next node or itinerary target in the current transportation itinerary.
[0007] Optionally, in an embodiment of the present application, after outputting a planned path that meets the preset operating cost condition, it further includes: controlling the target transport vehicle to complete the corresponding transport task based on the planned path; obtaining the satisfaction evaluation, actual energy consumption cost, and actual charging cost of the transport task; and optimizing the path planning problem model by using the satisfaction evaluation, the actual energy consumption cost, and the actual charging cost, so as to generate a new planned path for the target transport vehicle by using the optimized path planning problem model.
[0008] Optionally, in an embodiment of the present application, the expression of the optimized path planning problem model is:
[0009]
[0010] where, π new is the optimized path planning problem model, π old is the path planning problem model, R is the cumulative updated reward function, and β is the KL divergence hyperparameter balance coefficient between the optimization objective and the optimized path planning problem model and the path planning problem model.
[0011] Optionally, in an embodiment of the present application, the step of inputting the current position information, current transport itinerary, and current status data into the path planning problem model updated by the previous transport task to output a planned path that meets the preset operating cost condition includes: integrating the road network data involved in the current transport itinerary and the current status data to generate a corresponding observation vector; and generating a probability distribution for selecting a transport node based on the observation vector, so as to determine the next target node based on the probability distribution.
[0012] Optionally, in an embodiment of the present application, the step of inputting the current position information, current transport itinerary, and current status data into the path planning problem model updated by the previous transport task to output a planned path that meets the preset operating cost condition includes: obtaining the estimated charging cost based on at least one charging station data involved in the current transport itinerary and the current status data.
[0013] A path planning device for a transport vehicle according to an embodiment of the second aspect of the present application includes: a first acquisition module configured to acquire the current transport task of a target transport vehicle; a second acquisition module configured to acquire the current position information, current transport itinerary, and current state data of the target transport vehicle under the current transport condition based on the current transport task; a planning module configured to input the current position information, current transport itinerary, and current state data into a path planning problem model updated by the previous transport task to output a planned path that meets a preset operating cost condition, wherein the path planning problem model is obtained by cumulatively calculating a reward function corresponding to the historical transport tasks of the target transport vehicle.
[0014] Optionally, in an embodiment of the present application, the planning module includes: a first construction unit configured to construct a state space of the path planning problem model by using at least one of the current position information, power data in the current state data, node data in the current transport itinerary, and current electricity price data; a second construction unit configured to construct an action space of the path planning problem model by using at least one of data of the next target node in the current transport itinerary and itinerary objectives; a third construction unit configured to cumulatively update the reward function of the path planning problem model by using at least one of the estimated transport cost, estimated charging cost, and estimated power consumption data for traveling to the next node or itinerary objective in the current transport itinerary.
[0015] Optionally, in an embodiment of the present application, the path planning device for a transport vehicle further includes: a control module configured to control the target transport vehicle to complete the corresponding transport task based on the planned path; a third acquisition module configured to acquire a satisfaction evaluation, actual energy consumption cost, and actual charging cost of the transport task; an optimization module configured to optimize the path planning problem model by using the satisfaction evaluation, the actual energy consumption cost, and the actual charging cost, so as to generate a new planned path for the target transport vehicle by using the optimized path planning problem model.
[0016] Optionally, in an embodiment of the present application, the expression of the optimized path planning problem model is:
[0017]
[0018] wherein, π new is the optimized path planning problem model, π old is the path planning problem model, R is the cumulatively updated reward function, and β is a KL divergence hyperparameter balance coefficient between the optimization objective and the optimized path planning problem model and the path planning problem model.
[0019] Optionally, in an embodiment of the present application, the planning module includes: a generating unit configured to integrate the road network data and the current state data involved in the current transportation trip to generate a corresponding observation vector; a determining unit configured to generate a probability distribution for the selection of transportation nodes based on the observation vector, and determine the next target node based on the probability distribution.
[0020] Optionally, in an embodiment of the present application, the planning module includes: a calculating unit configured to obtain the estimated charging cost based on at least one charging station data and the current state data involved in the current transportation trip.
[0021] An embodiment of the third aspect of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the program to implement the path planning method for a transportation vehicle as described in the above embodiments.
[0022] An embodiment of the fourth aspect of the present application provides a computer-readable storage medium storing computer instructions for causing a computer to execute the path planning method for a transportation vehicle as described in the above embodiments.
[0023] An embodiment of the fifth aspect of the present application provides a computer program product including a computer program, which when executed, is used to implement the path planning method for a transportation vehicle as described above.
[0024] Embodiments of the present application can pre-construct a path planning problem model and optimize and adjust the path planning problem model each time a target transportation vehicle performs a transportation task. Then, after obtaining the current position information, the current transportation trip, and the current state data of the target transportation vehicle under the current transportation condition, the optimized path planning problem model is used to output a planned path that meets the preset operating cost conditions, so as to perform path planning according to the dynamic data under the current transportation condition, comprehensively consider the task objective and the dynamic cost, and obtain a transportation path that balances efficiency and cost, which is convenient for popularization and application. Thereby, it solves the technical problems in the related art that the transportation efficiency improved by only considering static geographical information is limited, and there is a lack of dynamic consideration of transportation costs, resulting in a relatively high overall transportation cost and inconvenience for popularization and application.
[0025] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0027] Figure 1 It is a flowchart of a path planning method for a transport vehicle provided according to an embodiment of the present application;
[0028] Figure 2 It is a schematic diagram of the principle of a path planning method for a transport vehicle provided according to an embodiment of the present application;
[0029] Figure 3 It is a schematic structural diagram of a path planning device for a transport vehicle provided according to an embodiment of the present application;
[0030] Figure 4 It is a schematic structural diagram of an electronic device provided according to an embodiment of the present application. Specific embodiments
[0031] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0032] The path planning method, device, electronic device, and storage medium of the transport vehicle according to the embodiments of the present application will be described below with reference to the accompanying drawings. In view of the technical problem in the related art mentioned in the above background art that only static geographical information is considered, the improved transport efficiency is limited, and there is a lack of dynamic consideration of the transport cost, resulting in a relatively high overall transport cost and inconvenience for popularization and application, the present application provides a path planning method for a transport vehicle. In this method, a path planning problem model can be pre-constructed, and the path planning problem model can be optimized and adjusted each time the target transport vehicle performs a transport task. Then, after obtaining the current position information, current transport itinerary, and current state data of the target transport vehicle under the current transport condition, the optimized path planning problem model is used to output a planned path that meets the preset operating cost conditions, so as to perform path planning based on the dynamic data under the current transport condition, comprehensively consider the task objectives and dynamic costs, and obtain a transport path that balances efficiency and cost, which is convenient for popularization and application. Thus, the technical problem in the related art that only static geographical information is considered, the improved transport efficiency is limited, and there is a lack of dynamic consideration of the transport cost, resulting in a relatively high overall transport cost and inconvenience for popularization and application is solved.
[0033] Specifically, Figure 1 It is a schematic flow diagram of a path planning method for a transport vehicle provided by an embodiment of the present application.
[0034] As Figure 1 shown, the path planning method for the transport vehicle includes the following steps:
[0035] In step S101, obtain the current transportation task of the target transportation vehicle.
[0036] In the actual execution process, embodiments of the present application can obtain the current transportation task of the target transportation vehicle through methods such as schedule arrangement and terminal reading, including all transportation nodes in the current transportation task, the location of each transportation node, the loading or unloading task of each transportation node, etc. Based on the current transportation task, embodiments of the present application can determine data such as the task progress in combination with the actual transportation working conditions of the vehicle, so as to perform path planning subsequently.
[0037] In step S102, based on the current transportation task, obtain the current position information, current transportation itinerary, and current status data of the target transportation vehicle under the current transportation working conditions.
[0038] In the actual execution process, embodiments of the present application can collect the current position information, current transportation itinerary (such as the overall transportation task, completed nodes, uncompleted nodes, next transportation target, etc.), and current status data (such as remaining power, current load, current power consumption rate, etc.) of the target transportation vehicle under the current transportation working conditions according to the current transportation task, so as to input the data under the current transportation working conditions into the model to achieve path planning.
[0039] In step S103, input the current position information, current transportation itinerary, and current status data into the path planning problem model updated by the previous transportation task, so as to output a planned path that meets the preset operating cost conditions, where the path planning problem model is obtained by accumulating the reward functions corresponding to the historical transportation tasks of the target transportation vehicle.
[0040] It can be understood that the solution involved in embodiments of the present application is a dynamically optimized model.
[0041] The control strategy of embodiments of the present application itself is a neural network. Using its non-linear mapping ability, it can fit any function. Therefore, in the initialization model, only by setting random parameters can a preliminary model be obtained. Then, through the method of reinforcement learning, the parameters of the neural network are iteratively optimized to finally approximate the optimal strategy.
[0042] Therefore, the path planning problem model used in embodiments of the present application is a path planning problem model obtained by continuously optimizing the target transportation vehicle through multiple transportation tasks.
[0043] Embodiments of the present application can input the obtained data into the path planning problem model updated by the previous transportation task, thereby updating the state space, action space, and reward function of the path planning problem model according to the current transportation working conditions, so as to output a planned path that minimizes the operating cost.
[0044] Optionally, in an embodiment of the present application, the current location information, the current transportation itinerary, and the current status data are input into the path planning problem model updated by the previous transportation task to output a planned path that meets the preset operating cost conditions, including: constructing the state space of the path planning problem model by using at least one of the current location information, the power data in the current status data, the node data in the current transportation itinerary, and the current electricity price data; constructing the action space of the path planning problem model by using at least one of the data of the next target node in the current transportation itinerary and the itinerary target; and cumulatively updating the reward function of the path planning problem model by using at least one of the estimated transportation cost, the estimated charging cost, and the estimated power consumption data for traveling to the next node or the itinerary target in the current transportation itinerary.
[0045] As a possible implementation manner, the embodiment of the present application can construct a path planning problem model, which can consider the electricity price differences of charging stations in different regions at different time periods.
[0046] The embodiment of the present application can transform the path planning problem model into a constrained Markov decision process, where the state space can include the current location of the target transport vehicle, the remaining power, the visited customer nodes and charging station information, and the electricity price information at the current moment, etc.; the action space can include traveling to the next target node, the itinerary target, such as traveling to the charging station to charge or returning to the warehouse; the reward function considers factors such as transportation cost, charging cost, and power consumption, and is adjusted according to the dynamic electricity price.
[0047] Optionally, in an embodiment of the present application, the current location information, the current transportation itinerary, and the current status data are input into the path planning problem model updated by the previous transportation task to output a planned path that meets the preset operating cost conditions, including: obtaining the estimated charging cost based on at least one charging station data involved in the current transportation itinerary and the current status data.
[0048] In some embodiments, the embodiment of the present application can design a policy network based on LSTM-Transformer for predicting future electricity prices and generating the optimal path planning of the target transport vehicle; among them, the LSTM network is used to capture the time series relationship of electricity prices, and the Transformer network is used to capture the complex coupling relationship between nodes in the path planning; the LSTM network can also include an encoder module to encode the road network state and integrate it with the current status data of the target transport vehicle.
[0049] Furthermore, the embodiment of the present application can use the generator module to predict the electricity price when the target transport vehicle arrives at the charging station, and consider the electricity price change trend and the remaining power of the electric vehicle to optimize the charging decision.
[0050] Among them, the embodiments of the present application can be trained through machine learning algorithms using historical electricity price data, current state data of the target transport vehicle, current transport itinerary, etc., to predict the future change trend of electricity prices, so as to estimate the charging cost.
[0051] Optionally, in an embodiment of the present application, the current location information, current transport itinerary, and current state data are input into the path planning problem model updated by the previous transport task to output a planned path that meets the preset operating cost conditions, including: integrating the road network data and current state data involved in the current transport itinerary to generate a corresponding observation vector; generating a probability distribution for the selection of transport nodes based on the observation vector to determine the next target node based on the probability distribution.
[0052] In some other embodiments, the embodiments of the present application can integrate the road network data and current state data involved in the current transport itinerary of the target transport vehicle through an embedding module to generate a final observation vector, and input it into the decoder module; the decoder module generates a probability distribution for node selection according to the observation vector, so as to determine the next target node.
[0053] Among them, the embedding module can be implemented by a fully connected network, which is used to integrate the states of different attributes and output the final observation vector.
[0054] After the above process, the embodiments of the present application can, through self-training, enable the policy network to learn to select reasonable charging times and paths according to the electricity price fluctuation trend to minimize the operating cost.
[0055] Optionally, in an embodiment of the present application, after outputting the planned path that meets the preset operating cost conditions, it further includes: controlling the target transport vehicle to complete the corresponding transport task based on the planned path; obtaining the satisfaction evaluation, actual energy consumption cost, and actual charging cost of the transport task; using the satisfaction evaluation, actual energy consumption cost, and actual charging cost to optimize the path planning problem model, so as to generate a new planned path for the target transport vehicle using the optimized path planning problem model.
[0056] Among them, the expression of the optimized path planning problem model is:
[0057]
[0058] Among them, π new is the optimized path planning problem model, π old is the path planning problem model, R is the cumulative updated reward function, and β is the KL divergence hyperparameter balance coefficient between the optimization target and the optimized path planning problem model and the path planning problem model.
[0059] During the actual execution process, the target transport vehicle in the embodiment of the present application can travel according to the planned path until it reaches the next target node or transport target, that is, it travels to the next transport point, returns to the warehouse, or travels to the charging station.
[0060] As a possible implementation, the embodiment of the present application can utilize the actual results of the current planned path, that is, the obtained customer satisfaction evaluation, actual energy consumption cost (power consumption), and actual charging cost to optimize the path planning problem model, such as optimizing the reward function, so that the optimized path planning problem model can output a better planning scheme when the target transport vehicle performs a new path planning to ensure effectiveness and economy.
[0061] To verify the effectiveness and economy of the embodiment of the present application, the embodiment of the present application can also perform simulation through experiments, that is, by simulating path planning problems in different scenarios, and performing simulation and calculation on the planned path and cost problems of the transport vehicle.
[0062] Combined with Figure 2 As shown, a detailed description of the path planning method for the transport vehicle in the embodiment of the present application is given in an embodiment.
[0063] As Figure 2 shown, the embodiment of the present application may include the following steps:
[0064] Step S1: Initialize the deep reinforcement learning model (transport path planning model), including setting the state space, action space, reward function, etc.
[0065] Among them, the state space includes the current position data of the target transport vehicle, the current remaining power, the node data in the transport itinerary (including the completed node data and uncompleted node data, etc.), the charging station information, and the current electricity price information, etc.;
[0066] The action space includes the next node data, travel target (such as going to the charging station to charge or returning to the warehouse);
[0067] The reward function considers factors such as transport cost, charging cost, power consumption, and customer satisfaction.
[0068] Step S2: Collect the remaining power of the target transport vehicle, the position information of the nodes, the layout of the charging stations, and the dynamic electricity price information.
[0069] Among them, the embodiment of the present application can collect the remaining power, the position information of the nodes, and the layout of the charging stations through in-vehicle sensors or vehicle networking technology;
[0070] The real-time electricity price information can be obtained through the electricity market trading platform or the smart grid system.
[0071] Step S3: Use the collected information as the input of the deep reinforcement learning model, and obtain the optimal path selection strategy by training the model.
[0072] Among them, in the embodiment of the present application, the collected information can be used as the input of the deep reinforcement learning model; by training the model, the model can output the optimal path selection strategy according to the input information; during the training process, the reward function and model parameters are continuously optimized to improve the accuracy and generalization ability of the model.
[0073] Step S4: Plan the driving path of the target transport vehicle according to the optimal path selection strategy.
[0074] Step S5: During the driving process, dynamically adjust the charging strategy according to the real-time electricity price information and the remaining power.
[0075] Among them, in the embodiment of the present application, the driving path of the target transport vehicle can be planned according to the optimal path selection strategy output by the deep reinforcement learning model; during the driving process, the charging strategy is dynamically adjusted according to the real-time electricity price information and the remaining power to ensure that the target transport vehicle can complete the distribution task on time and minimize the energy consumption cost as much as possible.
[0076] Step S6: Practical application and effect evaluation.
[0077] The embodiment of the present application can be applied to an actual logistics distribution system; through comparative experiments, evaluate the effects of the embodiment of the present application in reducing energy consumption costs, improving customer satisfaction, and optimizing the supply and demand balance of the power market.
[0078] According to the path planning method of the transport vehicle proposed by the embodiment of the present application, a path planning problem model can be pre-constructed, and the path planning problem model can be optimized and adjusted each time the target transport vehicle performs a transport task. Then, after obtaining the current position information, current transport itinerary, and current state data of the target transport vehicle under the current transport working condition, use the optimized path planning problem model to output a planned path that meets the preset operating cost conditions, so as to perform path planning according to the dynamic data under the current transport working condition, comprehensively consider the task objectives and dynamic costs, and obtain a transport path that balances efficiency and cost, which is convenient for popularization and application. Thus, it solves the technical problems in the related art that only considering static geographical information improves the transport efficiency limitedly, and there is a lack of dynamic consideration of transport costs, resulting in a relatively high overall transport cost and inconvenience for popularization and application.
[0079] Next, describe the path planning device of the transport vehicle according to the embodiment of the present application with reference to the accompanying drawings.
[0080] Figure 3 It is a block diagram of the path planning device of the transport vehicle according to the embodiment of the present application.
[0081] As shown Figure 3 The path planning device 10 of the transport vehicle includes: a first acquisition module 100, a second acquisition module 200, and a planning module 300.
[0082] Specifically, the first acquisition module 100 is configured to acquire the current transportation task of the target transport vehicle.
[0083] The second acquisition module 200 is configured to acquire the current position information, current transportation itinerary, and current status data of the target transport vehicle under the current transportation condition based on the current transportation task.
[0084] The planning module 300 is configured to input the current position information, current transportation itinerary, and current status data into the path planning problem model updated by the previous transportation task, so as to output a planned path that meets the preset operating cost condition, where the path planning problem model is cumulatively obtained from the reward functions corresponding to the historical transportation tasks of the target transport vehicle.
[0085] Optionally, in an embodiment of the present application, the planning module 300 includes: a first construction unit, a second construction unit, and a second construction unit.
[0086] Among them, the first construction unit is configured to construct the state space of the path planning problem model by using at least one of the current position information, the power data in the current status data, the node data in the current transportation itinerary, and the current electricity price data.
[0087] The second construction unit is configured to construct the action space of the path planning problem model by using at least one of the data of the next target node in the current transportation itinerary and the itinerary target.
[0088] The third construction unit is configured to cumulatively update the reward function of the path planning problem model by using at least one of the estimated transportation cost, estimated charging cost, and estimated power consumption data for traveling to the next node or itinerary target in the current transportation itinerary.
[0089] Optionally, in an embodiment of the present application, the path planning device 10 of the transport vehicle further includes: a control module, a third acquisition module, and an optimization module.
[0090] Among them, the control module is configured to control the target transport vehicle to complete the corresponding transportation task based on the planned path.
[0091] The third acquisition module is configured to acquire the satisfaction evaluation of the transportation task, the actual energy consumption cost, and the actual charging cost.
[0092] The optimization module is configured to optimize the path planning problem model by using the satisfaction evaluation, the actual energy consumption cost, and the actual charging cost, so as to generate a new planned path for the target transport vehicle by using the optimized path planning problem model.
[0093] Optionally, in an embodiment of the present application, the expression of the optimized path planning problem model is:
[0094]
[0095] Among them, π new is the optimized path planning problem model, π old is the path planning problem model, R is the cumulative updated reward function, and β is the KL divergence hyperparameter balance coefficient between the optimization objective and the optimized path planning problem model and the path planning problem model.
[0096] Optionally, in an embodiment of the present application, the planning module 300 includes: a generating unit and a determining unit.
[0097] Among them, the generating unit is used to integrate the road network data and the current state data involved in the current transportation trip to generate a corresponding observation vector.
[0098] The determining unit is used to generate a probability distribution for the selection of transportation nodes based on the observation vector, and to determine the next target node based on the probability distribution.
[0099] Optionally, in an embodiment of the present application, the planning module 300 includes: a calculating unit.
[0100] Among them, the calculating unit is used to obtain an estimated charging cost based on at least one charging station data and the current state data involved in the current transportation trip.
[0101] According to the path planning device for a transportation vehicle provided by the embodiments of the present application, a path planning problem model can be pre-constructed, and the path planning problem model can be optimized and adjusted each time the target transportation vehicle performs a transportation task. Furthermore, after obtaining the current position information, the current transportation trip, and the current state data of the target transportation vehicle under the current transportation condition, the optimized path planning problem model is used to output a planned path that meets the preset operating cost condition, so as to perform path planning according to the dynamic data under the current transportation condition, comprehensively consider the task objective and the dynamic cost, and obtain a transportation path that balances efficiency and cost, which is convenient for popularization and application. Thus, the technical problem in the related art that only considering static geographical information improves the transportation efficiency limitedly, and there is a lack of dynamic consideration of the transportation cost, resulting in a relatively high overall transportation cost and inconvenience for popularization and application is solved.
[0102] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device may include:
[0103] A memory 401, a processor 402, and a computer program stored in the memory 401 and executable on the processor 402.
[0104] When the processor 402 executes the program, it implements the path planning method for a transport vehicle provided in the above embodiments.
[0105] Furthermore, the electronic device further includes:
[0106] A communication interface 403 for communication between the memory 401 and the processor 402.
[0107] The memory 401 is used to store a computer program executable on the processor 402.
[0108] The memory 401 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.
[0109] If the memory 401, the processor 402, and the communication interface 403 are implemented independently, the communication interface 403, the memory 401, and the processor 402 can be interconnected through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0110] Optionally, in a specific implementation, if the memory 401, the processor 402, and the communication interface 403 are integrated on a single chip, the memory 401, the processor 402, and the communication interface 403 can communicate with each other through an internal interface.
[0111] The processor 402 may be a Central Processing Unit (CPU), or an Application Specific Integrated Circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0112] This embodiment also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the path planning method of the transportation vehicle as described above.
[0113] The embodiment of the present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the path planning method of the transportation vehicle provided by the embodiment of the present invention.
[0114] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic descriptions of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0115] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.
[0116] Any process or method description shown in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or N executable instructions for implementing a customized logic function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.
[0117] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definable sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion (electronic device) having one or N wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.
[0118] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or combinations thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0119] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0120] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0121] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for transport vehicle path planning, characterized in that: The following steps are involved: Get the current transport task of the target transport vehicle; Based on the current transportation task, obtain the current position information, current transportation itinerary and current status data of the target transportation vehicle under the current transportation condition; The current location information, current transport itinerary and current status data are input into a path planning problem model updated by the last transport task to output a planned path that meets preset operating cost conditions, wherein the path planning problem model is accumulated by a reward function corresponding to the historical transport tasks of the target transport vehicle.
2. The method according to claim 1, characterized in that The current location information, current transportation itinerary and current status data are input into the path planning problem model updated by the last transportation task to output a planned path that meets the preset operating cost conditions, including: Constructing the state space of the path planning problem model using at least one of the current location information, the power data in the current state data, the node data in the current transportation itinerary, and the current electricity price data; Constructing the action space of the path planning problem model using the data of the next target node in the current transport trip and at least one of the trip goals; The reward function of the path planning problem model is cumulatively updated using at least one of the estimated transportation cost, the estimated charging cost, and the estimated power consumption data for traveling to the next node or the trip target in the current transportation trip.
3. The method according to claim 2, characterized in that After outputting the planning path that meets the preset operating cost conditions, it also includes: Controlling the target transport vehicle to complete the corresponding transport task based on the planned path; Obtaining satisfaction evaluation, actual energy consumption cost and actual charging cost of the transportation task; The path planning problem model is optimized using the satisfaction evaluation, the actual energy consumption cost, and the actual charging cost, so as to generate a new planned path for the target transport vehicle using the optimized path planning problem model.
4. The method according to claim 3, characterized in that The expression of the optimized path planning problem model is: Among them, π new is the optimized path planning problem model, π old is the path planning problem model, R is the cumulative updated reward function, and β is the KL divergence hyperparameter balance coefficient between the optimization objective and the optimized path planning problem model and the path planning problem model.
5. The method according to claim 2, characterized in that: The current location information, current transportation itinerary and current status data are input into the path planning problem model updated by the last transportation task to output a planned path that meets the preset operating cost conditions, including: Integrating the road network data involved in the current transport itinerary and the current state data to generate a corresponding observation vector; A probability distribution of transportation node selection is generated based on the observation vector to determine the next target node based on the probability distribution.
6. The method according to claim 2, characterized in that The current location information, current transportation itinerary and current status data are input into the path planning problem model updated by the last transportation task to output a planned path that meets the preset operating cost conditions, including: The estimated charging cost is obtained based on at least one charging station data involved in the current transportation trip and the current status data.
7. A path planning device for a transport vehicle, characterized in that: include: A first acquisition module is used to acquire the current transport task of the target transport vehicle; A second acquisition module is used to acquire the current position information, current transportation itinerary and current status data of the target transportation vehicle under the current transportation condition based on the current transportation task; A planning module is used to input the current location information, current transport itinerary and current status data into a path planning problem model updated by the last transport task to output a planned path that meets preset operating cost conditions, wherein the path planning problem model is accumulated by a reward function corresponding to the historical transport tasks of the target transport vehicle.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the path planning method for a transport vehicle as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the path planning method for a transport vehicle as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed, it is used to implement the path planning method for a transport vehicle as described in any one of claims 1-6.