Smart guided trolley regulation method, device and equipment considering charging strategy
By optimizing the charging strategy of the intelligent guided vehicle through reinforcement learning algorithms, the problem of ignoring energy consumption in existing technologies is solved, intelligent power management is realized, and terminal operation efficiency is improved.
Patent Information
- Application Number
- CN202410490859.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-04-23
AI Technical Summary
The existing intelligent guided vehicle control strategy has simple operating logic in urban logistics or warehousing factories, ignores energy consumption, and makes it difficult to determine when, where, and how much to charge, resulting in low transportation efficiency in complex and high-intensity terminal operation scenarios.
By adopting reinforcement learning algorithm and training intelligent control model, combined with intelligent guidance of vehicles, charging station information and operation tasks, the charging strategy is optimized, intelligent control is realized, and the appropriate charging time, location and power are selected.
It improves the operating efficiency of the intelligent guided vehicle in complex and high-intensity dock operation scenarios, realizes intelligent power management, and is suitable for complex dock operation environments.
Smart Images

Figure CN118466407B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device and equipment for intelligently guiding a vehicle to control the vehicle in consideration of charging strategies. Background Art
[0002] With the rapid development of information technology and intelligent control technologies, related applications have gradually become integrated into people's lives, providing a wide range of services. For example, intelligent guided vehicles are a key piece of equipment in logistics and transportation systems. In the context of port terminals, these vehicles connect the quay cranes at the terminal's front edge with the storage areas within the terminal, serving as the terminal's primary transport equipment. The cost of heavy machinery at terminals is extremely high, and the efficiency of quay cranes and yard cranes in the storage areas is difficult to improve. Therefore, improving the transport efficiency of intelligent guided vehicles has become the most direct and effective way to enhance the efficiency of container terminals.
[0003] Existing intelligent guided vehicle control strategies are typically applied in urban logistics or warehouse factory settings, with relatively simple operational logic. Furthermore, most current intelligent guided vehicles use electricity, which consumes power during transportation. Existing intelligent guided vehicle control strategies often only optimize routes, ignoring energy consumption constraints. This makes it difficult to determine when to charge at a charging station, which charging station to go to, and how much to charge. This results in poor practical application and makes it difficult to cope with complex, high-intensity terminal operations.
[0004] Therefore, the problems existing in the existing technology still need to be solved and optimized. Summary of the Invention
[0005] The purpose of this application is to solve one of the technical problems existing in the related art to at least a certain extent.
[0006] To this end, one purpose of the embodiments of the present application is to provide a method, device and equipment for intelligently guiding a vehicle to control the vehicle while taking into account the charging strategy.
[0007] In order to achieve the above technical objectives, the technical solutions adopted in the embodiments of the present application include:
[0008] On the one hand, an embodiment of the present application provides a method for controlling an intelligent guided vehicle taking into account a charging strategy, the method comprising:
[0009] Acquire first information of each first intelligent guided vehicle set at the target wharf, second information of the first charging station, and target operation tasks;
[0010] Establishing a first state matrix according to the first information, the second information and the target operation task;
[0011] The first action prediction result is used to represent a next to-be-executed task selected by the intelligent regulation and control model for each first intelligent guiding trolley; wherein, the type of to-be-executed task includes a transportation task and a charging task;
[0012] According to the first action prediction result, each first intelligent guiding trolley is regulated and controlled, and the step of returning to execute the first information of each first intelligent guiding trolley, the second information of the first charging station and the target operation task set at the target wharf is returned.
[0013] In addition, the intelligent guiding trolley regulation and control device considering the charging strategy according to the above-mentioned embodiments of the present application can further have the following additional technical features:
[0014] Further, in an embodiment of the present application, the intelligent regulation and control model is obtained by training in the following manner:
[0015] The third information of each second intelligent guiding trolley, the fourth information of the second charging station and the training operation task set at the reference wharf are obtained;
[0016] According to the current third information, the current fourth information and the training operation task, a second state matrix is established;
[0017] The second action prediction result is outputted by the to-be-optimized intelligent regulation and control model based on the second state matrix; the second action prediction result is used to represent a next to-be-executed task selected by the intelligent regulation and control model for each second intelligent guiding trolley; wherein, the type of to-be-executed task includes a transportation task and a charging task;
[0018] According to the second action prediction result, a corresponding task reward is determined by a reinforcement learning algorithm;
[0019] According to the task reward, the parameters of the intelligent regulation and control model are updated to obtain a trained intelligent regulation and control model.
[0020] Further, in an embodiment of the present application, the updating of the parameters of the intelligent regulation and control model according to the task reward to obtain the trained intelligent regulation and control model comprises:
[0021] According to the task reward, the parameters of the intelligent regulation and control model are updated, and it is detected whether the current iteration parameter reaches a predetermined value;
[0022] If the current iteration parameter reaches a predetermined value, the current intelligent regulation model is determined as a trained intelligent regulation model; or, if the current iteration parameter does not reach the predetermined value, new third information and new fourth information are determined based on the environment according to the second action prediction result, and the step of establishing the second state matrix according to the current third information, the current fourth information and the training task is returned to be executed.
[0023] Further, in an embodiment of the present application, the reinforcement learning algorithm adopts a proximal policy optimization algorithm, and the intelligent regulation model comprises an Actor network and a Critic network.
[0024] Further, in an embodiment of the present application, the corresponding task reward is determined according to the second action prediction result by a reinforcement learning algorithm, comprising:
[0025] detecting a current training iteration round;
[0026] If the training iteration round is less than a preset threshold, the maximum cumulative time length and the minimum cumulative time length of completing all to-be-executed tasks in each second intelligent guiding trolley are determined according to the second action prediction result;
[0027] The difference value of the maximum cumulative time length and the minimum cumulative time length is calculated, and the corresponding task reward is determined according to the reciprocal of the difference value.
[0028] Further, in an embodiment of the present application, the method further comprises:
[0029] If the training iteration round is greater than or equal to the preset threshold, the maximum cumulative time length of completing all to-be-executed tasks in each second intelligent guiding trolley is determined according to the second action prediction result;
[0030] The corresponding task reward is determined according to the reciprocal of the maximum cumulative time length.
[0031] Further, in an embodiment of the present application, the corresponding task reward is determined according to the second action prediction result by a reinforcement learning algorithm, comprising:
[0032] The corresponding task reward is determined by the following formula:
[0033]
[0034] In the formula, R represents a task reward, a represents a first parameter, max t i represents the maximum cumulative time length of completing all to-be-executed tasks in the second intelligent guiding trolley, min t i represents the minimum cumulative time length of completing all to-be-executed tasks in the second intelligent guiding trolley.
[0035] In another aspect, the embodiments of the present application provide a smart guided vehicle regulation device considering charging strategy, the device comprising:
[0036] An acquisition unit is configured to acquire first information of each first smart guided vehicle, second information of a first charging station, and a target operation task set at a target wharf;
[0037] A building unit is configured to build a first state matrix according to the first information, the second information, and the target operation task;
[0038] A prediction unit is configured to output a first action prediction result based on the first state matrix by using a trained intelligent regulation model, wherein the first action prediction result is used to represent a next to-be-executed task selected by the intelligent regulation model for each first smart guided vehicle, and the type of the to-be-executed task includes a transportation task and a charging task;
[0039] A processing unit is configured to regulate each first smart guided vehicle according to the first action prediction result, and return to the step of acquiring the first information of each first smart guided vehicle, the second information of the first charging station, and the target operation task set at the target wharf.
[0040] In another aspect, the embodiments of the present application provide a computer device, comprising:
[0041] At least one processor;
[0042] At least one memory configured to store at least one program;
[0043] When the at least one program is executed by the at least one processor, the at least one processor is caused to implement the above-mentioned smart guided vehicle regulation method considering charging strategy.
[0044] In another aspect, the embodiments of the present application further provide a computer readable storage medium, wherein the computer readable storage medium stores a processor executable program, and the processor executable program is used to implement the above-mentioned smart guided vehicle regulation method considering charging strategy when executed by a processor.
[0045] The advantages and beneficial effects of the present application will be partially given in the following description, partially will become obvious from the following description, or will be understood by the practice of the present application:
[0046] The embodiment of the application discloses an intelligent guiding trolley regulation method considering charging strategy, obtains first information of each first intelligent guiding trolley, second information of a first charging station and a target operation task arranged at a target wharf; a first state matrix is established according to the first information, the second information and the target operation task; a first action prediction result is output based on the first state matrix through a trained intelligent regulation model; the first action prediction result is used to represent a next to-be-executed task selected by the intelligent regulation model for each first intelligent guiding trolley; wherein the type of the to-be-executed task includes a transportation task and a charging task; each first intelligent guiding trolley is regulated according to the first action prediction result, and the step of obtaining the first information of each first intelligent guiding trolley, the second information of the first charging station and the target operation task arranged at the target wharf is returned. The method can realize intelligent regulation of the intelligent guiding trolley by the strategy of reinforcement learning, can improve the operation efficiency of the intelligent guiding trolley under the condition of considering the power consumption of the intelligent guiding trolley, is suitable for complex and high-intensity wharf operation scenes, and is better in practicability. BRIEF DESCRIPTION OF DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the following introduces the drawings of the related technical solutions in the embodiments of the application or the prior art. It should be understood that the drawings in the following introduction are only for the convenience of clearly describing some embodiments in the technical solutions of the application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0048] Figure 1 An implementation environment schematic diagram of the intelligent guiding trolley regulation method considering charging strategy provided in the embodiments of the application;
[0049] Figure 2 A flowchart schematic diagram of the intelligent guiding trolley regulation method considering charging strategy provided in the embodiments of the application;
[0050] Figure 3 A structure schematic diagram of the intelligent regulation model provided in the embodiments of the application;
[0051] Figure 4 A structure schematic diagram of the computer device provided in the embodiments of the application. DETAILED DESCRIPTION
[0052] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not intended to limit the present application. When the following description refers to the accompanying drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with embodiments of the present application. They are merely examples of apparatuses and methods consistent with some aspects of the embodiments of the present application as detailed in the appended claims.
[0053] It can be understood that the terms "first", "second", and the like used in the present application can be used herein to describe various concepts, but unless specifically stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of the present application, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information. Depending on the context, the word "if" as used herein can be interpreted as "when" or "when" or "in response to determining".
[0054] The terms "at least one", "multiple", "each", "any" and the like used in the present application include one, two or more than two, multiple includes two or more than two, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as understood by those skilled in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0056] First, the terms involved in the present application are analyzed:
[0057] Artificial intelligence is a new technical science that studies, develops and applies systems for simulating, extending and expanding human intelligence. Artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, to perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0058] Machine learning, a multi-disciplinary subject, involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc., which is dedicated to studying how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. Its applications are widespread in various fields of artificial intelligence. Machine learning usually includes artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.
[0059] Reinforcement learning, a machine learning method, trains agents to learn from the environment and improve their behavior through trial and error and reward mechanisms. In reinforcement learning, agents learn the optimal behavior strategy to maximize cumulative rewards by interacting with the environment, choosing different actions, and receiving feedback from the environment.
[0060] Currently, with the rapid development of information technology and intelligent control technology, related applications have gradually integrated into people's lives, providing a variety of services. For example, intelligent guided vehicles are the main equipment in logistics transportation systems. In the context of ports, intelligent guided vehicles connect the shore cranes in the front of the terminal and the yard cranes inside the terminal, and are the main transportation equipment of the terminal. The cost of related heavy machinery at the terminal is very high, and the working efficiency of the shore cranes and the yard cranes is difficult to improve. Therefore, how to improve the transportation efficiency of the intelligent guided vehicles has become the most direct and effective way to improve the working efficiency of the container terminal.
[0061] In related technologies, the application of existing intelligent guided vehicle control strategies is usually in the context of urban logistics or warehouse factories, and the logic of operation is relatively simple. Moreover, most current intelligent guided vehicles use electric energy, which consumes electricity during transportation. The existing intelligent guided vehicle control strategy often optimizes scheduling only for route problems, ignores the limitation of energy consumption, and is difficult to determine when to charge, where to charge, and how much to charge, resulting in poor actual application results and difficulty in coping with complex and high-intensity terminal operation scenarios.
[0062] Therefore, in the embodiments of the present application, an intelligent guided vehicle control method considering charging strategy is provided. The method can realize intelligent control of the intelligent guided vehicle by considering the power consumption of the intelligent guided vehicle through the strategy of reinforcement learning, improve the operation efficiency of the intelligent guided vehicle, and be suitable for complex and high-intensity terminal operation scenarios, thus being more practical.
[0063] Next, the implementation environment related to the intelligent guided vehicle control method considering charging strategy provided in the embodiments of the present application is introduced. Referring to Figure 1 , Figure 1An implementation environment schematic diagram of a smart guided trolley regulation method considering charging strategy is given. The main body of the hardware and software of the implementation environment mainly includes a terminal device 110 and a server 120, and the terminal device 110 is in communication connection with the server 120. The smart guided trolley regulation method considering charging strategy can be configured on the terminal device 110 side, or the smart guided trolley regulation method considering charging strategy can be realized through the interaction of the terminal device 110 and the server 120.
[0064] Specifically, in the embodiment of the present application, the terminal device 110 can include but is not limited to any one or more of a smart watch, a smart phone, a computer, a personal digital assistant (PDA), a smart voice interaction device, or a vehicle-mounted terminal. The server 120 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal device 110 and the server 120 can establish a communication connection through a wireless network or a wired network, which uses standard communication technology and / or protocol. The network can be set as the Internet, or any other network, such as any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network.
[0065] Of course, it can be understood that Figure 1 The implementation environment in the foregoing embodiment is only one optional application scenario of the smart guided trolley regulation method considering charging strategy provided in the embodiment of the present application, and the actual application is not fixed to the hardware and software environment shown in Figure 1 The smart guided trolley regulation method considering charging strategy provided in the embodiment of the present application is described in detail below in combination with the implementation environment shown in Figure 1
[0066] Please refer to Figure 2 , Figure 2 is a flowchart of a smart guided trolley regulation method considering charging strategy provided in the embodiment of the present application. Referring to Figure 2 , the smart guided trolley regulation method considering charging strategy provided in the present application includes but is not limited to:
[0067] Step 210, obtaining first information of each first intelligent guiding trolley, second information of a first charging station and a target operation task arranged at a target wharf;
[0068] Step 220, establishing a first state matrix according to the first information, the second information and the target operation task;
[0069] Step 230, outputting a first action prediction result based on the first state matrix through a trained intelligent regulation and control model; the first action prediction result is used to represent a next to-be-executed task selected by the intelligent regulation and control model for each first intelligent guiding trolley; wherein the type of to-be-executed task includes a transportation task and a charging task;
[0070] Step 240, regulating and controlling each first intelligent guiding trolley according to the first action prediction result, and returning to the step of obtaining the first information of each first intelligent guiding trolley, the second information of the first charging station and the target operation task arranged at the target wharf.
[0071] In the embodiment of the application, an intelligent guiding trolley regulation and control method considering charging strategy is provided. The method can realize automatic regulation and control of the intelligent guiding trolley through reinforcement learning, so as to improve the work efficiency of the intelligent guiding trolley. Specifically, in the embodiment of the application, first, the related configuration information of a wharf that needs to be regulated and controlled can be obtained, and the wharf is recorded as a target wharf. The information of each intelligent guiding trolley arranged at the target wharf can be obtained, and the intelligent guiding trolleys are recorded as first intelligent guiding trolleys, and the information of the intelligent guiding trolleys is recorded as first information. The first information can include the number information of the first intelligent guiding trolleys, the position information of each first intelligent guiding trolley, the current power and the charging threshold power, etc. In the embodiment of the application, the power of the intelligent guiding trolley can include three gears. When the power of the intelligent guiding trolley is in the highest gear (such as 80%-100%), no charging is arranged (i.e., only the transportation task can be executed and the charging task cannot be executed); when the power of the intelligent guiding trolley is in the lowest gear (such as 0%-20%), forced charging is arranged (i.e., only the charging task can be executed and the transportation task cannot be executed); and when the power is in the middle gear (such as 20%-80%), whether to charge is determined by the strategy learned by the intelligent agent of reinforcement learning.
[0072] In the embodiment of the present application, the information of the charging station set up at the target terminal is also obtained, and the charging station is recorded as the first charging station, and the information of the first charging station is recorded as the second information. The second information may include the location of the first charging station, the number of charging vehicles, the available spaces, etc. In addition, the target operation task is also obtained. The target operation task refers to the task that needs to be completed by the intelligent guided vehicle in this round of regulation. It can be represented by a container operation task. For example, it can be recorded where each container needs to be transported from and to, thereby obtaining the target operation task.
[0073] In the embodiment of the present application, after obtaining the above information, the control of the intelligent guided car can be realized through the reinforcement learning algorithm. Below, the intelligent agent, state, action, reward, environment, etc. involved in this scenario are introduced.
[0074] (1) Intelligent Agent
[0075] The agent is set as the system that controls the intelligent guided vehicles, and through this system, all intelligent guided vehicles are dispatched in a unified manner, rather than having each intelligent guided vehicle be a separate agent. This means that this problem is a single-agent problem, not a multi-agent problem.
[0076] (2) Status
[0077] S = (S1, S2, S3), the state consists of three parts: the intelligent guided vehicle state, the task state, and the charging station state, which contains important information for interacting with the environment.
[0078] The state of the intelligent guided vehicle is a matrix formed by the aggregation of the states of all intelligent guided vehicles. Each intelligent guided vehicle state contains four pieces of information: current coordinates, accumulated working time (the total time of the transport task and the charging task), and power level. Therefore, the state matrix of the intelligent guided vehicle is as follows:
[0079]
[0080] Among them, (x i y i t i e i ) represents the current position of the i-th intelligent guided vehicle (x i ,y i ), at this time, the accumulated working time t i , the remaining power is e i .
[0081] The optimization objective in the embodiments of the present application is to make the container transport operation end time shortest. From the state of the intelligent guiding trolley described above, the working time of each intelligent guiding trolley is accumulated separately and does not interfere with each other. Therefore, only the intelligent guiding trolley with the maximum accumulated working time needs to be found after all tasks are completed, and the accumulated working time is the completion time. Therefore, the optimization objective is converted to make the intelligent guiding trolley with the maximum accumulated working time reach the minimum time. Therefore, the objective function of the embodiments of the present application is:
[0082] T = min max t i
[0083] In the embodiments of the present application, the task state contains information of all container operation tasks. Each task information contains three information of task starting point, task ending point and whether it has been completed. Therefore, the state matrix corresponding to the task is as follows:
[0084]
[0085] Wherein, (x i1 y i1 x i2 y i2 f i ) represents that the i-th task is from the starting point (x i1 , y i1 ) to the ending point (x i2 , y i2 ). If f i =1, it indicates that the task is completed, and f i =0 is not completed.
[0086] The charging station state is a collection of all charging station related information. Each charging station state contains two information of charging station coordinates and the number of charging vehicles in the charging station. Therefore, the state matrix corresponding to the charging station is:
[0087]
[0088] Wherein, (p j q j r j ) represents that the j-th charging station is currently at (p j , q j ), and there are ri intelligent guiding trolleys in the charging station.
[0089] (Three) Action
[0090] In the embodiments of the present application, the action is designed as a two-dimensional action A=(A1, A2), which contains two actions.
[0091] A1 is the first step action, which means selecting a vehicle, and the formula is:
[0092] A1 = (a1, a2,..., a n )
[0093] where a i means selecting the i-th vehicle to work.
[0094] A2 is the second step action, which means selecting a task, including transportation task and charging task. If the selected task is transportation task, the container transportation task will be executed, and if the selected task is charging task, the intelligent guiding trolley will go to the charging station to charge. The selected charging station and charging level are determined by the intelligent agent itself. The formula is:
[0095] A2 = (b1, b2,..., b m , b 11 , b 12 ,..., b 1t ,..., b ct )
[0096] where b i means selecting the i-th task for transportation, and b jk means going to the j-th charging station to charge to level k.
[0097] (Four) Reward
[0098] As mentioned above, the optimization goal in the embodiment of the present application is to minimize the maximum cumulative working time of the intelligent guiding trolley, so directly taking the reciprocal of this time as the reward of reinforcement learning is the most concise and effective method. However, the environment of the intelligent guiding trolley scheduling problem in the embodiment of the present application is complex, and the reward needs to be further designed. Because the parameters of the neural network are randomly initialized at the beginning of reinforcement learning, the learning of the intelligent agent is not intelligent enough, so the feedback obtained from the reward is less, and the learning is difficult to form a strategy, so directly applying the maximum cumulative working time of the intelligent guiding trolley as the reward is not effective. If the intelligent agent takes the reciprocal of the difference between the maximum cumulative working time and the minimum cumulative working time as the reward, it can learn a better strategy in the learning process: to make the cumulative working time of the intelligent guiding trolley as evenly distributed as possible. Although this is not the optimal strategy, it can serve as a preliminary strategy for the intelligent agent to help it have a good transition in the early stage.
[0099] (Five) Environment
[0100] In reinforcement learning, the agent is responsible for selecting actions based on the state, while the environment changes its state based on the output action. In the intelligent guided vehicle scheduling problem, if the agent performs a transport action, the state changes as follows: For the intelligent guided vehicle state, the current position of the intelligent guided vehicle changes to the transport destination, the time accumulated from the current position to the task starting point and then to the task destination, and the battery level decreases with the distance traveled. For the task state, the completion status of the task is marked as completed. For the charging station state, if the intelligent guided vehicle's previous action was charging, the number of vehicles at the corresponding charging station decreases by one. If it was charging, the state changes as follows: For the intelligent guided vehicle state, the current position of the intelligent guided vehicle changes to the charging station position, the time accumulated to travel to the charging station and the charging time to the specified level, and the battery level is restored to the specified level. For the task state, there is no change. For the charging station state, the number of vehicles charging at the charging station increases by one.
[0101] In an embodiment of the present application, during the application phase, relevant state matrices can be established based on the acquired first information, second information, and target operation tasks, and these matrices are recorded as first state matrices. Then, the trained intelligent control model can be used to predict based on the first state matrix to obtain a first action prediction result. The first action prediction result can be used to characterize the next task to be performed selected for each first intelligent guided vehicle. These tasks to be performed can be transportation tasks or charging tasks, and this application does not limit this. Then, each first intelligent guided vehicle can be controlled based on the first action prediction result. After control, the status of each intelligent guided vehicle, charging station, and task will change. Therefore, in an embodiment of the present application, step 210 can be returned to and the execution can be repeated until all tasks are completed.
[0102] It can be understood that the intelligent guided vehicle control method considering the charging strategy provided in the embodiment of the present application can realize the intelligent control of the intelligent guided vehicle by taking into account the power consumption of the intelligent guided vehicle through the reinforcement learning strategy, thereby improving the operating efficiency of the intelligent guided vehicle, being suitable for complex and high-intensity terminal operation scenarios, and having better practicality.
[0103] In some embodiments, the intelligent control model is trained in the following manner:
[0104] Obtaining third information of each second intelligent guided vehicle set at the reference dock, fourth information of the second charging station, and a training task;
[0105] Establishing a second state matrix according to the current third information, the current fourth information, and the training job task;
[0106] The second action prediction result is output by the intelligent regulation and control model to be optimized based on the second state matrix, and the second action prediction result is used to represent a next to-be-executed task selected by each second intelligent guiding trolley according to the intelligent regulation and control model; wherein, the type of the to-be-executed task includes a transportation task and a charging task.
[0107] According to the second action prediction result, a corresponding task reward is determined by a reinforcement learning algorithm.
[0108] According to the task reward, the parameters of the intelligent regulation and control model are updated to obtain a trained intelligent regulation and control model.
[0109] In the embodiment of the application, when the intelligent regulation and control model is used, it needs to be trained first. Specifically, during the training, the setting information of the related wharf can be obtained, and the wharf is referred to as a reference wharf. It can be understood that in the embodiment of the application, the wharf used during the training and the wharf corresponding to the actual application can be the same or different. The intelligent guiding trolley at the reference wharf is referred to as a second intelligent guiding trolley, and the charging station is referred to as a second charging station. The third information of each second intelligent guiding trolley, the fourth information of the second charging station, and the training operation task can be obtained, and then the corresponding state matrix can be established according to the current third information, the fourth information, and the training operation task. In the embodiment of the application, it is referred to as a second state matrix. The second action prediction result can be predicted and output by the intelligent regulation and control model to be optimized according to the second state matrix. The second action prediction result represents the next to-be-executed task selected by each second intelligent guiding trolley according to the intelligent regulation and control model. The specific content can be implemented by referring to the foregoing target wharf, which will not be described herein.
[0110] In the embodiment of the application, the corresponding task reward can be determined by a reinforcement learning algorithm according to the second action prediction result, and then the parameters of the intelligent regulation and control model are updated according to the task reward to obtain a trained intelligent regulation and control model. Specifically, in the embodiment of the application, the training of the intelligent regulation and control model can be performed in multiple iterations. After the parameters of the intelligent regulation and control model are updated each time, it can be detected whether the current iteration parameter reaches a predetermined value, such as whether the current iteration round reaches a predetermined round or whether the current certain index (such as a loss value) meets a predetermined threshold. If so, it can be considered that the training is completed, and the current intelligent regulation and control model can be determined as a trained intelligent regulation and control model. If the current iteration parameter does not reach the predetermined value, the new third information and the fourth information can be determined based on the environment according to the second action prediction result, so as to re-determine the second state matrix and iteratively update the parameters of the intelligent regulation and control model.
[0111] In some embodiments, the determining, according to the second action prediction result, the corresponding task reward through the reinforcement learning algorithm comprises:
[0112] detecting a current training iteration round;
[0113] if the training iteration round is less than a preset threshold, determining, according to the second action prediction result, a maximum cumulative time length and a minimum cumulative time length in which all the second intelligent guiding trolleys complete all the to-be-executed tasks;
[0114] calculating a difference between the maximum cumulative time length and the minimum cumulative time length, and determining the corresponding task reward according to an inverse of the difference;
[0115] if the training iteration round is greater than or equal to the preset threshold, determining, according to the second action prediction result, a maximum cumulative time length in which all the second intelligent guiding trolleys complete all the to-be-executed tasks;
[0116] determining the corresponding task reward according to an inverse of the maximum cumulative time length.
[0117] In the embodiments of the present application, when the task reward is determined through the reinforcement learning algorithm, as mentioned above, the inverse of the difference between the maximum cumulative working time and the minimum cumulative working time is taken as the reward, and a relatively optimal strategy can be learned in the learning process: the cumulative working time of the intelligent guiding trolley is arranged as evenly as possible. Although this is not the optimal strategy, it can be used as a preliminary strategy of the intelligent agent to help the intelligent agent have a good transition in the early stage. Therefore, in the embodiments of the present application, the current training iteration round can be detected, and it is compared with the preset threshold. If the training iteration round is less than the preset threshold, the maximum cumulative time length and the minimum cumulative time length in which all the second intelligent guiding trolleys complete all the to-be-executed tasks can be determined according to the second action prediction result, and then the difference between the two is calculated, and the inverse of the difference is used to determine the corresponding task reward. If the training iteration round is greater than or equal to the preset threshold, the inverse of the maximum cumulative time length is used to determine the corresponding task reward. In some embodiments, the corresponding task reward can also be determined through the following formula:
[0118]
[0119] In the formula, R represents the task reward, a represents a first parameter, max t i represents the maximum cumulative time length in which all the second intelligent guiding trolleys complete all the to-be-executed tasks, min t i represents the minimum cumulative time length in which all the second intelligent guiding trolleys complete all the to-be-executed tasks.
[0120] In reinforcement learning algorithms, the most commonly used method is DQN, which has good results in many simple application scenarios. However, the intelligent guiding trolley scheduling problem in the present application is a complex problem, and both its state space and action space are very large, so it is not suitable to use DQN and other methods to solve this problem. Therefore, the PPO (Proximal Policy Optimization) algorithm is used in the present application. The advantage of the PPO algorithm is that it achieves a good balance between ease of implementation, performance, and algorithm efficiency. Specifically, the PPO algorithm is implemented based on the Actor-Critic framework, which consists of two parts: the Actor network and the Critic network. The Actor network is responsible for forming the policy (i.e., outputting the action), while the Critic network is responsible for evaluating the quality of the policy to provide guidance for the parameter update of the Actor network. Both networks are fully connected networks, and the Critic network directly outputs the evaluation value V(s), while the Actor network is different. The Actor network outputs the mean and variance of the action, which forms a Gaussian distribution, and then a sample is taken from the distribution to obtain an action.
[0121] In combination with the content of reinforcement learning, the overall framework is as shown in FIG. 1. During training, the environment generates the state (intelligent guiding trolley state, task state, charging station state) and reward (maximum cumulative working time) to the agent according to the action. Then the agent executes the PPO algorithm according to the state and reward, and outputs the action again, which is one iteration. The PPO algorithm updates the network parameters in the agent every T iterations. Figure 3
[0122] In the agent, the first level network is the Actor network. Its input is the state S, and the output is the action A, which is to generate the optimal action according to the state. The Actor network approaches the best action by maximizing the following loss function:
[0123]
[0124]
[0125]
[0126] where θ is the parameter of the policy π θ . represents the expected experience over a period of time steps. is the advantage estimate at time t. It is used to compare with the historical average action to evaluate the quality of the action. t (θ) is the ratio of the probabilities under the new policy to the old policy. ε is a hyperparameter, typically set to 0.1 or 0.2. The clip function is a truncation function that returns the upper or lower bound when the importance sampling exceeds the specified upper or lower bound, ensuring that the new and old policies are sufficiently similar and do not change significantly.
[0127] The second-level network is the Critic network. It receives the state S and outputs the evaluation value V(s). The Critic network approaches the true cumulative reward by minimizing the following loss function:
[0128]
[0129]
[0130] in is the value function of state st, and V target is the target value function. is the parameter of the value function, and γ is the reduction coefficient.
[0131] The specific algorithm flow of the PPO algorithm is as follows: In each iteration, the agent interacts with the environment until all tasks are completed. Assume that there are T interactions in total. Then we can calculate the reward r according to the reward. t and the value function V θ (s t ) calculate the odds estimate Similarly, we can also get the target value function V according to the formula target The next step is to maximize the loss function L clip (θ) to update the parameters of the Actor network, that is, the strategy π θ θ in . This partially ensures that the update between the old and new policies will not be too large. Finally, minimize the loss function To update the parameters of the Critic network. This completes one iteration.
[0132] The following describes an intelligent guided vehicle control device that takes into account charging strategies according to an embodiment of the present application.
[0133] The intelligent guided vehicle control device considering the charging strategy proposed in the embodiment of the present application includes:
[0134] an acquiring unit, configured to acquire first information of each first intelligent guided vehicle provided at a target terminal, second information of the first charging station, and a target operation task;
[0135] an establishing unit, configured to establish a first state matrix according to the first information, the second information and the target operation task;
[0136] a prediction unit, configured to output a first action prediction result based on the first state matrix using a trained intelligent control model; the first action prediction result is used to represent the next task to be performed selected by the intelligent control model for each of the first intelligent guided vehicles; wherein the types of tasks to be performed include transportation tasks and charging tasks;
[0137] A processing unit is used to control each of the first intelligent guided vehicles according to the first action prediction result, and return to execute the step of obtaining the first information of each first intelligent guided vehicle set at the target terminal, the second information of the first charging station and the target operation task.
[0138] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0139] Reference Figure 4 , an embodiment of the present application provides a computer device, including:
[0140] at least one processor 410;
[0141] at least one memory 420, for storing at least one program;
[0142] When at least one program is executed by at least one processor 410, the at least one processor 410 implements Figure 2 An intelligent guided vehicle control method considering charging strategy is shown.
[0143] Similarly, the contents of the above method embodiments are applicable to the present computer device embodiment. The functions specifically implemented by the present computer device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0144] An embodiment of the present application also provides a computer-readable storage medium, which stores a program executable by the processor 410. When the program executable by the processor 410 is executed by the processor 410, it is used to execute the above-mentioned intelligent guided vehicle control method considering the charging strategy.
[0145] The present application also discloses a computer-readable storage medium in which a program executable by a processor is stored. When the program is executed by the processor, it is used to implement the following Figure 2 An embodiment of an intelligent guided vehicle control method considering charging strategy is shown.
[0146] It is understandable that if Figure 2The contents of the embodiments of the intelligent guided vehicle regulation method considering charging strategy shown are applicable to the embodiments of the computer readable storage medium, the embodiments of the computer readable storage medium specifically implement the functions as Figure 2 The embodiments of the intelligent guided vehicle regulation method considering charging strategy shown are the same as the embodiments of the computer readable storage medium, and the beneficial effects achieved are the same as Figure 2 The embodiments of the intelligent guided vehicle regulation method considering charging strategy shown are the same as the embodiments of the computer readable storage medium, and the beneficial effects achieved are the same as
[0147] In some alternative embodiments, the functions / operations mentioned in the block diagram can not occur in the order mentioned in the operation diagram. For example, depending on the functions / operations involved, two blocks shown in succession can actually be executed substantially concurrently or the blocks can sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts in this application are provided by way of example only. The disclosed methods are not limited to the operations and logical flows presented in this application. Alternative embodiments are contemplated in which the order of various operations is changed and in which sub-operations described as part of a larger operation are executed independently.
[0148] Furthermore, although the present application is described in the context of functional modules, it is to be understood that one or more of the functions and / or features can be integrated in a single physical system and / or software module, or one or more functions and / or features can be implemented in separate physical systems or software modules. It is also to be understood that detailed discussion of the actual implementation of each module is unnecessary to an understanding of the present application. Rather, the actual implementation is within the routine of an engineer's knowledge given the property, functionality and internal relationships of the various functional modules in the system disclosed herein. Accordingly, the present application is not limited to purely hardware or software implementations, but rather encompasses hybrid implementations within the scope of the appended claims and their equivalents.
[0149] If the functions are implemented in software, the functions can be stored in or implemented as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage medium can be any available medium that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, or twisted pair, then the coaxial cable, fiber optic cable, or twisted pair are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), and Blu-Ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0150] In other words, like a human driver of a vehicle, an autonomous vehicle can be programmed to follow traffic laws and to make decisions based on its environment. For example, an autonomous vehicle can be programmed to follow a speed limit, to stop at a stop sign, to yield to a pedestrian, to merge onto a highway, to change lanes, to park, and so on. In some embodiments, an autonomous vehicle can be programmed to follow traffic laws and to make decisions based on its environment using a machine learning algorithm. For example, an autonomous vehicle can be programmed to follow a speed limit, to stop at a stop sign, to yield to a pedestrian, to merge onto a highway, to change lanes, to park, and so on using a machine learning algorithm.
[0151] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber (optical), and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via optical scanning of the paper or other medium, then compiled, interpreted or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.
[0152] It should be understood that aspects of the application can be implemented in hardware, software, firmware or a combination of them. In the above embodiments, various steps or methods can be implemented in software or firmware that is stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any of the following technologies, or combinations thereof, can be used: a discrete logic circuit having logic gates for implementing logic functions upon data signals, an application specific integrated circuit having appropriate combinational logic gates, a programmable gate array (PGA), a field programmable gate array (FPGA), and / or the like.
[0153] In the above description of the present specification, the description referring to the terms "one embodiment", "another embodiment" or "certain embodiments" or the like means that a specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. The illustrative expressions of the above terms do not necessarily refer to the same embodiment or example in the present specification. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.
[0154] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that changes, modifications, alternatives and variations to these embodiments can be made without departing from the principles and spirit of the application, and the scope of the application is defined by the claims and their equivalents.
[0155] The above is a specific description of the preferred embodiments of the present application, but the present application is not limited to the embodiments, and those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present application, and these equivalent modifications or substitutions are included in the scope defined by the claims of the present application
[0156] In the above description of the present specification, the description referring to the terms "one embodiment", "another embodiment" or "certain embodiments" or the like means that a specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the present application. The illustrative expressions of the above terms do not necessarily refer to the same embodiment or example in the present specification. Also, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in an appropriate manner.
[0157] Although the embodiments of the present application have been shown and described, it will be appreciated by those skilled in the art that changes, modifications, alternatives and variations to these embodiments can be made without departing from the principles and spirit of the application, and the scope of the application is defined by the claims and their equivalents.
Claims
1. A method for controlling an intelligent guided vehicle considering a charging strategy, characterized in that: The method comprises: Acquire first information of each first intelligent guided vehicle set at the target wharf, second information of the first charging station, and target operation tasks; Establishing a first state matrix according to the first information, the second information and the target operation task; Outputting a first action prediction result based on the first state matrix using the trained intelligent control model; the first action prediction result is used to represent the next task to be performed selected by the intelligent control model for each of the first intelligent guided vehicles; wherein the types of tasks to be performed include transportation tasks and charging tasks; Control each of the first intelligent guided vehicles according to the first action prediction result, and return to the step of obtaining the first information of each of the first intelligent guided vehicles set at the target terminal, the second information of the first charging station, and the target operation task; The intelligent control model is trained in the following way: Obtaining third information of each second intelligent guided vehicle set at the reference dock, fourth information of the second charging station, and a training task; Establishing a second state matrix according to the current third information, the current fourth information, and the training job task; Outputting a second action prediction result based on the second state matrix by the intelligent control model to be optimized; the second action prediction result is used to represent the next task to be performed selected by the intelligent control model for each second intelligent guided vehicle; wherein the types of tasks to be performed include transportation tasks and charging tasks; Determine the corresponding task reward through a reinforcement learning algorithm based on the second action prediction result; According to the task reward, the parameters of the intelligent control model are updated to obtain a trained intelligent control model; The reinforcement learning algorithm adopts a proximal strategy optimization algorithm, and the intelligent control model includes an Actor network and a Critic network; Determining a corresponding task reward through a reinforcement learning algorithm based on the second action prediction result includes: The corresponding task reward is determined by the following formula: Where, Indicates the task reward, Represents the first parameter, Indicates the maximum cumulative time to complete all pending tasks in the second intelligent guided vehicle. Indicates the minimum cumulative time required to complete all pending tasks in the second intelligent guided vehicle.
2. The intelligent guided vehicle control method considering charging strategy according to claim 1 is characterized in that: The updating of the parameters of the intelligent control model according to the task reward to obtain a trained intelligent control model includes: According to the task reward, the parameters of the intelligent control model are updated, and whether the current iteration parameters reach a predetermined value is detected; If the current iteration parameter reaches a predetermined value, the current intelligent control model is determined as a trained intelligent control model; or, if the current iteration parameter does not reach a predetermined value, new third information and new fourth information are determined based on the environment according to the second action prediction result, and the step of establishing a second state matrix based on the current third information, the current fourth information and the training job task is returned to be executed.
3. The intelligent guided vehicle control method considering charging strategy according to claim 1 is characterized in that: Determining a corresponding task reward through a reinforcement learning algorithm based on the second action prediction result includes: Detect the current training iteration round; If the number of training iterations is less than a preset threshold, determining the maximum cumulative time and the minimum cumulative time for completing all pending tasks in each of the second intelligent guided vehicles based on the second action prediction result; The difference between the maximum cumulative duration and the minimum cumulative duration is calculated, and the corresponding task reward is determined according to the inverse of the difference.
4. The intelligent guided vehicle control method considering charging strategy according to claim 3 is characterized in that: The method further comprises: If the number of training iterations is greater than or equal to a preset threshold, determining the maximum cumulative time for each of the second intelligent guided vehicles to complete all pending tasks based on the second action prediction result; The corresponding task reward is determined according to the inverse of the maximum accumulated time.
5. An intelligent guidance vehicle control device considering charging strategy, characterized in that: The device comprises: an acquiring unit, configured to acquire first information of each first intelligent guided vehicle provided at a target terminal, second information of the first charging station, and a target operation task; an establishing unit, configured to establish a first state matrix according to the first information, the second information and the target operation task; a prediction unit, configured to output a first action prediction result based on the first state matrix using a trained intelligent control model; the first action prediction result is used to represent the next task to be performed selected by the intelligent control model for each of the first intelligent guided vehicles; wherein the types of tasks to be performed include transportation tasks and charging tasks; a processing unit, configured to control each of the first intelligent guided vehicles according to the first action prediction result, and return to execute the step of obtaining the first information of each of the first intelligent guided vehicles set at the target terminal, the second information of the first charging station, and the target operation task; The intelligent control model is trained in the following way: Obtaining third information of each second intelligent guided vehicle set at the reference dock, fourth information of the second charging station, and a training task; Establishing a second state matrix according to the current third information, the current fourth information, and the training job task; Outputting a second action prediction result based on the second state matrix by the intelligent control model to be optimized; the second action prediction result is used to represent the next task to be performed selected by the intelligent control model for each second intelligent guided vehicle; wherein the types of tasks to be performed include transportation tasks and charging tasks; Determine the corresponding task reward through a reinforcement learning algorithm based on the second action prediction result; According to the task reward, the parameters of the intelligent control model are updated to obtain a trained intelligent control model; The reinforcement learning algorithm adopts a proximal strategy optimization algorithm, and the intelligent control model includes an Actor network and a Critic network; Determining a corresponding task reward through a reinforcement learning algorithm based on the second action prediction result includes: The corresponding task reward is determined by the following formula: Where, Indicates the task reward, Represents the first parameter, Indicates the maximum cumulative time to complete all pending tasks in the second intelligent guided vehicle. Indicates the minimum cumulative time required to complete all pending tasks in the second intelligent guided vehicle.
6. A computer device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the intelligent guided vehicle control method considering the charging strategy as described in any one of claims 1-4.
7. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement an intelligent guided vehicle control method considering a charging strategy as described in any one of claims 1 to 4 when executed by the processor.
Citation Information
Patent Citations
Cooperative charging method based on multi-agent reinforcement learning
CN114202168A
DDQN-based automatic container terminal AGV scheduling method
CN114912809A