Generation method and device of intelligent logistics vehicle optimization control model and storage medium
Through a deep deterministic strategy gradient algorithm model trained in the simulation environment and optimized in the actual environment, a smart logistics vehicle optimization control model is generated, which solves the safety problems of logistics vehicle path planning and control in complex environments, and achieves efficient and safe logistics vehicle operation.
Patent Information
- Application Number
- CN202510516099.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-12
AI Technical Summary
The existing logistics vehicle control methods are difficult to achieve efficient and safe path planning and vehicle control in complex dynamic environments. The traditional phased control methods cannot be feedback and adjusted in real time, resulting in safety hazards.
Based on the dynamic information and risk information of logistics vehicles in the simulation environment, deep deterministic strategy gradient algorithm model is trained, and the smart logistics vehicle optimization control model is generated, and environmental changes are perceived in real time and control strategies are optimized.
It realizes efficient and safe operation of logistics vehicles in complex environments, significantly improves decision-making efficiency and adaptability, and reduces safety risks.
Smart Images

Figure CN120469267A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of autonomous driving technology, and in particular to a method, device, and storage medium for generating an optimization control model for a smart logistics vehicle. Background Art
[0002] Most existing logistics vehicle control methods treat path planning and vehicle control as two separate phases. The path planning algorithm outputs a target trajectory, and the vehicle controller generates control instructions based on this trajectory. However, this phased approach presents numerous challenges in complex dynamic environments. For example, the target trajectory may become invalid due to unexpected obstacles or environmental changes, preventing the controller from making timely adjustments, thus posing a safety hazard.
[0003] In other words, as logistics scenarios become increasingly complex and dynamic, traditional control methods often rely on static environmental models for offline planning, lacking real-time feedback and adjustments. Unexpected environmental conditions (such as the sudden appearance of obstacles, changing traffic conditions, or equipment failures) can render existing control strategies ineffective, and traditional methods often lack the ability to quickly recalculate and dynamically adjust. This limitation of offline training results in insufficient model performance in dynamic and complex environments, making it difficult to meet the requirements for efficient and safe operation of logistics vehicles. Summary of the Invention
[0004] In view of this, it is necessary to provide a method, device and storage medium for generating an optimization control model of a smart logistics vehicle to solve the technical problem that the existing logistics vehicle control method is difficult to meet the requirements of efficient and safe operation of logistics vehicles.
[0005] In order to solve the above problems, in a first aspect, the present invention provides a method for generating an optimization control model of a smart logistics vehicle, comprising: Based on the dynamic information and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model; The dynamic information and risk information of the logistics vehicle in the actual environment are used as the input of the preliminary training model, and the logistics vehicle control instructions are used as the output of the preliminary training model. The preliminary training model is retrained, and the dynamics of the logistics vehicle are adjusted based on the logistics vehicle control instructions. The preliminary training model is optimized according to the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0006] In one possible implementation, the risk information of the logistics vehicle in the simulation environment is determined based on the driving state of the logistics vehicle and the relative positional relationship between the logistics vehicle and the road boundary and obstacles in the simulation environment; The risk information of logistics vehicles in the actual environment is determined based on the driving status of the logistics vehicles and the relative position relationship between the logistics vehicles and road boundaries and obstacles in the actual environment.
[0007] In one possible implementation, risk information of a logistics vehicle is determined based on the vehicle's driving status and the relative positional relationship between the logistics vehicle and road boundaries and obstacles, including: The road boundary is modeled as a polygonal curve composed of discrete points, and the drivable area of the logistics vehicle is defined in the polygonal curve to obtain a road boundary risk field model; Determine the current position of the logistics vehicle based on the driving state of the logistics vehicle, determine the minimum distance between the current position of the logistics vehicle and the road boundary point based on the road boundary risk field model and the current position of the logistics vehicle, and determine the risk of the logistics vehicle deviating from the lane based on the minimum distance between the current position of the logistics vehicle and the road boundary point; Taking the obstacle center as the center of the circle and combining it with the preset radius, an obstacle risk field model is constructed; Determining a minimum distance between the current position of the logistics vehicle and an obstacle based on the obstacle risk field model and the current position of the logistics vehicle, and determining a risk of collision between the logistics vehicle and the obstacle based on the minimum distance between the current position of the logistics vehicle and the obstacle; Based on the risk of the logistics vehicle deviating from the lane and the risk of the logistics vehicle colliding with an obstacle, risk information of the logistics vehicle is obtained.
[0008] In one possible implementation, the deep deterministic policy gradient algorithm model includes: an actor network model and a critic network model; Based on the dynamic and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model, including: The dynamic and risk information of logistics vehicles in the simulation environment is used as the state input of the Actor Network Model, and the control instructions of the logistics vehicles are used as the action output of the Actor Network Model. The Actor Network Model is trained, and the Q values corresponding to the input and output of the Actor Network Model are evaluated based on the Critic Network Model. The Actor Network Model is then optimized based on the Q values. Optimize the Critic network model with the goal of minimizing the TD target error corresponding to the action output of the Actor network model; Based on the trained Actor network model and the optimized Critic network model, a preliminary training model is obtained.
[0009] In one possible implementation, the method for generating the optimization control model of the smart logistics vehicle further includes: The road boundary point data and obstacle data collected by the sensors on the logistics vehicle are stored in a set experience pool, and data is collected from the experience pool through random sampling to determine the dynamic information and risk information of random logistics vehicles; Based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated.
[0010] In one possible implementation, based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated, including: Determine rewards based on the current dynamic information of the logistics vehicle; Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the logistics vehicle optimization control model is updated.
[0011] In one possible implementation, based on the reward and the dynamic information and risk information of the random logistics vehicles, the logistics vehicle optimization control model is updated, including: Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the main network in the logistics vehicle optimization control model is gradually approximated in a soft update manner to update the logistics vehicle optimization control model.
[0012] In a second aspect, the present invention further provides a smart logistics vehicle optimization control method, comprising: Obtain dynamic information and risk information in the environment where the logistics vehicle is located, and input the dynamic information and risk information into the logistics vehicle optimization control model obtained by the above method to obtain the control instructions of the logistics vehicle, and based on the control instructions, control the logistics vehicle to travel according to the target path.
[0013] In a third aspect, the present invention further provides a device for generating an optimization control model for a smart logistics vehicle, comprising: The simulation training module is used to train the deep deterministic policy gradient algorithm model based on the dynamic information and risk information of logistics vehicles in the simulation environment to obtain a preliminary training model; A dynamic optimization module is used to use the dynamic information and risk information of logistics vehicles in the actual environment as the input of the preliminary training model, and the logistics vehicle control instructions as the output of the preliminary training model, to retrain the preliminary training model, and to adjust the dynamics of the logistics vehicle based on the logistics vehicle control instructions, and to optimize the preliminary training model based on the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0014] In a fourth aspect, the present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method for generating the smart logistics vehicle optimization control model as described in any one of the above items, or the steps of the above-mentioned smart logistics vehicle optimization control method.
[0015] The beneficial effect of adopting the above-mentioned implementation method is: the method, device and storage medium for generating the intelligent logistics vehicle optimization control model provided by the present invention first train the deep deterministic policy gradient algorithm model based on the dynamic information and risk information of the logistics vehicle in the simulation environment to obtain a preliminary training model, and then the logistics vehicle retrains the preliminary training model in the actual environment, and adjusts the dynamics of the logistics vehicle based on the logistics vehicle control instructions, and optimizes the preliminary training model according to the dynamic information feedback from the logistics vehicle to obtain the logistics vehicle optimization control model.
[0016] The logistics vehicle optimization control model is obtained by training with a deep deterministic policy gradient algorithm model. It can generate control instructions for controlling the logistics vehicle to travel along the target path based on the dynamic information and risk information in the environment where the logistics vehicle is located. There is no need for an additional path planning module to plan the path of the logistics vehicle and then generate control instructions based on the planned path, which can significantly improve decision-making efficiency. Moreover, the process of training the logistics vehicle optimization control model in the present invention includes not only a training process in a simulation environment, but also an online training process in an actual environment. It can be optimized according to the real-time changes of the actual environment, that is, according to the dynamic information fed back by the logistics vehicle. The logistics vehicle optimization control model thus obtained can better adapt to changes in complex environments, thereby ensuring the safe operation of the logistics vehicle. Therefore, the present invention can solve the technical problem that the existing logistics vehicle control method is difficult to meet the requirements of efficient and safe operation of logistics vehicles. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A flowchart of an embodiment of a method for generating an optimization control model for a smart logistics vehicle provided by the present invention; Figure 2 A flowchart of another embodiment of the method for generating an optimization control model for a smart logistics vehicle provided by the present invention; Figure 3 A flowchart of another embodiment of the method for generating an optimization control model for a smart logistics vehicle provided by the present invention; Figure 4 This is a schematic diagram of the risk field factor calculation provided by the present invention; Figure 5 This is a flow chart of an embodiment of the intelligent logistics vehicle optimization control method provided by the present invention; Figure 6 This is a functional block diagram of an embodiment of a device for generating an optimization control model for a smart logistics vehicle provided by the present invention; Figure 7 This is a schematic structural diagram of an embodiment of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0019] The following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0020] In the description of the embodiments of the present application, unless otherwise specified, “a plurality of” means two or more.
[0021] The terms "including" and "having" and any variations thereof in the embodiments of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product or device comprising a series of steps or modules is not necessarily limited to those steps or modules explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products or devices.
[0022] The naming or numbering of the steps in the embodiments of the present invention does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0023] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present invention. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute a separate or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0024] The present invention provides a method, device and storage medium for generating an optimization control model for a smart logistics vehicle, which are described below respectively.
[0025] like Figure 1 As shown, the present invention provides a method for generating an optimization control model of a smart logistics vehicle, comprising: S101. Based on the dynamic information and risk information of the logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model.
[0026] It can be understood that the deep deterministic policy gradient algorithm model, that is, the deep deterministic policy gradient (DDPG) reinforcement learning algorithm model, is referred to as the reinforcement learning algorithm model.
[0027] It should be noted that dynamic information refers to the movement and status of the logistics vehicle, such as its current position, speed, and heading angle. Risk information represents the risks that may occur during the logistics vehicle's operation, such as the risk of deviation from the shipping lane or the risk of collision with an obstacle.
[0028] S102, using the dynamic information and risk information of the logistics vehicle in the actual environment as the input of the preliminary training model, and the logistics vehicle control instructions as the output of the preliminary training model, retraining the preliminary training model, adjusting the dynamics of the logistics vehicle based on the logistics vehicle control instructions, and optimizing the preliminary training model based on the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0029] It can be understood that in order to solve the path planning and control problems of logistics vehicles in complex dynamic environments, the present invention constructs a risk field model to perceive the safety of the dynamic environment in real time, and combines it with the deep deterministic policy gradient (DDPG) reinforcement learning algorithm model. It can achieve efficient and safe path planning and precise control, and significantly improve the adaptability of logistics vehicles in dynamic environments.
[0030] The method provided by the present invention can be executed by an application on a terminal or server. The terminal can be a mobile phone or a computer. In a preferred embodiment, the terminal can be an onboard terminal used in a logistics vehicle, or a mobile terminal connected to the logistics vehicle. The server can be an edge server or a cloud server. Dynamic information and risk information about the logistics vehicle can be collected and processed by sensors on the logistics vehicle.
[0031] In some embodiments, the present invention provides an end-to-end online optimization control method for logistics vehicles based on security reinforcement learning, such as Figure 2As shown, the specific steps include: Constructing a risk field model: By real-time sensing the relative position between the vehicle, road boundaries, and obstacles, a comprehensive environmental risk assessment system is established to provide safety input for the reinforcement learning model. The risk field model includes modeling of the road boundary risk field and the obstacle risk field. By dynamically assessing the vehicle's driving status, the risk of the vehicle's surrounding environment is quantified in real time to obtain risk information.
[0032] Building a reinforcement learning model: A deep deterministic policy gradient (DDPG) approach is used to build a reinforcement learning model. The model takes the vehicle's dynamic state and risk field information as input and outputs continuous control actions. By learning interaction strategies in different environments, the model optimizes the vehicle's speed and direction adjustments, achieving end-to-end control of path planning.
[0033] Basic training in a simulation environment: Initial training of the reinforcement learning model is conducted in a simulation environment. Through interaction with the simulation environment, the agent gradually optimizes its control strategy and develops basic response capabilities to different risk conditions in complex scenarios.
[0034] Real-vehicle optimization training: This involves migrating the reinforcement learning model to the actual logistics vehicle operating environment and utilizing the vehicle's real-time perception data for online optimization training. By combining real-time feedback with dynamic environmental adjustments, the model's adaptability to the real environment is enhanced, enabling optimized path planning in dynamic environments.
[0035] Through the above steps, the method of the present invention can not only efficiently plan the path of logistics vehicles in dynamic and complex environments, but also significantly reduce safety risks during driving, while improving training efficiency and model adaptability in practical applications.
[0036] In some embodiments, the risk information of the logistics vehicle in the simulation environment is determined based on the driving state of the logistics vehicle and the relative positional relationship between the logistics vehicle and the road boundary and obstacles in the simulation environment; The risk information of logistics vehicles in the actual environment is determined based on the driving status of the logistics vehicles and the relative position relationship between the logistics vehicles and road boundaries and obstacles in the actual environment.
[0037] It's understandable that the road boundary risk field model and obstacle risk field model are constructed. The road boundary risk field model primarily assesses the risk of vehicles departing from their lanes by calculating the distance between the logistics vehicle and the road boundary, based on the relative position of the vehicle. This risk field can help the autonomous driving system monitor the relative position of the vehicle and the lane boundary in real time, thereby preventing accidents such as vehicle deviating from the lane or rolling over.
[0038] The obstacle risk field model assesses the collision risk between the vehicle and surrounding obstacles. Obstacles can be other vehicles, pedestrians, traffic facilities, or any other object that could interfere with driving. By calculating the distance to obstacles and assessing their risk, the system can make decisions about avoiding or slowing down.
[0039] In some embodiments, determining risk information of a logistics vehicle based on the driving status of the logistics vehicle and the relative positional relationship between the logistics vehicle and road boundaries and obstacles includes: The road boundary is modeled as a polygonal curve composed of discrete points, and the drivable area of the logistics vehicle is defined in the polygonal curve to obtain a road boundary risk field model; Determine the current position of the logistics vehicle based on the driving state of the logistics vehicle, determine the minimum distance between the current position of the logistics vehicle and the road boundary point based on the road boundary risk field model and the current position of the logistics vehicle, and determine the risk of the logistics vehicle deviating from the lane based on the minimum distance between the current position of the logistics vehicle and the road boundary point; Taking the obstacle center as the center of the circle and combining it with the preset radius, an obstacle risk field model is constructed; Determining a minimum distance between the current position of the logistics vehicle and an obstacle based on the obstacle risk field model and the current position of the logistics vehicle, and determining a risk of collision between the logistics vehicle and the obstacle based on the minimum distance between the current position of the logistics vehicle and the obstacle; Based on the risk of the logistics vehicle deviating from the lane and the risk of the logistics vehicle colliding with an obstacle, risk information of the logistics vehicle is obtained.
[0040] It can be understood that the road boundary is modeled as a polygonal curve composed of a series of discrete points to define the drivable area of the vehicle, thereby obtaining a road boundary risk field model.
[0041] The proximity between the vehicle and the boundary is evaluated by calculating the minimum distance between the vehicle's current position and the road boundary point.
[0042] Based on the minimum distance between the vehicle and the road boundary, a road boundary risk factor is defined. The road boundary risk factor can represent the risk of the logistics vehicle deviating from the lane. When the vehicle approaches the boundary, the risk factor increases rapidly, indicating a high risk of deviating from the lane.
[0043] In this embodiment, a road boundary risk safety constraint condition can be set to ensure that the minimum distance between the vehicle and the road boundary is not less than a preset safety distance threshold to prevent the vehicle from deviating from the safe driving area.
[0044] The obstacle is modeled as a circle, its center point and radius are defined, and the obstacle risk field model is constructed.
[0045] The minimum safe distance between the vehicle and the obstacle is calculated based on the distance between the vehicle's current position and the obstacle's center point.
[0046] The obstacle risk factor is defined based on the distance between the vehicle and the obstacle. When the vehicle approaches the obstacle, the risk factor increases rapidly, indicating that the vehicle is at risk of collision.
[0047] Obstacle risk safety constraint: Ensures that the minimum distance between the vehicle and the obstacle is not less than the preset safety distance threshold, that is, the minimum safety distance, to avoid collision risk.
[0048] In this embodiment, the road boundary risk factor and the obstacle risk factor constitute the risk information of the logistics vehicle.
[0049] In some embodiments, a deep deterministic policy gradient algorithm model includes: an Actor network model and a Critic network model; Based on the dynamic and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model, including: The dynamic and risk information of logistics vehicles in the simulation environment is used as the state input of the Actor Network Model, and the control instructions of the logistics vehicles are used as the action output of the Actor Network Model. The Actor Network Model is trained, and the Q values corresponding to the input and output of the Actor Network Model are evaluated based on the Critic Network Model. The Actor Network Model is then optimized based on the Q values. Optimize the Critic network model with the goal of minimizing the TD target error corresponding to the action output of the Actor network model; Based on the trained Actor network model and the optimized Critic network model, a preliminary training model is obtained.
[0050] As you can understand, TD target error (Temporal Difference Error) is an important concept in reinforcement learning that is used to evaluate the accuracy of prediction models. TD error measures the performance of a model by comparing the difference between the model's predicted values and the actual values.
[0051] Actor network: inputs vehicle state information and outputs continuous actions, including speed increments and steering angle increments; Critic network: Provides guidance for the optimization of the Actor network by evaluating the Q values of states and actions.
[0052] In some embodiments, the method for generating an optimization control model for a smart logistics vehicle further includes: The road boundary point data and obstacle data collected by the sensors on the logistics vehicle are stored in a set experience pool, and data is collected from the experience pool through random sampling to determine the dynamic information and risk information of random logistics vehicles; Based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated.
[0053] Furthermore, based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated, including: Determine rewards based on the current dynamic information of the logistics vehicle; Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the logistics vehicle optimization control model is updated.
[0054] It is understandable that the reward is calculated based on the current state and action of the logistics vehicle , as a reference for strategy optimization. The experience pool stores data collected during actual vehicle operation. And update the Actor and Critic networks in the logistics vehicle optimization control model through random sampling. Updating the network by random sampling can improve the generalization ability of the model.
[0055] In some embodiments, based on the reward and the dynamic information and risk information of the random logistics vehicle, updating the logistics vehicle optimization control model includes: Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the main network in the logistics vehicle optimization control model is gradually approximated in a soft update manner to update the logistics vehicle optimization control model.
[0056] It is understandable that this embodiment uses a soft update method to optimize the model, which can ensure that the model optimization process is more stable.
[0057] In some embodiments, reference Figure 3 As shown, the method provided by the present invention includes: 1. Construction of risk field model (1) Road boundary risk areas refer to Figure 4 As shown in Figure 1, the road boundary risk field is used to assess the risk of vehicles leaving the lane in real time, which is accomplished through the following steps: Modeling: Modeling the road boundary as a polygonal curve composed of a series of discrete points to define the drivable area of the vehicle; Minimum distance calculation: Evaluate the proximity between the vehicle and the boundary by calculating the minimum distance between the vehicle's current position and the road boundary point; Definition of risk factor (i.e., risk information): A road boundary risk factor is defined based on the minimum distance between the vehicle and the road boundary. When the vehicle approaches the boundary, the risk factor increases rapidly, indicating a high risk of lane deviation.
[0058] Road boundary risk safety constraint: Ensure that the minimum distance between the vehicle and the road boundary is not less than the preset safety distance threshold to prevent the vehicle from deviating from the safe driving area.
[0059] (2) Obstacle risk area The obstacle risk field is used to assess the collision risk between the vehicle and the obstacle. This is accomplished through the following steps: Modeling: Model the obstacle as a circle, defining its center point and radius; Minimum distance calculation: Calculate the minimum safe distance between the vehicle and the obstacle based on the distance between the vehicle's current position and the obstacle's center point; Risk factor definition: The obstacle risk factor is defined based on the distance between the vehicle and the obstacle. When the vehicle approaches the obstacle, the risk factor increases rapidly, indicating that the vehicle is at risk of collision.
[0060] Obstacle risk safety constraint: Ensures that the minimum distance between the vehicle and the obstacle is not less than the preset safety distance threshold to avoid collision risk.
[0061] (3) Comprehensive risk assessment The risk information of the road boundary risk field and the obstacle risk field is integrated, combined with the road risk factor and the obstacle risk factor. The importance of the two is balanced by dynamically adjusting the weight factor to finally generate a comprehensive risk factor, that is, the final risk information, which provides environmental safety assessment input for the reinforcement learning model.
[0062] 2. DDPG-based reinforcement learning model (1) State space The state space is used to describe the dynamic information of the vehicle and the risk information of the surrounding environment as the input of the reinforcement learning model. The state space includes: Vehicle dynamic information, such as current position, speed, and heading angle; Road boundary risk factors and obstacle risk factors from the risk field model.
[0063] (2) Action Space The action space is the output of the reinforcement learning model and represents the vehicle control instructions. The action space adopts a continuous control mode, including: Speed increment: used to adjust the vehicle's speed; Steering angle increment: used to adjust the vehicle's driving direction.
[0064] (3) Reward Function The reward function is used to guide the reinforcement learning model to achieve the goal in path planning while ensuring the safety and stability of the vehicle. The reward function consists of the following parts: Mission reward: Encourage vehicles to reach the target location quickly; Safety penalty: Based on the risk factor output by the risk field, penalties are imposed for actions such as leaving the lane or approaching obstacles; Stability Reward: Encourages the vehicle to maintain stable driving and avoid sudden acceleration, sudden braking and sharp steering.
[0065] (4) Actor-Critic Network Construction The present invention adopts the deep deterministic policy gradient (DDPG) method to build a reinforcement learning model, including: Actor network: It inputs vehicle state information (vehicle dynamic information, i.e., vehicle action and state information) and outputs continuous actions, including speed increments and steering angle increments. Critic network: Provides guidance for the optimization of the Actor network by evaluating the Q values of states and actions.
[0066] (5) Target Network and Experience Replay To improve the stability of training, the present invention introduces a target network and experience replay mechanism: Target network: gradually approaches the main network parameters through a soft update mechanism to reduce the instability of model training; Experience replay: Stores historical data on the interaction between the vehicle and the environment, and updates the model through random sampling to avoid overfitting problems caused by sample correlation.
[0067] 3. Model online optimization training (1) Simulation training phase The initial training of the reinforcement learning model is completed through a simulation environment. The intelligent agent gradually optimizes the control strategy by interacting with the virtual scene, forming a basic ability to cope with different dynamic environments.
[0068] (2) Real-car optimization training phase The model trained through simulation is deployed to the actual operating environment. Through real-time perception of environmental data and combined with online optimization mechanisms, the model's adaptability and decision-making ability in dynamic and complex scenarios are continuously improved.
[0069] (3) Online optimization mechanism Online optimization is achieved through the following core steps: Real-time perception: Sensors detect dynamic changes in road boundaries and obstacles in real time and update the vehicle's comprehensive risk factor; Dynamic decision-making and action execution: The Actor network generates control actions based on real-time status, and the vehicle executes the control instructions to adjust the speed and direction of the logistics vehicle; Strategy optimization: Combined with real-time reward feedback, the actor and critic networks are updated online to continuously improve the performance of the model in real-world scenarios.
[0070] (1) Construction of risk field model Step 11: Road Boundary Risk Area The road boundary risk field primarily assesses the risk of a vehicle departing from its lane by calculating the distance between the vehicle and the road boundary. This risk field helps the autonomous driving system monitor the relative position of the vehicle and lane boundaries in real time, thereby preventing accidents such as vehicle departure from the lane and rollover.
[0071] Road boundary modeling: Assuming that the road boundary consists of a series of discrete points, the boundary can be approximately modeled as a polygonal curve or spline curve. The set of boundary points is: ,in For the Coordinates of road boundary points; is the total number of boundary points; The feasible area for the vehicle to travel is defined. The current position of the vehicle at a certain moment can be expressed as: ,in is the coordinate of the vehicle's center of mass.
[0072] Calculation of the minimum distance between the vehicle and the road boundary: the minimum distance between the vehicle and the road boundary Defined as vehicle position To all boundaries The minimum value of the Euclidean distance, that is ,Right now .
[0073] Definition of road risk factor: To quantify the risk of lane departure, a risk factor can be defined , this factor is inversely proportional to the distance between the vehicle and the boundary. Let ,in For safe distance, if Beyond this distance, the risk factor is 0; is the minimum distance between the vehicle and the boundary; is the adjustment parameter of the risk factor, which is used to control the rate of risk increase. When the risk factor is small, It will increase rapidly, indicating that the vehicle is at great risk of leaving the lane.
[0074] Generate safety constraints: Based on the minimum distance between the vehicle and the road boundary, generate the following safety constraints: ,in, The safe distance threshold between the vehicle and the road boundary. If the distance between the vehicle and the road boundary is less than this threshold, the system will generate a warning and take appropriate evasive measures.
[0075] Step 12: Obstacle risk field The obstacle risk field assesses the collision risk between the vehicle and surrounding obstacles. Obstacles can be other vehicles, pedestrians, traffic facilities, or any other objects that could interfere with driving. By calculating the distance to obstacles and assessing their risk, the system can make decisions about avoiding or slowing down.
[0076] Obstacle modeling: Assuming obstacles It is circular, and its center point coordinates , the radius is The obstacle set can be expressed as: ,in For the The center position of each obstacle; For the The radius of the obstacle; is the total number of obstacles.
[0077] Calculation of the minimum distance between the vehicle and the obstacle: The current position of the vehicle is , which is related to obstacles The minimum distance ,in is the Euclidean distance between the vehicle and the obstacle center; if Indicates that the vehicle has entered the obstacle range. .
[0078] Definition of obstacle risk factor: In order to quantify the risk of collision with obstacles, an obstacle risk factor can be defined in: The safe distance from obstacles; is the adjustment parameter of the risk factor; For vehicles and obstacles The total obstacle risk factor is the weighted sum of all obstacle risk factors: .
[0079] Generate safety constraints: Based on the minimum distance between the vehicle and the obstacle, generate the following safety constraints: .in, is the safety distance threshold of the obstacle. , the system will adjust the control action to avoid the obstacle.
[0080] Step 13: Comprehensive risk assessment In order to evaluate the overall safety of a vehicle, a comprehensive risk factor can be defined , which combines the road boundary risk and obstacle risk, in and is a weighting factor used to adjust the importance of road boundary risk and obstacle risk. The result of the comprehensive risk assessment is the final risk information.
[0081] Since the vehicle status (such as speed, direction, etc.) and environmental risks will change dynamically over time, the comprehensive risk factor should also be updated in real time: , through real-time calculation , which can provide dynamic risk input for reinforcement learning models.
[0082] End-to-end reinforcement learning model: Step 21: Building a reinforcement learning framework The reinforcement learning model is based on the following basic definitions: environment : The actual driving environment of the logistics vehicle, including dynamic scenes composed of road boundaries, obstacles, lane lines, etc.; : Logistics vehicle control system, selecting actions through strategy network Interaction with the environment; state : Description of the current environment and vehicle status, including the output of the risk field model; action The actions that the agent can take, such as accelerating, decelerating, turning, etc.; rewards : The score of the environment's feedback on the actions taken by the agent, which is used to optimize the strategy.
[0083] Reinforcement learning goal: find the optimal policy , so that each state Next action Can maximize the cumulative reward, that is ,in is a discount factor that measures the importance of future rewards.
[0084] Step 22: Define the state space
[0085] State Space It is the input of the agent's perception of the environment and its own state, including the dynamic information of the logistics vehicle and the output of the risk field model. It is specifically defined as: ,in is the dynamic state of the vehicle, , is the current position coordinate of the vehicle; is the vehicle speed; is the vehicle heading angle. is the risk factor of the road boundary, which comes from the risk field model. is the risk factor of the obstacle, which comes from the risk field model. Represents other environmental information, such as the distance and angle of obstacles.
[0086] Step 23: Define the action space
[0087] Action Space The output of the reinforcement learning model defines the control instructions that the logistics vehicle can take at each time step. A continuous action space provides more fine-grained control instructions and is suitable for complex dynamic environments, such as high-speed driving, complex obstacle avoidance, and precise path tracking. Given the complexities of logistics scenarios, a continuous action space is constructed.
[0088] In the continuous action space, the action vector can be defined as the vehicle's velocity and steering angle increments: ,in is the speed increment, which represents the speed adjustment value of the vehicle at the current time step; is the steering angle increment, which indicates the direction adjustment value of the vehicle; and Indicates the upper and lower limits of speed adjustment; and The upper and lower limits for steering angle adjustment.
[0089] The motion control of the vehicle can be described by the following dynamic formula: Velocity update ,in: is the current speed; Velocity increment for the current time step; position update ,in: is the current position of the vehicle; is the current speed; is the current heading angle; is the time step; steering angle update .
[0090] The action space should satisfy the constraints of vehicle physical characteristics and safety: speed constraint ; Steering angle constraint ; Safety constraints .
[0091] Step 24. Reward Function Build Reward Function The design goals are: to encourage vehicles to complete logistics tasks as quickly as possible (task goal reward); to punish unsafe behaviors such as lane deviation and collision with obstacles (safety penalty); to encourage vehicles to maintain stable driving and reduce sudden acceleration, sudden braking and sharp steering (smoothness reward); and to dynamically adjust the reward value according to the real-time environment to adapt to scenarios of different complexities. Taking all the above goals into consideration, the reward function The overall form of can be defined as: ,in Rewards for mission objectives; penalties for lane departure; Penalty for obstacle collision risk; Penalty for driving smoothness; 、 、 are the weights of lane risk, obstacle risk, and smoothness penalty, respectively, used to adjust the importance of each part.
[0092] Mission Target Reward Used to encourage vehicles to reach their destinations along the scheduled routes and complete logistics transportation tasks. , is the current position of the vehicle; is the target position; ϵ is the distance threshold for the target to be reached; It is a fixed reward for reaching the goal.
[0093] penalties for lane departure; For more information on obstacle collision risk penalties, see Risk Field Construction.
[0094] Driving stability penalty Used to encourage the vehicle to maintain stable driving and avoid sudden acceleration, sudden braking and sharp steering. ,in: is the velocity increment; is the steering angle increment; is the smoothness penalty coefficient. When the vehicle's speed or direction changes significantly, the smoothness penalty increases, encouraging the vehicle to choose smoother control commands.
[0095] Based on the above, the comprehensive reward function is: ,in 、 、 is a weight factor used to balance the task objective reward with the penalty for safety and stability.
[0096] Step 25: Actor network construction The goal of the Actor network is to generate deterministic policies ; Its input is the current state , including vehicle dynamic state (speed, position, heading angle, etc.), risk factors output by the risk field model and and unstable factors .
[0097] Its output is continuous action , which are used to adjust the vehicle's speed and steering angle respectively.
[0098] Speed Constraint ; Steering angle constraint ; Safety constraints etc. as activation functions.
[0099] ,in Represents the hidden layer operation of the neural network; is an activation function that constrains the output to the interval [−1, 1].
[0100] The goal of an actor is to find a strategy so that the state-action value maximize: .
[0101] Critic Network provides The evaluation value is used to guide the parameter update of the Actor network.
[0102] Step 26: Critic Network The goal of the Critic network is to estimate the state-action value function , used to evaluate the action quality output by the Actor network and guide the optimization of the Actor network through gradient backpropagation.
[0103] Its input is the state and actions , the output is the state-action value , indicating that in the state Next action The expected cumulative rewards that can be obtained.
[0104] in, .
[0105] Minimize TD error parameter update ,in .
[0106] Step 27: Target network update The target network is one of the core enhancement mechanisms of DDPG, which is used to improve the stability of training. The target network update rule is: ,in: Represents the parameters of the target network; Represents the parameters of the main network; τ∈[0,1] is the soft update coefficient, which is usually small.
[0107] Step 28: Experience Recovery Experience Pool Store the agent's interaction data at each time step: ,in: is the current state; For the current action; For immediate rewards; For the next state.
[0108] At each training step, a batch is randomly sampled from the experience pool. , use these samples to update the Actor and Critic networks.
[0109] Model online optimization training: Simulation training phase (i.e. basic training): The simulation phase is the first step in online optimization training. Its goal is to use a virtual environment to conduct preliminary training of the reinforcement learning model. During this phase, the agent gradually optimizes its strategy through interaction with the simulation environment and develops basic response capabilities to various risk scenarios.
[0110] Step 31: Construction of simulation environment Risk field environment simulation: Within the simulation environment, dynamic scenarios are constructed using road boundary risk field and obstacle risk field models. The environment includes road boundary points and obstacles. Road boundary points define the vehicle's driving area and lane departure risk. Obstacles include dynamic and static objects such as other vehicles, pedestrians, and roadside facilities, enabling simulation of diverse traffic scenarios.
[0111] Logistics mission simulation: Vehicles need to travel from a starting point to a designated destination. The scenario may include dynamic obstacles, complex intersections, and traffic restrictions, simulating the complexity of real-world logistics tasks.
[0112] Step 32: Simulation training of DDPG model State input: current state , including vehicle dynamic state and risk factors. Action input: action Generated by the Actor network, it controls the vehicle's velocity increment and direction adjustment. Reward function design: , the reward function guides the agent to maximize safety and stability while completing the task.
[0113] Actor network training: The actions generated by the Actor network are evaluated by the Critic network, and the strategy is optimized to maximize the Q value: .
[0114] Critic network training: Update by minimizing TD target error: .
[0115] Target network and experience replay: The target network uses soft updates to stabilize model training; the experience pool stores interaction data and randomly samples and updates the network to improve generalization capabilities.
[0116] Actual vehicle optimization stage: online optimization training.
[0117] Step 33: Model migration and initial deployment Simulation model migration: The Actor and Critic network models trained in simulation are directly applied to the real vehicle environment as the initial strategy.
[0118] Adaptive adjustment: In the actual environment, through real-time perception of vehicle status and risk field information , the model makes adaptive adjustments to the environment.
[0119] Step 34: Real-time online optimization mechanism During actual vehicle operation, the strategy is optimized using an online update method based on reinforcement learning. The specific mechanism includes the following steps: Real-time perception and status update: Use vehicle sensors (such as lidar, cameras, etc.) to obtain road boundary points and obstacle data and calculate the current risk factor Update status in real time based on the vehicle's current location and dynamic information .
[0120] Dynamic decision-making and action execution: Actor network based on the current state Output Action The action is executed in the real environment and the vehicle adjusts according to the control instructions.
[0121] Reward feedback and online update: Calculate rewards based on current state and action , as a reference for strategy optimization. The experience pool stores data collected during actual vehicle operation. , and update the Actor and Critic networks by random sampling.
[0122] Update of target network: During the online optimization process, the target network parameters are gradually approached to the main network in a soft update manner to ensure the stability of the optimization.
[0123] Step 35: Dynamically adjust strategy During the online optimization phase, the model's optimization strategy is dynamically adjusted based on the actual needs and operating environment of the logistics vehicle.
[0124] Risk factor weight adjustment: In high-density traffic environments, increase the obstacle risk factor weights to prioritize collision avoidance.
[0125] Reduce lane risk factors in open environments The weight of the vehicle is given priority to optimize driving efficiency.
[0126] Scenario-adaptive training: The system automatically adjusts the parameters of the reward function based on the characteristics of the environment. For example, it increases the reward weight for smooth driving in complex scenarios.
[0127] Real-time update of the experience pool: Dynamically add the currently collected environment and decision-making data to the experience pool, combine it with historical data for intensive training, and further optimize the model.
[0128] Compared with the prior art, the present invention has the following advantages: (1) Real-time strategy update and optimization. Through online optimization, the present invention enables logistics vehicles to perceive environmental changes in real time and update strategies immediately. In the event of sudden obstacles, driving parameters can be quickly adjusted to effectively avoid collisions, significantly improving dynamic environmental adaptability and safety, and breaking through the limitations of traditional methods that are not real-time enough.
[0129] (2) Enhance model adaptability and decision-making accuracy. After migrating the simulation training model to the real vehicle, use real-time data for online optimization to allow the model to continuously learn and adapt to complex actual scenarios, accurately assess risks, plan paths, and execute actions. Compared with models trained only by simulation, the effectiveness and reliability in actual applications are greatly improved.
[0130] (3) Improve training efficiency and resource utilization flexibility. Online optimization eliminates the need for extensive offline retraining when the environment changes, saving computing resources and time. The model can continuously learn and improve performance as it runs, achieving synchronization between training and application. At the same time, it enhances the versatility and portability of the model. A single model can be flexibly deployed in different logistics parks, quickly adapting to new environments, facilitating the widespread application and promotion of the technology, and having significant economic and social benefits.
[0131] like Figure 5 As shown, the present invention also provides a smart logistics vehicle optimization control method, including: S501. Obtain dynamic information and risk information in the environment where the logistics vehicle is located, and input the dynamic information and risk information into the logistics vehicle optimization control model obtained by the above method to obtain control instructions for the logistics vehicle, and based on the control instructions, control the logistics vehicle to travel along the target path.
[0132] like Figure 6 As shown, the present invention also provides a device 600 for generating an optimization control model for a smart logistics vehicle, comprising: A simulation training module 601 is used to train a deep deterministic policy gradient algorithm model based on the dynamic information and risk information of logistics vehicles in a simulation environment to obtain a preliminary training model; Dynamic optimization module 602, configured to use the dynamic information and risk information of logistics vehicles in the actual environment as input to the preliminary training model, and the logistics vehicle control instructions as output of the preliminary training model, retrain the preliminary training model, adjust the dynamics of the logistics vehicles based on the logistics vehicle control instructions, and optimize the preliminary training model based on the dynamic information fed back by the logistics vehicles to obtain an optimized control model for the logistics vehicles; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0133] The device for generating the optimization control model of a smart logistics vehicle provided in the above embodiment can implement the technical solution described in the embodiment of the method for generating the optimization control model of a smart logistics vehicle. The specific implementation principles of the above modules or units can be found in the corresponding content in the embodiment of the method for generating the optimization control model of a smart logistics vehicle, and will not be repeated here.
[0134] like Figure 7 As shown, the present invention also provides an electronic device 700. The electronic device 700 includes a processor 701, a memory 702 and a display 703. Figure 7 Only some of the components of the electronic device 700 are shown, but it should be understood that it is not required to implement all of the shown components, and more or fewer components may be implemented instead.
[0135] In some embodiments, the memory 702 may be an internal storage unit of the electronic device 700, such as a hard disk or memory of the electronic device 700. In other embodiments, the memory 702 may also be an external storage device of the electronic device 700, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 700.
[0136] Furthermore, the memory 702 may include both an internal storage unit of the electronic device 700 and an external storage device. The memory 702 is used to store application software installed in the electronic device 700 and various data.
[0137] In some embodiments, the processor 701 can be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run the program code or process data stored in the memory 702, such as the method for generating the smart logistics vehicle optimization control model or the smart logistics vehicle optimization control method in the present invention.
[0138] In some embodiments, display 703 can be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. Display 703 is used to display information on electronic device 700 and to display a visual user interface. Components 701-703 of electronic device 700 communicate with each other via a system bus.
[0139] In some embodiments of the present invention, when the processor 701 executes the program for generating the intelligent logistics vehicle optimization control model in the memory 702, the following steps may be implemented: Based on the dynamic information and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model; The dynamic information and risk information of the logistics vehicle in the actual environment are used as the input of the preliminary training model, and the logistics vehicle control instructions are used as the output of the preliminary training model. The preliminary training model is retrained, and the dynamics of the logistics vehicle are adjusted based on the logistics vehicle control instructions. The preliminary training model is optimized according to the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0140] When the processor 701 executes the smart logistics vehicle optimization control program in the memory 702, the following steps can be implemented: Obtain dynamic information and risk information in the environment where the logistics vehicle is located, and input the dynamic information and risk information into the logistics vehicle optimization control model obtained by the above method to obtain the control instructions of the logistics vehicle, and based on the control instructions, control the logistics vehicle to travel according to the target path.
[0141] It should be understood that when the processor 701 executes the generation program / smart logistics vehicle optimization control model in the memory 702, in addition to the above functions, it can also realize other functions. For details, please refer to the description of the corresponding method embodiment above.
[0142] Furthermore, the embodiments of the present invention do not specifically limit the type of electronic device 700 mentioned. Electronic device 700 may be a portable electronic device such as a mobile phone, tablet computer, personal digital assistant (PDA), wearable device, or laptop computer. Exemplary embodiments of portable electronic devices include, but are not limited to, portable electronic devices running iOS, Android, Microsoft, or other operating systems. The portable electronic devices mentioned above may also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in other embodiments of the present invention, electronic device 700 may not be a portable electronic device, but rather a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0143] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for generating an optimization control model for a smart logistics vehicle provided by the above methods is implemented. The method includes: Based on the dynamic information and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model; The dynamic information and risk information of the logistics vehicle in the actual environment are used as the input of the preliminary training model, and the logistics vehicle control instructions are used as the output of the preliminary training model. The preliminary training model is retrained, and the dynamics of the logistics vehicle are adjusted based on the logistics vehicle control instructions. The preliminary training model is optimized according to the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
[0144] When the computer program is executed by the processor, it is implemented to perform the smart logistics vehicle optimization control method provided by the above methods, which includes: Obtain dynamic information and risk information in the environment where the logistics vehicle is located, and input the dynamic information and risk information into the logistics vehicle optimization control model obtained by the above method to obtain the control instructions of the logistics vehicle, and based on the control instructions, control the logistics vehicle to travel according to the target path.
[0145] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0146] The above is a detailed introduction to the method, device and storage medium for generating the optimization control model of the smart logistics vehicle provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A method for generating an optimization control model for a smart logistics vehicle, characterized in that: include: Based on the dynamic information and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model; The dynamic information and risk information of the logistics vehicle in the actual environment are used as the input of the preliminary training model, and the logistics vehicle control instructions are used as the output of the preliminary training model. The preliminary training model is retrained, and the dynamics of the logistics vehicle are adjusted based on the logistics vehicle control instructions. The preliminary training model is optimized according to the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
2. The method for generating an optimization control model for a smart logistics vehicle according to claim 1 is characterized in that: The risk information of the logistics vehicle in the simulation environment is determined based on the driving status of the logistics vehicle and the relative position relationship between the logistics vehicle and the road boundaries and obstacles in the simulation environment; The risk information of logistics vehicles in the actual environment is determined based on the driving status of the logistics vehicles and the relative position relationship between the logistics vehicles and road boundaries and obstacles in the actual environment.
3. The method for generating an optimization control model for a smart logistics vehicle according to claim 2 is characterized in that: The risk information of the logistics vehicle is determined based on the driving status of the logistics vehicle and the relative position relationship between the logistics vehicle and road boundaries and obstacles, including: The road boundary is modeled as a polygonal curve composed of discrete points, and the drivable area of the logistics vehicle is defined in the polygonal curve to obtain a road boundary risk field model; Determine the current position of the logistics vehicle based on the driving state of the logistics vehicle, determine the minimum distance between the current position of the logistics vehicle and the road boundary point based on the road boundary risk field model and the current position of the logistics vehicle, and determine the risk of the logistics vehicle deviating from the lane based on the minimum distance between the current position of the logistics vehicle and the road boundary point; Taking the obstacle center as the center of the circle and combining it with the preset radius, an obstacle risk field model is constructed; Determining a minimum distance between the current position of the logistics vehicle and an obstacle based on the obstacle risk field model and the current position of the logistics vehicle, and determining a risk of collision between the logistics vehicle and the obstacle based on the minimum distance between the current position of the logistics vehicle and the obstacle; Based on the risk of the logistics vehicle deviating from the lane and the risk of the logistics vehicle colliding with an obstacle, risk information of the logistics vehicle is obtained.
4. The method for generating an optimization control model for a smart logistics vehicle according to claim 1 is characterized in that: Deep deterministic policy gradient algorithm model, including: Actor network model and Critic network model; Based on the dynamic and risk information of logistics vehicles in the simulation environment, a deep deterministic policy gradient algorithm model is trained to obtain a preliminary training model, including: The dynamic and risk information of logistics vehicles in the simulation environment is used as the state input of the Actor Network Model, and the control instructions of the logistics vehicles are used as the action output of the Actor Network Model. The Actor Network Model is trained, and the Q values corresponding to the input and output of the Actor Network Model are evaluated based on the Critic Network Model. The Actor Network Model is then optimized based on the Q values. Optimize the Critic network model with the goal of minimizing the TD target error corresponding to the action output of the Actor network model; Based on the trained Actor network model and the optimized Critic network model, a preliminary training model is obtained.
5. The method for generating an optimization control model for a smart logistics vehicle according to any one of claims 1 to 4, characterized in that: Also includes: The road boundary point data and obstacle data collected by the sensors on the logistics vehicle are stored in a set experience pool, and data is collected from the experience pool through random sampling to determine the dynamic information and risk information of random logistics vehicles; Based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated.
6. The method for generating an optimization control model for a smart logistics vehicle according to claim 5 is characterized in that: Based on the dynamic information and risk information of random logistics vehicles, the logistics vehicle optimization control model is updated, including: Determine rewards based on the current dynamic information of the logistics vehicle; Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the logistics vehicle optimization control model is updated.
7. The method for generating an optimization control model for a smart logistics vehicle according to claim 6 is characterized in that: Based on the reward and the dynamic and risk information of the random logistics vehicles, the logistics vehicle optimization control model is updated, including: Based on the reward, as well as the dynamic information and risk information of the random logistics vehicles, the main network in the logistics vehicle optimization control model is gradually approximated in a soft update manner to update the logistics vehicle optimization control model.
8. A smart logistics vehicle optimization control method, characterized in that: include: Obtain dynamic information and risk information in the environment where the logistics vehicle is located, and input the dynamic information and the risk information into the logistics vehicle optimization control model obtained by the method described in any one of claims 1-7 to obtain the control instructions of the logistics vehicle, and based on the control instructions, control the logistics vehicle to travel according to the target path.
9. A device for generating an optimization control model for a smart logistics vehicle, characterized in that: include: The simulation training module is used to train the deep deterministic policy gradient algorithm model based on the dynamic information and risk information of logistics vehicles in the simulation environment to obtain a preliminary training model; A dynamic optimization module is used to use the dynamic information and risk information of logistics vehicles in the actual environment as the input of the preliminary training model, and the logistics vehicle control instructions as the output of the preliminary training model, to retrain the preliminary training model, and to adjust the dynamics of the logistics vehicle based on the logistics vehicle control instructions, and to optimize the preliminary training model based on the dynamic information fed back by the logistics vehicle to obtain an optimized control model for the logistics vehicle; Among them, the logistics vehicle optimization control model is used to generate control instructions for controlling the logistics vehicle to travel along the target path based on dynamic information and risk information in the environment where the logistics vehicle is located.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for generating the intelligent logistics vehicle optimization control model as described in any one of claims 1 to 7, or the steps of the intelligent logistics vehicle optimization control method as described in claim 8 are implemented.