Delivery trolley and control method
By adopting technical solutions of track adaptive control and reinforcement learning optimization on tracked trolleys, the problem of insufficient traffic capacity and poor environmental adaptability of the trolley on complex terrain is solved, and more stable and efficient driving is achieved.
Patent Information
- Application Number
- CN202510303674.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-10
AI Technical Summary
The tracked trolleys have insufficient traffic capacity and poor environmental adaptability on complex terrains in the laboratory, especially under ramps, slippery grounds and diverse ground materials.
A delivery car is designed, using a technical solution that combines track adaptive control and reinforcement learning optimization. By adjusting the track spacing and height in real time, combining the use of multi-axis robotic arms and electromagnets, adaptive adjustment of different ground conditions is achieved, and path decisions are optimized through deep Q learning algorithms.
It significantly improves the driving stability and decision-making ability of the car in complex environments, solves the problem of poor adaptability of traditional fixed control parameters on heterogeneous terrain, and improves the passing capacity and traction efficiency.
Smart Images

Figure CN120117059A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent vehicles, and specifically to a delivery vehicle and a control method therefor. Background Art
[0002] Automated tracked vehicles are increasingly used in the transportation of laboratory samples. Especially in scenarios involving precision instruments and high safety requirements, higher requirements are put forward for the environmental adaptability of the vehicles. Although the laboratory environment is relatively enclosed, there are still complex terrains such as slopes, slippery floors, and floor gaps. Most traditional tracked vehicles adopt fixed control parameters and operate relatively stably in a single environment. However, when facing complex terrain changes, it is difficult to achieve precise adaptation, and the passing ability is greatly limited.
[0003] The setting of the tracked tension directly affects the passing ability of the vehicle on different ground surfaces. Most existing vehicles adopt fixed tension control. On hard ground, it may be insufficient, resulting in a decrease in the adhesion of the track and affecting the driving stability. On soft ground or slippery areas, the fixed tension is likely to cause the track to slip, causing the vehicle to lose traction. The ground materials in the laboratory are diverse, and there may be anti-static floor mats or buffer rubber mats in some areas. Traditional control methods are difficult to balance the requirements of different ground surfaces, resulting in the restricted passage of tracked vehicles in some areas and affecting the transportation efficiency.
[0004] Secondly, most existing tracked drive methods adopt constant torque or constant power output. On ground with low adhesion, the track is prone to idling, resulting in energy waste and at the same time reducing the passing ability. When the vehicle travels on a slope or terrain with a slight height difference, the fixed drive method may not be able to provide sufficient traction, affecting the uphill ability and even causing slipping and backward movement. Some solutions adopt simple power compensation, but lack feedback adjustment of the real-time slip rate and cannot achieve precise drive control in complex terrain, resulting in insufficient passing ability of tracked vehicles in the changing terrain of the laboratory. Summary of the Invention
[0005] Aiming at the deficiencies of the prior art, the present invention provides a delivery vehicle and a control method therefor, which solve the problems of insufficient passing ability and poor environmental adaptability of tracked vehicles on complex terrain in the laboratory.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: A delivery vehicle includes a base, a body is installed on the top of the base, a taking component is installed inside the body, a frame is installed at the bottom of the base, side frames are installed on the outside of the frame through a width adjustment component, a track group is installed on the outside of the side frames, and a height adjustment component is installed on the top of the frame.
[0007] Preferably, the width adjustment component includes an electric push rod III, which is fixedly connected to the outside of the vehicle frame. The output end of the vehicle frame is fixedly connected to the inside of the side frame. A motor is fixedly connected to the end of the vehicle frame. The output end of the motor is fixedly connected to a polygonal tube sleeve. A polygonal column is slidably connected inside the polygonal tube sleeve. The other end of the polygonal column is fixedly connected to the middle of the crawler group drive wheel.
[0008] Preferably, the height adjustment component includes an electric push rod I. One end of the electric push rod I is rotatably connected to the bottom of the machine base, and the other end of the electric push rod I is rotatably connected to the top of the vehicle frame.
[0009] Preferably, the height adjustment component further includes a rotating beam. One end of the rotating beam is rotatably connected to the bottom of the machine base, and the other end of the rotating beam is rotatably connected to the top of the vehicle frame. An electric push rod II is rotatably connected to the middle of the rotating beam, and the other end of the electric push rod II is rotatably connected to the bottom of the machine base.
[0010] Preferably, the picking component includes a base, which is fixedly connected to the inside of the machine body. A multi-axis robotic arm is installed on the top of the base. A fixing ring is installed at the free end of the multi-axis robotic arm. A telescopic cylinder II is fixedly connected to the middle of the fixing ring. The output end of the telescopic cylinder II is fixedly connected to a support table, and an electromagnet is fixedly connected to the middle of the support table.
[0011] Preferably, a partition board is fixedly connected to the inside of the machine body. The partition board divides the inside of the machine body into two areas: a storage area and an activity area. A telescopic cylinder I is rotatably connected to the bottom inside the machine body, and the other end of the telescopic cylinder I is rotatably connected to a warehouse door. The warehouse door is connected to the outside of the machine body through a hinge.
[0012] A control method for a delivery trolley includes the following steps: Physical modeling of the trolley: Establish a kinematic model of the tracked trolley, including traction force calculation, friction analysis, ramp stability analysis, and steering control model, to ensure the stable driving of the trolley on different terrains; Path planning optimization: Use the rapidly-exploring random tree for global path planning and combine it with the dynamic window algorithm for local path adjustment to ensure obstacle avoidance and optimal path of the trolley; Track adaptive control: According to terrain changes, adjust the track spacing and track height in real time to make the trolley adapt to different ground conditions and improve passability and stability; Reinforcement learning optimization: Adopt the deep Q-learning algorithm to adjust path decisions according to the current state of the trolley and perform multi-vehicle collaboration through V2X communication to improve scheduling efficiency.
[0013] Preferably, the physical modeling of the trolley further includes the following content: The calculation of traction force is based on the trolley mass, driving force, frictional force, and ramp resistance to ensure that the trolley has sufficient driving force under different slope conditions; The frictional force is modeled using the Coulomb friction model, and the friction coefficient is adjusted based on the terrain conditions; A steering control model based on the crawler speed difference is adopted to calculate the turning radius and angular velocity of the trolley, enabling the trolley to accurately control the turning amplitude and trajectory.
[0014] Preferably, the path planning optimization specifically includes the following: In the global path planning stage, the Rapidly-Exploring Random Tree (RRT) algorithm is adopted. Based on the random sampling and path extension strategies, a collision-free global driving path is generated; In the local path optimization stage, the Dynamic Window Approach (DWA) algorithm is utilized. Obstacles are detected through sensor data, and the optimal driving path is selected within the velocity space to minimize the collision risk; A local path optimization objective function is constructed using three elements: the target distance, the obstacle distance, and the speed. The driving trajectory is dynamically adjusted to improve the movement efficiency of the trolley.
[0015] Preferably, the crawler adaptive control includes the following: By adjusting the crawler spacing, the adaptability of the trolley in narrow channels and on ramps is improved. The adjustment of the crawler spacing is calculated in real time based on the terrain detection results; A terrain feedback mechanism is adopted. The crawler height is adjusted according to the parameters of the slope, step height, and surface friction force to ensure that the trolley can smoothly pass through terrains with height differences and optimize the grounding area on smooth ground to enhance the traction force; The crawler control strategy is optimized by combining reinforcement learning. The crawler configuration is adjusted based on different environmental data, enabling the trolley to adapt to different terrains and improving the stability and passing ability during long-term use.
[0016] Working principle: By driving the multi-axis robotic arm to work, the fixed ring is driven to move. In cooperation with the telescopic cylinder two, the turntable and the electromagnet are driven to move. Then, after placing the experimental sample to be delivered on the upper part of the electromagnet, the electromagnet is powered to generate magnetic force to adsorb the metal shell externally equipped with the sample. Then, the sample is transported into the storage area again by the multi-axis robotic arm, and the telescopic cylinder one is started to close the warehouse door.
[0017] The driving motor works to drive the polygonal pipe sleeve to rotate, which cooperates with the polygonal column to drive the inner driving wheel of the crawler group to rotate, so as to realize the movement of the trolley. When encountering a slope obstacle, one of the electric push rods 1 or 2 works to drive the frame to rotate downward, so that the crawler group forms an angle with the machine base to adapt to the movement on the inclined plane. When there are relatively high obstacles in the path, driving the electric push rod 1 and the electric push rod 2 at the same time can push the whole frame downward. At this time, the distance between the crawler group and the machine base will change, so as to lift the whole body to cross the relatively high obstacle.
[0018] When facing a relatively wide obstacle, the electric push rod 3 is driven to work to push the side frame outwards. At this time, the distance between the side frame and the frame will change. At the same time, the polygonal column extends from the middle of the polygonal pipe sleeve to ensure the normal operation of the crawler group.
[0019] The present invention provides a delivery trolley and a control method. It has the following beneficial effects: 1. By adopting the technical scheme combining crawler adaptive control and reinforcement learning optimization, the present invention realizes the intelligent adjustment of the tracked delivery trolley, significantly improving the driving stability and decision-making ability in complex environments. Compared with the existing crawler control methods that rely on fixed control parameters, the present invention can adaptively adjust the control strategy, solving the problem of poor adaptability of the fixed parameter method in heterogeneous terrains.
[0020] 2. The present invention adopts methods such as crawler tension adjustment, slip ratio control and ground pressure adaptive adjustment to ensure that the trolley can maintain stable driving under different terrain conditions. Compared with the problem of traditional crawler control methods relying on fixed tension settings, the present invention solves the problems of slipping or loosening of the crawler on soft or rough terrains by adjusting the tension in real time through feedback, enabling the tracked trolley to operate efficiently on slopes, sandy lands and slippery roads.
[0021] 3. Through the slip ratio adjustment based on PID control, the present invention effectively reduces the energy loss during the crawler driving process and improves the traction efficiency of the trolley. Compared with the technical scheme of traditional constant torque driving mode that is prone to slip on low adhesion ground, the present invention corrects the driving torque in real time according to the slip ratio feedback, solving the problem of energy waste caused by crawler idling, reducing the crawler wear at the same time, and extending the service life of the equipment.
[0022] 4. The present invention uses deep reinforcement learning to optimize the strategy, enabling the trolley to autonomously learn and adapt to complex environments without relying on manually preset rules. Different from the existing navigation schemes based on fixed path planning, the present invention enables the trolley to dynamically adjust the control strategy through the deep deterministic policy gradient algorithm, solving the problem that traditional methods cannot effectively cope with dynamic obstacles and unknown terrains, and improving the autonomous decision-making ability. Brief Description of the Drawings
[0023] Figure 1 is a three-dimensional structure diagram of the present invention; Figure 2 is a schematic diagram of the internal structure of the body in the present invention; Figure 3 is a schematic diagram of the structure of the taking component in the present invention; Figure 4 is a schematic diagram of the structure of the height adjustment component in the present invention.
[0024] Among them, 1, machine base; 2, body; 3, partition board; 4, storage area; 5, activity area; 6, first telescopic cylinder; 7, bin door; 8, base; 9, multi-axis robotic arm; 11, fixing ring; 12, second telescopic cylinder; 13, support table; 14, electromagnet; 15, rotating beam; 16, vehicle frame; 17, motor; 18, polygonal pipe sleeve; 19, polygonal column; 20, side frame; 21, track group; 22, first electric push rod; 23, second electric push rod; 24, third electric push rod. Detailed Embodiments
[0025] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0026] Please refer to the attached Figure 1 - attached Figure 4, an embodiment of the present invention provides a delivery cart, including a base 1, a body 2 is installed on the top of the base 1, and a taking component is installed inside the body 2. The taking component is used to conveniently send samples out of the body 2 or take external samples into the body 2 for temporary storage. The taking component includes a base 8, the base 8 is fixedly connected inside the body 2, a multi-axis robotic arm 9 is installed on the top of the base 8, a fixing ring 11 is installed at the free end of the multi-axis robotic arm 9, the multi-axis robotic arm 9 is assembled by multiple trunks, and each trunk is driven through a drive connection. The fixing ring 11 is fixed on the trunk at the outermost end of the multi-axis robotic arm 9. The structure and circuit connection of the multi-axis robotic arm 9 are prior arts and will not be elaborated here. A telescopic cylinder two 12 is fixedly connected to the middle of the fixing ring 11, and a support table 13 is fixedly connected to the output end of the telescopic cylinder two 12. The overall position of the support table 13 is adjusted by the telescopic cylinder two 12, and the multi-axis robotic arm 9 is used to realize the multi-angle movement of the support table 13 while keeping the support table 13 horizontal. An electromagnet 14 is fixedly connected to the middle of the support table 13. The sample is assembled in a matching metal container, and then the container is placed on the upper part of the support table 13. Then, the electromagnet 14 is driven to work to generate a magnetic force, which attracts the metal container, thereby ensuring the stability of taking the sample.
[0027] Please refer to the attached Figure 4 , a frame 16 is installed at the bottom of the base 1, a side frame 20 is installed outside the frame 16 through a width adjustment component, and a crawler group 21 is installed outside the side frame 20. The crawler group 21 includes a driving wheel, a driven wheel and a crawler sleeved outside, which are all prior arts and will not be described in detail here. The distance between the two crawler groups 21 can be adjusted through the width adjustment component, so as to adapt to the obstacle-crossing operation of wider obstacles. The width adjustment component includes an electric push rod three 24, the electric push rod three 24 is fixedly connected to the outside of the frame 16, and the output end of the frame 16 is fixedly connected to the inside of the side frame 20. The side frame 20 is pushed outwards by driving the electric push rod three 24 to increase the distance between the frame 16 and the side frame 20. A motor 17 is fixedly connected to the end of the frame 16, a polygonal pipe sleeve 18 is fixedly connected to the output end of the motor 17, a polygonal column 19 is slidably connected inside the polygonal pipe sleeve 18, and the other end of the polygonal column 19 is fixedly connected to the middle of the driving wheel of the crawler group 21. The motor 17 provides power for the crawler group 21, and the motor 17 and the driving wheel inside the crawler group 21 are connected through the polygonal pipe sleeve 18 and the polygonal column 19. When the electric push rod three 24 pushes the side frame 20 outwards, the polygonal column 19 synchronously extends from the middle of the polygonal pipe sleeve 18, and both the polygonal column 19 and the polygonal pipe sleeve 18 are polygonal structures, ensuring that the power of the motor 17 can be transmitted to the crawler group 21 while adapting to the width adjustment.
[0028] Please refer to the attached Figure 4, a height adjustment component is installed on the top of the frame 16. The overall height of the device is changed through the height adjustment component to cope with the obstacle-crossing operation of higher obstacles. The height adjustment component includes an electric push rod 22. One end of the electric push rod 22 is rotatably connected to the bottom of the machine base 1, and the other end of the electric push rod 22 is rotatably connected to the top of the frame 16. Driving the electric push rod 22 to work can make the frame 16 rotate. At this time, the included angle between the frame 16 and the machine base 1 can be changed. When encountering a slope, this operation can meet the passing needs of the inclined plane. The height adjustment component also includes a rotating beam 15. One end of the rotating beam 15 is rotatably connected to the bottom of the machine base 1, and the other end of the rotating beam 15 is rotatably connected to the top of the frame 16. The middle of the rotating beam 15 is rotatably connected to an electric push rod 23. The other end of the electric push rod 23 is rotatably connected to the bottom of the machine base 1. The rotating beam 15 provides support for the frame 16. At the same time, only driving the electric push rod 23 to work can also realize the angle adjustment operation of the frame 16. When it is necessary to change the distance between the frame 16 and the machine base 1, only need to drive the electric push rod 22 and the electric push rod 23 at the same time. In this way, the machine base 1 is raised. When passing through higher obstacles, the adaptive adjustment of the electric push rod 22 and the electric push rod 23 can be used to ensure the smooth operation of the trolley.
[0029] Please refer to the appendix Figure 2 , a partition 3 is fixedly connected inside the body 2. The partition 3 divides the inside of the body 2 into two areas, a storage area 4 and an activity area 5. The inside of the storage area 4 can be arranged according to needs into a space that meets the requirements of temperature, humidity, light, sterility, etc. for the storage needs of different samples. This is prior art and does not belong to the main improvement of this solution, so it will not be described in detail here.
[0030] A telescopic cylinder 6 is rotatably connected to the inner bottom of the body 2. The other end of the telescopic cylinder 6 is rotatably connected to a hatch 7. The hatch 7 is connected to the outside of the body 2 through a hinge. Driving the telescopic cylinder 6 to work realizes the opening and closing of the hatch 7, and cooperates with the picking component to realize the automatic picking of samples.
[0031] The structural design of the tracked trolley ensures its basic driving ability under different terrain conditions. However, it is difficult to fully meet the complex passing requirements of the laboratory environment only by optimizing the mechanical structure. In order to achieve stable and efficient sample transportation, it is necessary to combine intelligent control methods to dynamically adjust the tension of the track, the driving strategy and the path planning. By optimizing the control method, the trolley can adaptively adjust the track state according to the real-time environmental changes, improve the adhesion and traction performance, and at the same time optimize the path selection to ensure smooth passing on complex terrains.
[0032] A control method for a delivery trolley includes the following steps: Physical modeling of the vehicle: Establish a kinematic model for the tracked vehicle, including traction force calculation, friction analysis, ramp stability analysis, and steering control model, to ensure the stable driving of the vehicle on different terrains; The physical modeling of the vehicle further includes the following: The calculation of the traction force is based on the vehicle mass, driving force, friction force, and ramp resistance to ensure that the vehicle has sufficient driving force under different slope conditions; The friction force is modeled using the Coulomb friction model, and the friction coefficient is adjusted based on the terrain conditions; Adopt a steering control model based on the difference in tracked speeds to calculate the turning radius and angular velocity of the vehicle, enabling the vehicle to accurately control the turning amplitude and trajectory.
[0033] In this embodiment, for the motion characteristics of the tracked delivery vehicle, the present invention proposes a comprehensive physical modeling method. This method takes into account various physical factors such as traction force, friction force, ramp torque, and steering control to ensure that the vehicle can operate stably in a complex laboratory environment. The core objective of the physical modeling steps is to accurately simulate the dynamic behavior of the vehicle and provide a scientific basis for subsequent path planning and control strategies.
[0034] Modeling of traction force and friction force: Generally, the traction force of a tracked vehicle comes from the output thrust of the motor. The magnitude of the traction force directly affects the motion stability and driving ability of the vehicle. The traction force F d The relationship with the vehicle mass m and acceleration a can be calculated by the following formula: F d = m·a; where m is the mass of the vehicle and a is the acceleration of the vehicle. Specifically, the acceleration is determined by the vehicle's drive system, and its calculation takes into account factors such as motor output power, load, and ground friction.
[0035] Friction force is an important factor affecting the traction force, and it is usually modeled using the Coulomb friction model. On a flat ground, the friction force F f is proportional to the normal force N and can be expressed as: F f = μ·N; where μ is the friction coefficient and N is the normal force. Specifically, the variation of the normal force N under different terrains is the key point of this modeling. On flat ground, the normal force N is mg, where g is the acceleration due to gravity; while on a ramp, the expression of the normal force is N = mgcosθ, where θ is the inclination angle of the ramp.
[0036] On a ramp, the vehicle is not only affected by the friction force but also needs to consider the downhill force of the ramp. The downhill force of the ramp F gRelated to the slope θ, the calculation formula is: F g = mgsinθ; In this case, the traction force needs to be greater than the sum of the frictional force and the ramp downward force to ensure the stable forward movement of the trolley, that is: F d ≥F g +F f ; In some embodiments, the calculation of the traction force will be adjusted in real time according to the terrain where the trolley is located to ensure the stable driving of the trolley on the ramp or complex terrain.
[0037] Establishment of the steering model: The steering principle of the tracked trolley is achieved by adjusting the speed difference between the left and right tracks. The calculation of the turning radius and the angular velocity are important parts in this embodiment. Specifically, let the speed of the left track be v L , and the speed of the right track be v R , and the distance between the tracks be w t , then the turning radius R of the trolley can be expressed as: where, w t is the distance between the left and right tracks, v L and v R are the linear velocities of the left and right tracks respectively. The turning radius R controls the turning amplitude of the trolley. When the difference between v L and v R is large, the trolley can achieve a smaller turning radius, and vice versa, the turning radius is larger.
[0038] The angular velocity ω is the speed of the trolley's rotation, and its calculation formula is: By precisely adjusting v L and v R , the trolley can flexibly change its driving direction and adapt to different path planning requirements. Especially in narrow channels, adjusting the speed difference between the left and right tracks can help the trolley turn better and avoid collisions.
[0039] Mechanical and stability analysis: In some embodiments, it is also necessary to consider the stability of the trolley under different slopes and different surface conditions. Through the comprehensive analysis of the traction force, frictional force, ramp downward force and steering control, the motion state of the trolley can be adjusted in real time to ensure its stable operation in various complex environments.
[0040] For example, when the trolley needs to climb a slope, it is necessary to increase the output power of the motor to provide sufficient traction force to overcome the ramp downward force. When the trolley encounters steps or uneven ground, the adaptive adjustment of the tracks (including the adjustment of the track spacing and track height) can improve the stability and passability of the trolley.
[0041] Through these physical models and combined with sensor feedback, the trolley can judge its driving state in real time and automatically adjust its motion mode to cope with different environmental challenges.
[0042] In this embodiment, by establishing the physical model of the trolley, the traction force, frictional force, ramp stability and steering ability of the trolley can be effectively predicted and controlled. These models provide an accurate dynamic basis for subsequent path planning, obstacle avoidance, track control, etc., enabling the trolley to have good motion stability and adaptability in a complex laboratory environment.
[0043] Path planning optimization: Use the Rapidly-exploring Random Tree (RRT) for global path planning and combine it with the Dynamic Window Approach (DWA) for local path adjustment to ensure obstacle avoidance and optimal path for the trolley; In the global path planning stage, the RRT algorithm is adopted. Based on the random sampling and path extension strategy, a collision-free global driving path is generated. In the local path optimization stage, the DWA is used to detect obstacles through sensor data and screen out the optimal driving path in the velocity space to minimize the collision risk. A local path optimization objective function is constructed using three elements: target distance, obstacle distance, and speed, to dynamically adjust the driving trajectory and improve the motion efficiency of the trolley.
[0044] In this embodiment, the goal of path planning optimization is to enable the tracked delivery trolley to autonomously decide the driving path in a complex environment, while taking into account obstacle avoidance and driving efficiency. Path planning includes both the generation of the global path and the adjustment of the local path, enabling the trolley to adapt to the changes in the dynamic environment in real time. This method combines a sampling-based path search strategy and a local trajectory optimization algorithm to ensure the rationality and feasibility of the path.
[0045] Global path planning: Generally, the core of global path planning is to find a collision-free path for the trolley from the starting point to the target point and make the path length as short as possible. For this purpose, the RRT algorithm is adopted in this embodiment for global path generation. The RRT algorithm is based on the principle of random sampling and can efficiently construct a feasible path in a complex environment, avoiding being trapped in local optima.
[0046] Specifically, the process of generating the global path includes the following steps: Random sampling is performed in the configuration space to generate new nodes and connect them to the existing tree structure to form an expanding path network; The path is optimized using a heuristic method to remove redundant points, reduce unnecessary turns, and improve the smoothness of the path; for an environment with dynamic obstacles, an improved RRT algorithm is used to enable the path to perform local replanning when new obstacles appear and maintain the feasibility of the path.
[0047] In a possible implementation, when the path length exceeds a preset threshold, the path generated by RRT is optimized twice in combination with the A algorithm to further reduce the driving cost.
[0048] Local path optimization: As an option, local path optimization performs real-time trajectory adjustment based on the Dynamic Window Algorithm (DWA). The core of the DWA algorithm is to select the optimal control command in the velocity space according to the kinematic constraints of the vehicle, enabling the vehicle to avoid obstacles while maintaining reasonable motion efficiency.
[0049] Specifically, the optimization objective function of DWA includes the following three parts: Goal distance: Measures the proximity of the vehicle's current position to the target point, and preferentially selects a trajectory that can quickly approach the target; Obstacle distance: Calculates the safe distance between the vehicle and the nearest obstacle to ensure that the selected trajectory will not cause a collision; Velocity cost: Considers the vehicle's current speed to avoid a decrease in stability due to sudden speed changes.
[0050] In some embodiments, to improve the obstacle avoidance ability, grid modeling of the environment can be combined with lidar data, and environmental potential field information can be added during the DWA calculation, enabling the vehicle to make a better path selection in areas with dense obstacles.
[0051] Path dynamic adjustment: In a complex environment, the vehicle needs to have the ability to dynamically adjust the path. In this embodiment, the environmental changes are sensed in real time through sensor data, and the driving trajectory is dynamically corrected in combination with the path planning algorithm.
[0052] Specifically, when the vehicle detects an obstacle in the front path, the following steps will be executed: Calculate the intersection point of the current path and the obstacle, and predict the collision time; Adopt the local path optimization method to select a feasible new path; Through the crawler adaptive adjustment strategy, combined with path correction, improve the passability.
[0053] In a possible implementation, the path correction process combines the deep reinforcement learning method, enabling the vehicle to optimize the obstacle avoidance strategy based on historical driving data and improve the long-term adaptability.
[0054] In this embodiment, through RRT global path search, DWA local optimization, and a dynamic adjustment mechanism, the trolley can autonomously plan a driving path in a complex environment. Combining sensor data and reinforcement learning optimization, this method improves the obstacle avoidance ability and driving stability of the trolley, ensuring the efficient execution of delivery tasks.
[0055] Track adaptive control: According to terrain changes, the track spacing and track height are adjusted in real time to enable the trolley to adapt to different ground conditions, improving passability and stability. By adjusting the track spacing, the adaptability of the trolley in narrow channels and ramps is improved, and the adjustment of the track spacing is calculated in real time based on terrain detection results. A terrain feedback mechanism is adopted to adjust the track height according to parameters such as slope, step height, and surface friction, ensuring that the trolley can smoothly pass through terrains with height differences and optimizing the grounding area on smooth ground to enhance traction.
[0056] In this embodiment, the track adaptive control aims to achieve real-time adjustment of the track motion state for the driving characteristics of the tracked delivery trolley. This control method realizes the adaptive adjustment of track tension, track slip ratio, and track grounding pressure by integrating terrain perception data and track drive characteristics, ensuring the driving stability and passability of the trolley in a complex environment. The track adaptive control is closely related to path planning optimization and physical modeling results, and can further improve the environmental adaptability of the trolley.
[0057] Track tension control: Generally, the track tension has a direct impact on the grounding performance and driving efficiency of the trolley. The adjustment of the track tension is achieved through an active tensioning mechanism, and the magnitude of the tension is related to the force state of the track. The tension F t can be expressed as: F t = k t (l t - l 0 ); where k t is the stiffness coefficient of the tensioning mechanism, l t is the current track length, and l 0 is the initial track length. Specifically, when the track length becomes loose or too tight due to environmental changes, the active tensioning mechanism can perform dynamic adjustment according to the tension feedback.
[0058] In a possible implementation, the tension adjustment is adaptively corrected in combination with terrain slope information to avoid track slippage caused by looseness in a ramp environment.
[0059] Track slip ratio control: The track slip ratio reflects the difference between the actual driving speed and the theoretical driving speed of the track, directly affecting the traction efficiency of the trolley. The calculation formula for the slip ratio s is: Among them, v t is the theoretical linear speed of the crawler, and v a is the actual traveling speed of the trolley. Generally, a large slip ratio will cause energy loss and unstable traveling.
[0060] As an option, the slip ratio control uses a proportional-integral-derivative (PID) algorithm to correct the output torque of the crawler drive motor in real time. The PID control output torque M is expressed as: Among them, K p is the proportional coefficient, which is used to represent the response degree of the controller to the current error; K i is the integral coefficient, which is used to consider the cumulative effect of all past errors; K d is the differential coefficient, which is used to consider the influence of the error change rate. e is the error between the target slip ratio and the actual slip ratio; dt is the time difference, which is usually used to calculate the change rate of the error (the calculation of the differential term); ∫edt is the integral of the error with respect to time, which represents the cumulative effect of the error; is the derivative of the error with respect to time, which represents the rate of change of the error over time.
[0061] In some embodiments, the PID control parameters can be adjusted online through an adaptive learning algorithm to further improve the robustness of the crawler slip ratio control.
[0062] Crawler ground pressure adjustment: The crawler ground pressure directly affects the ground adhesion and passability of the trolley. The ground pressure P g is expressed as: Among them, F n is the normal support force, and A is the crawler ground contact area. Generally, a larger ground pressure helps to enhance the adhesion, but an excessive ground pressure may cause the crawler to sink.
[0063] Specifically, the ground pressure adjustment is achieved by adjusting the stiffness of the crawler suspension mechanism. In a possible implementation, the suspension stiffness is adaptively adjusted according to the terrain characteristics to ensure that the ground pressure is within a reasonable range.
[0064] In some embodiments, by combining a terrain recognition algorithm to estimate the ground hardness and slope in real time, the crawler ground pressure can be dynamically adjusted to enhance the passability of the trolley in soft ground or ramp environments.
[0065] Torque distribution and differential control: In order to further improve the crawler adaptability, in this embodiment, a differential control is performed on the output power of the left and right crawler drive motors through a torque distribution algorithm. Let the output torque of the left crawler be ML , the output torque of the right track is M R , then the torque difference ΔM can be expressed as: ΔM = M R - M L ; Generally, the torque difference is related to the turning radius of the trolley and the terrain slope. In a ramp environment, by appropriately increasing the torque of the uphill track, the climbing stability of the trolley can be effectively improved.
[0066] As an option, the torque distribution can also be combined with the slip ratio feedback signal to dynamically adjust the driving power of the left and right tracks to reduce the track slipping phenomenon.
[0067] Track adaptive adjustment strategy: In a complex environment, the track adaptive control adopts a hierarchical control strategy to achieve the collaborative optimization of multiple control objectives. Specifically, the track adaptive adjustment strategy includes: Real-time tension adjustment based on sensor feedback; Dynamic torque distribution combined with slip ratio control; Ground pressure adjustment driven by terrain recognition.
[0068] In a possible implementation, the track adaptive control performs parameter adaptive optimization through a deep reinforcement learning algorithm, enabling the trolley to continuously optimize the control strategy according to historical driving data.
[0069] In this embodiment, through track tension control, slip ratio adjustment, and ground pressure adaptive adjustment, the dynamic optimization of the track motion state is achieved. Combined with the torque distribution and differential control strategy, this method significantly enhances the environmental adaptability of the tracked delivery trolley and provides a reliable technical guarantee for autonomous driving in complex environments.
[0070] Reinforcement learning optimization: Adopt the deep Q-learning algorithm to adjust the path decision according to the current state of the trolley, and perform multi-vehicle collaboration through V2X communication to improve the scheduling efficiency.
[0071] Combine reinforcement learning to optimize the track control strategy, adjust the track configuration based on different environmental data, enable the trolley to adapt to different terrains, and improve the stability and passing ability during long-term use.
[0072] In this embodiment, the reinforcement learning optimization is mainly used to improve the autonomous decision-making ability of the tracked delivery vehicle in complex environments. Through the reinforcement learning framework, the vehicle can continuously adjust its strategy in a changing environment to optimize path planning, tracked control, and obstacle avoidance behavior. The reinforcement learning optimization not only relies on real-time sensing data but also combines historical experience to form an optimal control scheme for different scenarios. This method enables the vehicle to continuously enhance its decision-making ability and adapt to various complex terrains and dynamic obstacles through state space modeling, reward function design, and policy update mechanisms.
[0073] State space modeling: Generally, the state space S of reinforcement learning needs to completely describe the running state of the vehicle and the surrounding environment information. In this embodiment, the state vector s t consists of the following parts: The current position of the vehicle (x t , y t , θ t ); Speed state (v t , ω t ); Track slip ratio (s L , s R ); Local environment information E sensed by sensors t .
[0074] Specifically, E t is composed of data collected by lidar, depth camera, and inertial measurement unit (IMU), and represents the distribution of surrounding obstacles and terrain features in a grid-like manner.
[0075] In some embodiments, to improve the generalization ability of the model, the state space also includes historical trajectory information, enabling reinforcement learning to remember past driving experiences to optimize long-term decision-making effects.
[0076] Reward function design: As an option, the design of the reward function R(s t , a t ) needs to balance path optimality, obstacle avoidance ability, and driving stability. In this embodiment, the reward function includes the following key parts: R = w 1 R goal + w 2 R safe + w 3 R eff ; Where: The target approaching reward R goal , used to encourage the vehicle to approach the target point, is defined as: R goal = -∥p t - pg ∥; Among them, p t is the current coordinate, and p g is the target point coordinate.
[0077] Safety reward R safe , to avoid collision risks, is defined as: Among them, d min is the distance to the nearest obstacle, and d th is the safety threshold.
[0078] Driving efficiency reward R eff , to encourage smooth driving, is defined as: R eff = -(|a t | + |j t |); Among them, a t is the acceleration, and j t is the jerk.
[0079] In a possible implementation, the weights w 1 , w 2 , w 3 of the reward function adopt an adaptive adjustment strategy to enable the trolley to dynamically adjust the optimization goal in different environments.
[0080] Policy update mechanism: The core of reinforcement learning lies in policy optimization to maximize the long-term cumulative reward. In this embodiment, deep deterministic policy gradient (DDPG) is used for policy update, enabling the trolley to optimize control commands in a continuous action space.
[0081] Specifically, the DDPG training process includes the following steps: Adopt the Experience Replay mechanism to store historical interaction data and avoid sample correlation problems; Use the Target Network for policy update to stabilize the training process; Optimize the control policy π θ through policy gradient update, where the policy gradient is: Among them, Q φ is the state-action value function.
[0082] In some embodiments, to improve the convergence speed of reinforcement learning, the policy network adopts the Attention Mechanism to enable the trolley to focus on key environmental information and improve decision-making accuracy.
[0083] In a complex environment, a single reinforcement learning strategy may not be able to adapt to multiple task requirements. This embodiment combines multiple strategy optimization methods to improve the environmental adaptability of the trolley.
[0084] Specifically, the multi-strategy fusion includes: Combined with the rule-driven method, heuristic rules are used for correction when the reinforcement learning decision is unstable; Utilize imitation learning, and optimize the initial strategy through expert demonstration data to accelerate the training process; Adopt multi-agent reinforcement learning (MARL) so that multiple trolleys can collaboratively plan paths and avoid conflicts.
[0085] In a possible implementation, the multi-strategy fusion uses meta-learning for weight adaptive adjustment, enabling the trolley to automatically select the optimal control strategy in different environments.
[0086] Adaptive adjustment of reinforcement learning optimization In different environments, the hyperparameters of reinforcement learning optimization may need to be dynamically adjusted. This embodiment performs hyperparameter search through Bayesian optimization to improve the model training efficiency.
[0087] Generally, the key hyperparameters of reinforcement learning include: Learning rate α, which controls the gradient update step size; Discount factor γ, which weighs short-term and long-term rewards; Exploration rate ∈, which determines the balance between random exploration and exploitation.
[0088] In some embodiments, the hyperparameter optimization combines an evolutionary algorithm for adaptive adjustment, enabling the reinforcement learning model to maintain high efficiency in different task scenarios.
[0089] This embodiment realizes the intelligent decision-making ability of the tracked delivery trolley through reinforcement learning optimization. Combining state space modeling, reward function design, and policy update mechanism, this method can optimize path planning, obstacle avoidance strategies, and driving control. By adopting multi-strategy fusion and adaptive adjustment, the trolley can autonomously learn in complex environments and continuously improve the task execution efficiency.
[0090] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A delivery vehicle, comprising a base (1), characterized in that: A body (2) is mounted on the top of the machine base (1), a picking assembly is mounted inside the body (2), a frame (16) is mounted on the bottom of the machine base (1), a side frame (20) is mounted on the outside of the frame (16) via a width adjustment assembly, a track assembly (21) is mounted on the outside of the side frame (20), and a height adjustment assembly is mounted on the top of the frame (16).
2. A delivery vehicle according to claim 1, characterized in that: The width adjustment assembly comprises an electric push rod three (24), the electric push rod three (24) being fixedly connected to the outside of the frame (16), the output end of the frame (16) being fixedly connected to the inside of the side frame (20), the end of the frame (16) being fixedly connected to a motor (17), the output end of the motor (17) being fixedly connected to a polygonal sleeve (18), the polygonal sleeve (18) being slidably connected to a polygonal column (19), the other end of the polygonal column (19) being fixedly connected to the middle of a driving wheel of a track assembly (21).
3. A delivery vehicle according to claim 1, characterized in that: The height adjustment assembly comprises an electric push rod (22), one end of which is rotatably connected to the bottom of the base (1), and the other end of which is rotatably connected to the top of the frame (16).
4. A delivery vehicle according to claim 1, characterized in that: The height adjustment assembly further comprises a rotating beam (15), one end of the rotating beam (15) being rotatably connected to the bottom of the machine base (1), the other end of the rotating beam (15) being rotatably connected to the top of the vehicle frame (16), a second electric push rod (23) being rotatably connected to the middle of the rotating beam (15), the other end of the second electric push rod (23) being rotatably connected to the bottom of the machine base (1).
5. A delivery vehicle according to claim 1, characterized in that: The picking assembly comprises a base (8), the base (8) being fixedly connected to the inside of the machine body (2), a multi-axis mechanical arm (9) being installed on the top of the base (8), a fixing ring (11) being installed on the free end of the multi-axis mechanical arm (9), a telescopic cylinder 2 (12) being fixedly connected to the middle of the fixing ring (11), a support platform (13) being fixedly connected to the output end of the telescopic cylinder 2 (12), and an electromagnet (14) being fixedly connected to the middle of the support platform (13).
6. A delivery vehicle according to claim 1, characterized in that: A partition (3) is fixedly connected to the interior of the machine body (2), and the partition (3) divides the interior of the machine body (2) into two areas, a storage area (4) and an activity area (5). A telescopic cylinder (6) is rotatably connected to the bottom of the machine body (2), and a door (7) is rotatably connected to the other end of the telescopic cylinder (6). The door (7) is connected to the outside of the machine body (2) via a hinge.
7. A method for controlling a delivery vehicle, characterized in that: A delivery vehicle as claimed in any one of claims 1 to 6, comprising the following steps: Physical modeling of the tracked vehicle: Establish the kinematic model of the tracked vehicle, including traction calculation, friction analysis, slope stability analysis, and steering control model to ensure the vehicle's stable driving on different terrains; Path planning optimization: Use fast exploration random trees for global path planning, and combine dynamic window algorithms for local path adjustment to ensure the car avoids obstacles and optimizes the path; Track adaptive control: according to the terrain changes, the track spacing and track height are adjusted in real time to make the trolley adapt to different ground conditions and improve the passability and stability; Reinforcement learning optimization: Using deep Q learning algorithm, the path decision is adjusted according to the current state of the car, and multi-vehicle collaboration is carried out through V2X communication to improve scheduling efficiency.
8. A method for controlling a delivery vehicle according to claim 7, characterized in that: The physical modeling of the car further includes the following contents: The calculation of traction is based on the mass of the trolley, driving force, friction and slope resistance, ensuring that the trolley has sufficient driving force under different slope conditions; Friction is modeled using the Coulomb friction model, and the friction coefficient is adjusted based on terrain conditions; A steering control model based on track speed difference is used to calculate the turning radius and angular velocity of the vehicle, so that the vehicle can accurately control the turning amplitude and trajectory.
9. A method for controlling a delivery vehicle according to claim 7, characterized in that: The path planning optimization specifically includes the following contents: In the global path planning stage, a fast exploration random tree algorithm is used to generate a collision-free global driving path based on random sampling and path extension strategies; In the local path optimization stage, a dynamic window algorithm is used to detect obstacles through sensor data and select the optimal driving path in the speed space to minimize the risk of collision; The three elements of target distance, obstacle distance and speed are used to construct the local path optimization objective function, dynamically adjust the driving trajectory, and improve the movement efficiency of the car.
10. A method for controlling a delivery vehicle according to claim 7, characterized in that: The crawler track adaptive control includes the following contents: By adjusting the track spacing, the adaptability of the trolley in narrow passages and ramps is improved. The adjustment of track spacing is calculated in real time based on terrain detection results. The terrain feedback mechanism is used to adjust the track height according to the parameters of slope, step height, and surface friction, ensuring that the trolley can smoothly pass through the terrain with high differences, and optimize the contact area on smooth ground to improve traction; Combined with reinforcement learning to optimize the track control strategy, the track configuration is adjusted based on different environmental data, so that the vehicle can adapt to different terrains and improve its stability and passability during long-term use.