Intelligent logistics method based on deep reinforcement learning

By optimizing express delivery decisions through deep reinforcement learning and combining real-time traffic and environmental information, the problem of reliance on manual scheduling in existing technologies has been solved, realizing intelligent and automated express delivery and improving efficiency and user experience.

CN121937014APending Publication Date: 2026-04-28GUANGDONG VOCATIONAL & TECHNICAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG VOCATIONAL & TECHNICAL COLLEGE
Filing Date
2026-01-17
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The existing express delivery model relies on manual scheduling, which makes it difficult to cope with complex and ever-changing logistics scenarios. The route selection is not intelligent enough and the response to environmental changes is insufficient, resulting in low delivery efficiency, high costs and poor user experience.

Method used

By employing deep reinforcement learning methods and combining express delivery information, user information, and environmental data, delivery decisions are optimized through deep neural networks and reinforcement learning algorithms. Navigation instructions are generated and delivery routes and speeds are adjusted in real time. Dynamic optimization is also performed by combining GPS and sensor data.

Benefits of technology

It has enabled intelligent and automated express delivery, improved delivery efficiency, reduced operating costs, enhanced user experience, and ensured stable system operation through fault prediction and remote repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121937014A_ABST
    Figure CN121937014A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent logistics method based on deep reinforcement learning, and relates to the technical field of logistics control, and the method comprises the steps: obtaining express information, user information and environment data, transmitting the information to a deep neural network and a reinforcement learning algorithm model, and generating delivery decision information; the path planning and navigation equipment generates a navigation instruction according to the delivery decision information, the current express delivery vehicle position information, the map data and the real-time traffic information, and controls the express delivery vehicle to deliver the express delivery vehicle according to the optimal path and speed; the data updating device obtains the operation data and sends the operation data to the system database, and the system database updates the logistics information and provides the logistics information to the deep reinforcement learning server for continuous training and parameter updating, so that the delivery decision information is dynamically adjusted. According to the method, the problems that in the express delivery process in the prior art, path selection is not intelligent enough, response to environment changes is insufficient, dependence on manual scheduling is high, and global optimization is difficult to achieve are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent logistics control technology, and in particular to a smart logistics method based on deep reinforcement learning. Background Technology

[0002] With the rapid development of e-commerce and the continuous growth of express delivery volume, express delivery companies are under tremendous pressure in terms of order volume, service scope, and service timeliness. The existing express delivery model largely relies on manual scheduling, experience-based route selection, and manual execution of delivery tasks. In scenarios with complex road conditions, scattered delivery points, and strict time requirements, this easily leads to low delivery efficiency, high operating costs, unstable delivery times, and poor user experience. Especially against the backdrop of increasingly congested urban traffic, frequent weather changes, and increasingly personalized and time-sensitive user needs, traditional delivery methods struggle to respond promptly to environmental changes and to optimize the entire delivery process globally.

[0003] To improve the efficiency and management of express delivery, various technical solutions have been proposed in the industry. One solution uses preset rules and route planning algorithms, combined with static or semi-static map data, to plan vehicle routes, achieving some success in relatively simple and fixed environments. However, such rule-based and static data-driven systems struggle to promptly detect real-time traffic flow, road construction, and sudden weather changes, lacking comprehensive analysis capabilities for multi-source environmental data and exhibiting limited overall optimization capabilities in complex and ever-changing logistics scenarios. Another solution still relies primarily on manual dispatching, with dispatchers arranging delivery sequences and routes based on experience. This approach is highly dependent on human intervention and significantly limited by individual experience and energy levels, making it difficult to maintain consistently high efficiency and stability in large-scale business scenarios. Furthermore, some existing systems rely solely on a single data source for decision-making, failing to comprehensively utilize location data, traffic information, weather information, and user behavior information, resulting in insufficient intelligence and adaptability in the delivery system.

[0004] With the development of artificial intelligence technology, deep learning and reinforcement learning have demonstrated strong advantages in intelligent decision-making, path planning, and complex system control. They can continuously learn and optimize decision-making strategies through continuous interaction with the environment in a large-scale state space. Addressing issues such as inaccurate path selection, slow response to dynamic environmental changes, and high reliance on manual scheduling in express delivery, it is necessary to combine deep reinforcement learning with data collection, automatic control, and logistics business scenarios to construct a smart logistics method that can comprehensively utilize historical logistics data and real-time environmental information and possess continuous learning capabilities. Against this backdrop, this invention proposes to identify express delivery information and user information, collect and update environmental data, introduce a deep reinforcement learning server for delivery decisions, and combine it with path planning and navigation control to achieve intelligent, automated, and refined management of the delivery process, thereby improving overall delivery efficiency, reducing operating costs, and enhancing user experience. Summary of the Invention

[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a smart logistics method based on deep reinforcement learning. This invention solves the problems in the existing express delivery process, such as insufficiently intelligent path selection, insufficient response to environmental changes, high dependence on manual scheduling, and difficulty in achieving global optimization.

[0006] To achieve the above objectives, the present invention provides the following solution: A smart logistics method based on deep reinforcement learning includes: Utilize information identification and verification equipment to obtain express delivery information and user information; Use data acquisition equipment to acquire environmental data; The acquired express delivery information, user information, and environmental data are sent to the deep reinforcement learning server respectively. The deep reinforcement learning server inputs express delivery information, user information, and environmental data into a pre-built deep neural network and reinforcement learning algorithm model to generate delivery decision information; Based on delivery decision information and the current location of the delivery vehicle, combined with map data and real-time traffic information, navigation instructions are generated. Based on navigation instructions, the delivery vehicle executes delivery according to the optimal delivery route and vehicle speed determined in the delivery decision information. Obtain operational data during the delivery process of the express truck; The operational data is sent to the system database so that the system database can update and store the logistics information, thus obtaining the updated logistics information; The updated logistics information is provided to the deep reinforcement learning server, which then continuously trains and updates the parameters of the deep neural network and reinforcement learning algorithm model based on the updated logistics information, thereby dynamically adjusting the delivery decision information.

[0007] The present invention discloses the following technical effects: This invention provides a smart logistics method based on deep reinforcement learning. By optimizing delivery strategies through deep reinforcement learning and combining real-time traffic, environmental information, and user needs, the method can adjust delivery routes and speeds in real time, thereby effectively reducing delivery time, improving delivery accuracy, maximizing resource savings, and reducing operating costs. The combination of deep neural networks and reinforcement learning algorithms enables the system to automatically extract effective information from complex historical data and real-time information, predict and optimize delivery routes, times, and sequences, achieve dynamic decision-making, and adapt to the needs of different scenarios, thus improving the system's adaptability and stability. With the continuous collection of environmental data such as GPS positioning, traffic flow, and weather, the system can make timely adjustments to road conditions and weather changes, such as avoiding congested sections and rationally planning driving speeds, further optimizing express delivery efficiency and cargo safety. Through automated identification and inspection equipment, route planning, and navigation functions, manual intervention is significantly reduced, improving the automation level and reliability of express delivery. In addition, the intelligent maintenance and fault prediction system enables remote fault repair and regular maintenance, extending equipment lifespan and ensuring stable system operation. Users can check the status of their packages in real time through the intelligent interactive interface and enjoy convenient and efficient delivery services. At the same time, the fault prediction and remote repair functions can minimize system downtime and ensure uninterrupted service. Attached Figure Description

[0008] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0009] Figure 1 This is a flowchart of a smart logistics method based on deep reinforcement learning, provided as an embodiment of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0012] like Figure 1 As shown, this invention provides a smart logistics method based on deep reinforcement learning, comprising: Step 100: Obtain express delivery information and user information using information identification and verification equipment; Specifically, the implementing entities are information identification and verification equipment (such as cameras, RFID readers, barcode / QR code scanners, etc.).

[0013] When the delivery vehicle arrives at the delivery location, the identification and inspection equipment is automatically activated. It uses an RFID reader to read the RFID tag information carried by the package, or scans the barcode / QR code provided by the user to collect user information, and transmits the collected information to the deep reinforcement learning server.

[0014] These devices work together to transmit the collected express delivery information to a deep reinforcement learning server for processing. Their connection to the mobile chassis should ensure stable operation while the vehicle is in motion, and facilitate scanning of express deliveries at different positions and angles.

[0015] Step 200: Acquire environmental data using data acquisition equipment; Specifically, data acquisition equipment (including GPS devices).

[0016] The data acquisition equipment works continuously while the delivery vehicle is in motion. The GPS device acquires the delivery vehicle's geographical location information in real time, while other sensors collect surrounding environmental data, such as road conditions, traffic flow, and weather conditions, and transmit this information to the deep reinforcement learning server in real time.

[0017] This provides deep reinforcement learning servers with comprehensive environmental information, enabling them to consider various factors and formulate more reasonable delivery strategies. For example, adjusting routes based on traffic flow to avoid congested areas and improve delivery efficiency; and adjusting vehicle speed and protective measures based on weather conditions to ensure cargo safety.

[0018] Step 300: Send the acquired express delivery information, user information, and environmental data to the deep reinforcement learning server respectively; Step 400: The deep reinforcement learning server inputs the express delivery information, user information, and environmental data into a pre-built deep neural network and reinforcement learning algorithm model to generate delivery decision information; Specifically, after receiving user information from the user identification and verification device and environmental information from the data acquisition device, the deep reinforcement learning server performs calculations based on a pre-trained deep neural network and reinforcement learning algorithm. The deep neural network analyzes and processes a large amount of historical logistics data and real-time information, while the reinforcement learning algorithm seeks the optimal delivery strategy through continuous trial and error and reward mechanisms, such as determining the best delivery route, delivery order, and vehicle speed.

[0019] Specifically, the algorithm model includes a state set, an action set, a reward function, and state transitions. The state set includes the current path and the set of nodes to be delivered. The action set is the node to be selected for delivery. The reward function is the negative value of the travel path as the reward after the action is executed. The state transition is the jump to the next state after selecting an action in the current state.

[0020] State set S: Represented in a multidimensional structured form, specifically s t =(P t N t E t U t ), where: P t Let N be the sequence of routes traveled by the delivery vehicle at time t (including its current real-time location coordinates); t Let E be the set of nodes whose deliveries have not been completed at time t. Each node in the set is associated with attributes such as node coordinates, delivery requirements (e.g., type of goods, whether expedited), and user pickup preferences. t This refers to real-time environmental data at time t, including road conditions, traffic flow, and weather conditions; U t For the core user information corresponding to the delivery node (such as user pickup time window and priority identifier), the static attributes and dynamic environmental variables required for delivery decision-making are fully covered by the state set, providing data support for accurate decision-making.

[0021] Action set A: The action space changes dynamically with the state, specifically a t ∈A(s t )=N t That is, the action a at time t t To start from the current set of nodes to be delivered, N t The next delivery target node selected is directly associated with the delivery order optimization, and the decision-making efficiency is improved by limiting the scope of actions to avoid invalid exploration.

[0022] Reward function R: It adopts a composite design of "basic reward + constraint penalty", with the core basic reward being r. base =-α·dist(pos t ,a t ), where α is the distance weighting coefficient, dist(pos t ,a t ) is the current location of the delivery vehicle. t To target node a t The actual driving distance (corrected by real-time traffic information) is used, with negative distance values ​​as the base reward to ensure the model tends to choose the shortest path; at the same time, a constraint adaptation term r is incorporated. constraint =β·priority(a t )-γ·delay(at), where β is the user priority weight, priority(a t) For target node a t The corresponding user priority level, γ is the delay penalty coefficient, delay(a t To determine the potential delivery delay duration for this node, the final total reward r is... t =r base +r constraint This achieves dual optimization of "shortest path" and "constraint satisfaction"; when all nodes are delivered, an additional final reward r is set. final =-λ·total_dist (λ is the global distance weight, and total_dist is the total distance traveled), which strengthens the recognition of the global optimal path.

[0023] State transition to T: Execute action a t Afterwards, the state is dynamically updated according to fixed rules. The specific transition logic is as follows: P t+1 =P t ∪(post→a t ), that is, the travel segment from the current position to the target node is supplemented from the already traveled path sequence; N t+1 =N\{a t}, that is, the set of nodes to be delivered removes the selected target node; pos t+1 =at, meaning the delivery vehicle's current location is updated to the target node's location; E t+1 with U t+1 The system synchronously updates real-time environmental data and user information to time t+1, ensuring that the status at the next moment accurately reflects the current delivery progress and environmental changes, thus achieving real-time synchronization between the status and the actual delivery scenario.

[0024] By fully leveraging the advantages of deep reinforcement learning technology, intelligent decision-making in the logistics delivery process can be achieved. Through the analysis and learning of complex data, it can quickly adapt to different user needs and environmental changes, improving the accuracy and scientific nature of decision-making, thereby optimizing the entire logistics delivery process.

[0025] Step 500: Based on the delivery decision information and the current location of the delivery vehicle, and combined with map data and real-time traffic information, generate navigation instructions; Step 600: Based on navigation instructions, the delivery vehicle executes delivery according to the optimal delivery route and vehicle speed determined in the delivery decision information; Step 700: Obtain operational data during the delivery process of the courier vehicle; Step 800: Send the running data to the system database so that the system database can update and store the logistics information, and obtain the updated logistics information; Step 900: Provide the updated logistics information to the deep reinforcement learning server so that the deep reinforcement learning server can continuously train and update the parameters of the deep neural network and reinforcement learning algorithm model based on the updated logistics information, thereby dynamically adjusting the delivery decision information.

[0026] Specifically, the implementing entity is the route planning and navigation equipment.

[0027] The path planning and navigation equipment obtains the decision results calculated by the deep reinforcement learning server, and plans the optimal delivery route based on the current location information of the delivery vehicle and the target delivery location, combined with map data and real-time traffic information.

[0028] It provides precise route guidance for delivery vehicles, improving delivery efficiency and reducing travel time and fuel consumption. Meanwhile, the real-time navigation function can adjust routes promptly based on changes in road conditions, ensuring delivery vehicles always travel on the optimal route and further improving the timeliness of logistics delivery.

[0029] Execution entity: Data update device.

[0030] During the delivery process, data update equipment collects real-time information on the delivery vehicle's operational status (such as location, speed, and mileage), delivery task completion status (such as successful delivery and reasons for delivery failure), and user feedback (such as user ratings, complaints, and suggestions). This information is then promptly fed back to the system database. The system database updates and stores the logistics information based on the received data, providing the latest data support for subsequent decision-making and logistics management.

[0031] This enables dynamic optimization and continuous improvement of the logistics system. By updating data in a timely manner, the deep reinforcement learning server can continuously adjust delivery strategies based on the latest realities, improving the system's adaptability to complex and ever-changing logistics environments. Simultaneously, the complete logistics information stored in the system's database facilitates data analysis and management decision-making for enterprises, contributing to improved overall logistics service quality.

[0032] Furthermore, it also includes: The user interface of the smart delivery vehicle interacts with users through a touch screen, obtains the personal information and express delivery query requests entered by the user, and sends the personal information and express delivery query requests to the cloud center to return the express delivery query results to the user; The system monitors the operational status parameters of the intelligent delivery vehicle in real time, including battery level, motor temperature, and sensor operating status. Based on the operating status parameters, the fault prediction algorithm is called to analyze the operating status parameters, predict potential faults of the intelligent express delivery vehicle, and generate fault diagnosis results. A maintenance reminder will be issued based on the fault diagnosis results.

[0033] Specifically, user information acquisition and query: Implementing entity: The user interface of the intelligent delivery vehicle; Users interact with the smart delivery vehicle via touchscreen or voice recognition system, entering personal information and checking their packages.

[0034] Provides a user data foundation for personalized services.

[0035] Sensors and logistics systems for intelligent delivery vehicles: The sensors in the smart delivery vehicle monitor the vehicle's operating status in real time (such as battery level, motor temperature, sensor status, etc.) and send the data to the logistics system.

[0036] Provides real-time data support for fault detection.

[0037] Fault prediction and diagnosis: Logistics system and fault prediction algorithm for intelligent delivery vehicles; The logistics system uses fault prediction algorithms to analyze real-time monitored data, predict potential fault points, and generate fault diagnosis reports.

[0038] Early detection of potential faults can prevent them from affecting package pickup services.

[0039] Maintenance reminders and remote repair: Implementing entity: Logistics system and remote maintenance platform for intelligent delivery vehicles Once the logistics system detects a fault or predicts a potential fault, it will alert users or courier station managers through a digital human + VR interactive interface for maintenance. For certain faults that can be repaired remotely, the logistics system will attempt to repair them through a remote maintenance platform.

[0040] Ensure the long-term stable operation of intelligent delivery vehicles and reduce service interruptions caused by malfunctions.

[0041] This embodiment also provides the specific hardware details for applying the above method: A smart delivery vehicle comprises a chassis, a delivery compartment module, an RFID reader / writer, a user interface, a battery compartment, a drive system, and a logistics system. The chassis serves as the supporting structure for the entire vehicle, bearing the weight of all other components and ensuring stability and load-bearing capacity. The delivery compartment module, mounted on the chassis, stores packages and is typically designed as a multi-layered structure, with each layer independently adjustable in height, providing ample storage space and facilitating package retrieval and storage. The RFID reader / writer, installed at the front or side of the compartment module, scans and reads RFID tags on packages, enabling automatic package identification and status updates. The user interface, usually located at the front or side of the vehicle, includes a touchscreen and a voice recognition system, serving as the interface for user interaction, allowing users to check delivery records, input personal information, and retrieve packages. The battery compartment, located under or at the rear of the chassis, stores batteries to provide power to the smart delivery vehicle and ensure continuous operation. The drive system, installed under the chassis, includes a motor, wheels, and drive controller to control the movement of the smart delivery vehicle, enabling free movement and navigation on the road. The logistics system, located inside the smart delivery vehicle, is its "brain" and control system, responsible for processing various information and controlling the coordinated operation of various components to ensure the normal operation of the entire device.

[0042] The delivery vehicle's compartment module is fixed to the chassis by bolts or welding, ensuring the compartment's stability and load-bearing capacity. The battery compartment is fixed to the underside or rear of the chassis by bolts or clips, ensuring safe storage and easy replacement of the battery. The drive system's motor and wheels are fixed to the underside of the chassis by bolts or welding, and the drive controller is connected to the control system via wires. The user interface is connected to the chassis via a mounting bracket or embedded method, ensuring convenient operation for the user. The RFID reader, as the RFID tag scanning and inspection device for express delivery, is connected to the logistics system to collect and identify user information and send the data to the logistics system's database. The logistics system is connected to the RFID reader, user interface, drive system, and other components via wires or wirelessly to achieve information transmission and control command issuance, thereby enabling collaborative work between the various structures of the intelligent delivery vehicle.

[0043] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0044] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A smart logistics method based on deep reinforcement learning, characterized in that, include: Utilize information identification and verification equipment to obtain express delivery information and user information; Use data acquisition equipment to acquire environmental data; The acquired express delivery information, user information, and environmental data are sent to the deep reinforcement learning server respectively. The deep reinforcement learning server inputs express delivery information, user information, and environmental data into a pre-built deep neural network and reinforcement learning algorithm model to generate delivery decision information; Based on delivery decision information and the current location of the delivery vehicle, combined with map data and real-time traffic information, navigation instructions are generated. Based on navigation instructions, the delivery vehicle executes delivery according to the optimal delivery route and vehicle speed determined in the delivery decision information. Obtain operational data during the delivery process of the express truck; The operational data is sent to the system database so that the system database can update and store the logistics information, thus obtaining the updated logistics information; The updated logistics information is provided to the deep reinforcement learning server, which then continuously trains and updates the parameters of the deep neural network and reinforcement learning algorithm model based on the updated logistics information, thereby dynamically adjusting the delivery decision information.

2. The intelligent logistics method based on deep reinforcement learning according to claim 1, characterized in that, The information identification and verification equipment includes: Cameras, RFID readers, barcode and QR code scanners.

3. The intelligent logistics method based on deep reinforcement learning according to claim 1, characterized in that, The data acquisition device is a GPS device.

4. The intelligent logistics method based on deep reinforcement learning according to claim 1, characterized in that, The environmental data includes: Road condition information, traffic flow information, and weather information.

5. The intelligent logistics method based on deep reinforcement learning according to claim 1, characterized in that, The delivery decision information includes: Optimal delivery route, delivery sequence, and vehicle speed.

6. The intelligent logistics method based on deep reinforcement learning according to claim 1, characterized in that, Also includes: The user interface of the smart delivery vehicle interacts with users through a touch screen, obtains the personal information and express delivery query requests entered by the user, and sends the personal information and express delivery query requests to the cloud center to return the express delivery query results to the user; The system monitors the operational status parameters of the intelligent delivery vehicle in real time, including battery level, motor temperature, and sensor operating status. Based on the operating status parameters, the fault prediction algorithm is called to analyze the operating status parameters, predict potential faults of the intelligent express delivery vehicle, and generate fault diagnosis results. A maintenance reminder will be issued based on the fault diagnosis results.