Community unmanned delivery control method, device, equipment and medium
By using an LSTM-CNN fusion network to predict delivery demand heatmaps and a dynamic resource scheduling model in a community unmanned delivery system, the resource allocation of parcel lockers and drones is optimized. This solves the problems of low efficiency, uneven resource allocation, frequent coordination conflicts, insufficient energy management, and weak system robustness in existing technologies, and achieves efficient, stable, and flexible community unmanned delivery.
Patent Information
- Application Number
- CN202511086977.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-18
AI Technical Summary
Existing unmanned delivery control methods in communities suffer from problems such as low efficiency, uneven resource allocation, frequent coordination conflicts, insufficient energy management, and weak system robustness.
This paper adopts a community unmanned delivery control method. By acquiring historical time series data, it uses an LSTM-CNN fusion network to predict delivery demand heatmaps. Combined with a dynamic resource scheduling model and an improved path planning algorithm, it optimizes the allocation of express locker resources and the path of drones. It also uses a deep deterministic policy gradient algorithm for energy management, constructs an energy sharing network, and optimizes energy use.
It significantly improves delivery efficiency and parcel locker utilization, reduces resource conflicts, lowers energy consumption, enhances system robustness, enables continuous operation in the event of a power outage, adapts to complex environmental changes, and improves system flexibility and stability.
Smart Images

Figure CN120975478A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of unmanned delivery, and in particular to a community unmanned delivery control method, a corresponding device, an electronic device, and a computer readable storage medium. BACKGROUND
[0002] Traditional express logistics, cargo transportation, sorting, delivery, and sending are all achieved by relying on manpower. Each of these links requires considerable manpower and resources. During the delivery process, the courier not only needs to deliver the express to the designated location, but also needs to send a message to the customer one by one to come and pick up the package. In the process of sending, sometimes the customer needs to go to a special sending point to send, and sometimes the courier needs to personally go to the designated place of the sender to collect the package. The manpower and resources consumed in these two processes are particularly huge. In order to save costs, express cabinets gradually enter people's field of vision.
[0003] At present, the traditional technology has the following technical defects, mainly including:
[0004] Firstly, the traditional fixed express cabinet has the problem of static allocation of resources. Especially during peak hours, the overflow rate of express cabinet compartments is high, which makes users unable to pick up the package in time. At the same time, the regional distribution of express cabinets is uneven, and some users need to walk more than 500 meters to pick up the package, which reduces the convenience of delivery and user experience. In addition, the traditional express cabinet relies on power grid power supply and cannot operate normally during power failure, further affecting the continuity and stability of delivery.
[0005] Secondly, although the direct delivery mode of unmanned aerial vehicles has high delivery efficiency, the response time delay of the centralized control architecture is high, and conflicts are easy to occur in the process of multi-machine cooperation, affecting the overall delivery effect. In addition, the endurance of unmanned aerial vehicles is limited, and they rely on fixed charging piles for energy supply, which limits their application in large-scale delivery, especially in long-distance and high-frequency delivery scenarios.
[0006] Thirdly, the mixed artificial-automated delivery mode relies on manual end delivery, which can adapt to certain delivery needs, but due to the high labor cost, it is difficult to achieve significant cost savings. Moreover, this mode lacks adaptability in extreme scenarios (such as bad weather or special terrain, etc.), and cannot meet the diversified delivery needs.
[0007] Fourthly, the main defect of the delivery robot mode is the strong dependence on environmental data, which lacks adaptability in actual environment. The endurance problem of the robot in a long time is also a prominent problem. When the endurance is insufficient, the robot cannot complete the scheduled delivery task, and the limitation of the energy consumption model limits the delivery radius.
[0008] In summary, the community unmanned delivery control method in the prior art generally faces problems such as low efficiency, unbalanced resource allocation, frequent coordination conflicts, insufficient energy management, and weak system robustness. The present applicant makes corresponding explorations to solve these problems. SUMMARY
[0009] The present application aims to solve the above problems and provide a community unmanned delivery control method, a corresponding device, an electronic device, and a computer readable storage medium.
[0010] To achieve the various purposes of the present application, the present application adopts the following technical solutions:
[0011] A community unmanned delivery control method is proposed to achieve one of the purposes of the present application, comprising:
[0012] In response to the community unmanned delivery control instruction, historical time series data corresponding to the community to be delivered is obtained, wherein the historical time series data includes delivery feature vectors corresponding to a plurality of historical time steps, and the delivery feature vectors include allocated package data, community demographic data, meteorological data, point of interest spatial data, and e-commerce promotion time data.
[0013] The historical time series data is input into a delivery demand prediction model that has been trained to a converged state to predict a delivery demand heat map for a first predetermined time period in the future, wherein the delivery demand heat map represents the distribution of express cabinet resources required by the to-be-allocated package.
[0014] Based on the delivery demand heat map, the space utilization rate of each express cabinet in the community to be delivered and the to-be-allocated package data are obtained every second predetermined time period, a preset dynamic resource scheduling model is used to determine the reward function value according to the space utilization rate of the express cabinet and the to-be-allocated package data, and the target express cabinet number to which the to-be-allocated package is allocated and the size specification of the corresponding compartment are determined according to the reward function value, wherein the second predetermined time period is less than the first predetermined time period, and the to-be-allocated package data includes the volume, weight, and priority of the to-be-allocated package.
[0015] In response to the unmanned aerial vehicle delivery instruction, a delivery path planning algorithm is used to perform global path planning according to the target express cabinet number to which the to-be-allocated package is allocated, and based on an improved contract net protocol, a bidding value is calculated according to a preset bidding value calculation model when task allocation is performed, the remaining power of the unmanned aerial vehicle, the distance to the target express cabinet number, the estimated completion time, and the deviation degree value of the to-be-allocated package from the current task plan of the unmanned aerial vehicle are comprehensively evaluated, and a bidding value is generated to match the optimal execution unmanned aerial vehicle to complete the community unmanned delivery control.
[0016] Optionally, before the step of obtaining historical time series data corresponding to the community to be dispatched, the method comprises:
[0017] The collected raw data is preprocessed, wherein the raw data comprises allocated package data, community demographic data, meteorological data, point of interest spatial data, and e-commerce promotion time data;
[0018] The raw data is aligned in space-time by being unified to a preset size of a spatial grid and a preset time granularity, and is standardized by Z-score normalization, wherein the preset size comprises 300m x 300m or 500m x 500m, and the preset time granularity comprises 1 hour or 2 hours;
[0019] The missing values are filled by using a space-time K-nearest neighbor algorithm, and the outliers are detected by using an isolation forest algorithm, and the preprocessed raw data is used to construct the historical time series data corresponding to the community to be dispatched, so as to complete the preprocessing of the raw data.
[0020] Optionally, the step of inputting the historical time series data into a delivery demand prediction model trained to a converged state to predict a delivery demand heat map for a first preset time length comprises:
[0021] The historical time series data of the community to be dispatched is converted into a 5-dimensional space-time feature tensor, wherein the 5-dimensional space-time feature tensor comprises a batch size, a time dimension, a spatial grid dimension, and a feature channel dimension;
[0022] The 5-dimensional space-time feature tensor is input into the delivery demand prediction model trained to a converged state, sequentially passes through a bidirectional LSTM unit to capture time series dependency, passes through a 3D convolution kernel space-time convolution layer to extract local space-time patterns, passes through a spatial attention weight matrix activated by Sigmoid to dynamically adjust feature importance, and finally passes through a fully connected layer to generate a delivery demand heat map for a first preset time length.
[0023] Optionally, the step of constructing a reward function of a dynamic resource scheduling model comprises:
[0024] A first weight coefficient corresponding to a space utilization rate of a delivery cabinet of the community to be dispatched, a second weight coefficient corresponding to an average pickup distance of a user, a third weight coefficient corresponding to system energy consumption, and a fourth weight coefficient corresponding to a flexibility value for responding to sudden delivery demand are obtained.
[0025] The first product between the express cabinet space utilization and the first weight coefficient, the second product between the user average pick-up distance and the second weight coefficient, the third product between the system energy consumption and the third weight coefficient, and the fourth product between the burst delivery demand response flexibility value and the fourth weight coefficient are calculated and determined.
[0026] The first difference between the first product and the second product is calculated and determined, the second difference between the first difference and the third product is calculated and determined, and the first sum between the second difference and the fourth product is used to construct the dynamic resource scheduling model.
[0027] Optionally, the step of constructing the bid value calculation model comprises:
[0028] The fifth weight coefficient corresponding to the remaining power of the unmanned aerial vehicle, the sixth weight coefficient corresponding to the distance to the target express cabinet number, the seventh weight coefficient corresponding to the predicted delivery completion time, and the eighth weight coefficient corresponding to the deviation degree value of the delivery to-be-assigned package from the current task plan of the unmanned aerial vehicle are obtained.
[0029] The fifth product between the remaining power of the unmanned aerial vehicle and the fifth weight coefficient is calculated and determined, the first ratio between the sixth weight coefficient and the distance to the target express cabinet number is calculated and determined, the sixth product between the predicted delivery completion time and the seventh weight coefficient is calculated and determined, and the seventh product between the deviation degree value of the delivery to-be-assigned package from the current task plan of the unmanned aerial vehicle and the eighth weight coefficient is calculated and determined.
[0030] The second sum between the fifth product, the first ratio, and the sixth product is calculated and determined, and the third difference between the second sum and the seventh product is calculated and determined, so as to construct the bid value calculation model.
[0031] Optionally, after the step of generating a bid value to match the optimal execution unmanned aerial vehicle by comprehensively evaluating the remaining power of the unmanned aerial vehicle, the distance to the target express cabinet number, the predicted delivery completion time, and the deviation degree value of the delivery to-be-assigned package from the current task plan of the unmanned aerial vehicle according to the preset bid value calculation model during task allocation, the step comprises:
[0032] An energy state space of the energy sharing network is constructed, wherein the state space comprises a state of charge of the super capacitor, a state of charge of the solid-state lithium battery, a load power, an ambient temperature, and a solar power prediction value in a future preset time length.
[0033] An energy action space of the energy sharing network is constructed, wherein the energy action space comprises a super capacitor output power, a lithium battery output power, and an energy sharing power.
[0034] Based on the deep deterministic policy gradient algorithm, energy scheduling is performed according to the energy state space, with the optimization objective of maximizing a preset energy management reward function, to output the energy action space. The expression for the energy management reward function is:
[0035] R = 0.7·(1-||SOC) bat -0.5|)+0.2·η system -0.1·P share ,
[0036] Where, η system P represents the overall energy efficiency of the system. share State of Charge (SOC) represents the amount of energy transferred between the drone and the parcel locker, or between different energy storage units. bat This indicates the state of charge of a solid-state lithium battery.
[0037] Optionally, the allocated package data includes the volume and weight of the allocated package, as well as the express locker number where the allocated package is placed and the size of its corresponding compartment. The point of interest spatial data represents the spatial distribution location and type information of the community to be delivered and the surrounding commercial areas, residential areas, and public service facilities.
[0038] The basic network architecture of the delivery demand prediction model is an LSTM-CNN fusion network, which includes an input layer, an LSTM layer, a spatiotemporal convolutional layer, an attention mechanism, and a fully connected layer. The LSTM layer includes multiple bidirectional LSTM units. The delivery route planning algorithm includes an improved A* algorithm.
[0039] A community unmanned delivery control device provided for another purpose of this application includes:
[0040] The data acquisition module is configured to respond to community unmanned delivery control commands and acquire historical time series data corresponding to the community to be delivered. The historical time series data includes delivery feature vectors corresponding to multiple historical time steps. The delivery feature vectors include allocated package data, community population statistics, meteorological data, point of interest spatial data, and e-commerce promotion time data.
[0041] The delivery heat map prediction module is configured to input the historical time series data into a delivery demand prediction model that has been trained to a convergent state, so as to predict the delivery demand heat map for the first preset time period in the future, wherein the delivery demand heat map represents the distribution of express cabinet resources required for the parcels to be allocated;
[0042] The dynamic resource scheduling module is configured to acquire the space utilization rate of each express locker and the data of parcels to be allocated in the community to be delivered every second preset time interval based on the delivery demand heat map. The module uses a preset dynamic resource scheduling model to determine the reward function value based on the express locker space utilization rate and the data of parcels to be allocated. Based on the reward function value, the module determines the target express locker number to which the parcels to be allocated are assigned and the size specifications of their corresponding compartments. The second preset time interval is shorter than the first preset time interval. The data of parcels to be allocated includes the volume, weight and priority of the parcels to be allocated.
[0043] The unmanned delivery control module is configured to respond to drone delivery commands. It employs a delivery path planning algorithm to perform global path planning based on the target parcel locker number assigned to the package to be delivered. Based on an improved contract network protocol and a preset bid value calculation model, during task allocation, it comprehensively evaluates the drone's remaining battery power, distance to the target parcel locker number, estimated delivery time, and deviation of the package to be delivered from the drone's current task plan to generate a bid value to match the optimal drone for execution, thereby completing the unmanned delivery control in the community.
[0044] An electronic device provided for another purpose of this application includes a central processing unit and a memory, the central processing unit being configured to invoke and run a computer program stored in the memory to perform the steps of the community unmanned delivery control method of this application.
[0045] A computer-readable storage medium is provided for another purpose of this application, which stores, in the form of computer-readable instructions, a computer program implemented according to the community unmanned delivery control method, which, when invoked by a computer, performs the steps included in the corresponding method.
[0046] Compared to existing technologies, this application addresses the common problems faced by existing community unmanned delivery control methods, such as low efficiency, uneven resource allocation, frequent coordination conflicts, insufficient energy management, and weak system robustness. This application offers the following beneficial effects, including but not limited to:
[0047] Firstly, the community unmanned delivery control method of this application can greatly improve delivery efficiency, with a 172.7% increase in last-mile delivery efficiency. Compared with the traditional model, this application achieves a significant improvement in delivery efficiency. By rationally allocating the space of the express cabinet, optimizing the drone path planning and task allocation, the delivery time is greatly shortened, improving the response speed and timeliness of delivery.
[0048] Secondly, the community unmanned delivery control method proposed in this application improves the space utilization rate of express delivery lockers by 93.6%. It adopts dual-timescale reinforcement learning to optimize the storage space of express delivery lockers, making the space use of express delivery lockers more reasonable and efficient, avoiding the problems of resource waste and insufficient space, and improving the storage capacity of each express delivery locker.
[0049] Third, the community unmanned delivery control method of this application can reduce resource conflicts and energy consumption, and reduce the drone task conflict rate by 90.6%. Through the improved A* algorithm and contract network protocol, the system effectively resolves conflicts in multi-drone collaboration, ensures smooth multi-drone task allocation, and improves the overall delivery collaboration efficiency.
[0050] Fourth, the community unmanned delivery control method of this application reduces energy consumption by 44.3%. The application of the deep deterministic policy gradient algorithm in energy scheduling effectively optimizes energy use. Through intelligent scheduling and control, unnecessary energy consumption is reduced, resulting in a significant improvement in the system's energy efficiency compared to traditional methods.
[0051] Fourth, the community unmanned delivery control method proposed in this application enhances the system's robustness and sustainability. It has a 96-120 hour off-grid operation capability, and the hybrid energy supply device integrates multiple energy harvesting and storage technologies, ensuring the system can operate continuously and stably for extended periods without grid support. This feature greatly enhances the system's robustness, especially in extreme situations such as power outages, guaranteeing the reliability of the unmanned delivery system.
[0052] Fifth, the community unmanned delivery control method of this application integrates an innovative architecture of distributed intelligent express cabinet clusters, drone swarms, and a collaborative control platform, enabling the system to flexibly respond to various delivery needs and complex environmental changes, exhibiting stronger adaptability and stability. Accurate delivery demand prediction is achieved through an LSTM-CNN fusion network. Utilizing historical data and multi-dimensional features (such as allocated package data, demographic data, and meteorological data), future delivery can be effectively predicted. Based on the prediction results and dynamic express cabinet space utilization, the system employs a dynamic resource scheduling model to intelligently allocate packages, ensuring maximum resource utilization while meeting various requirements such as package volume, weight, and priority.
[0053] In summary, this application demonstrates significant advantages in improving last-mile delivery efficiency, resource utilization, collaborative operations, energy management, and system robustness. It addresses bottlenecks in traditional delivery models by optimizing delivery processes, increasing resource utilization, and reducing energy consumption, thereby significantly lowering overall operating costs. This more efficient delivery model not only reduces enterprise costs but also enhances the consumer experience, driving socio-economic benefits. Intelligent energy scheduling and energy consumption optimization further reduce carbon emissions, aligning with green and sustainable development requirements, contributing to environmental protection, and possessing broad application prospects and substantial socio-economic benefits. Attached Figure Description
[0054] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0055] Figure 1 This is an exemplary network architecture used in the community unmanned delivery control method of this application;
[0056] Figure 2 This is a schematic block diagram of the community unmanned delivery control device in the embodiments of this application;
[0057] Figure 3 This is a schematic diagram of the structure of the computer device in the embodiments of this application. Detailed Implementation
[0058] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain this application, and should not be construed as limiting this application.
[0059] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this application means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is “connected” or “coupled” to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. Furthermore, “connected” or “coupled” as used herein can include wireless connections or wireless coupling. The term “and / or” as used herein includes all or any units and all combinations of one or more associated listed items.
[0060] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the same meaning as in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as herein.
[0061] Those skilled in the art will understand that the terms "client," "terminal," and "terminal device" as used herein include both devices that receive wireless signals, devices that only possess wireless signal receiver capabilities without transmission capabilities, and devices with receiving and transmitting hardware, devices that have receiving and transmitting hardware capable of bidirectional communication over a bidirectional communication link. Such devices may include: cellular or other communication devices such as personal computers or tablets, having single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) that can combine voice, data processing, fax, and / or data communication capabilities; PDA (Personal Digital Assistant) that may include a radio frequency receiver, pager, internet / intranet access, web browser, notepad, calendar, and / or GPS (Global Positioning System) receiver; and conventional laptops and / or handheld computers or other devices that have and / or include radio frequency receivers. As used herein, "client," "terminal," and "terminal device" can be portable, transportable, installed in a means of transportation (air, sea, and / or land), or suitable and / or configured to operate locally and / or in a distributed manner, operating in any other location on Earth and / or in space. "Client," "terminal," and "terminal device" as used herein can also be a communication terminal, an internet access terminal, or a music / video playback terminal, such as a PDA, a MID (Mobile Internet Device), and / or a mobile phone with music / video playback capabilities, or a smart TV, set-top box, etc.
[0062] The hardware referred to by the names "server," "client," and "service node" in this application is essentially an electronic device with the equivalent capabilities of a personal computer. It is a hardware device with the necessary components revealed by the von Neumann architecture, such as a central processing unit (including an arithmetic logic unit and a control unit), memory, input devices, and output devices. The computer program is stored in its memory, and the central processing unit loads the program stored in the secondary storage into the main memory to run it, execute the instructions in the program, and interact with the input and output devices to complete specific functions.
[0063] It should be noted that the concept of "server" used in this application can also be extended to the case of server clusters. Based on the network deployment principles understood by those skilled in the art, the servers should be logically divided. Physically, these servers can be independent of each other but accessible through interfaces, or they can be integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method in this application.
[0064] One or more of the technical features of this application, unless explicitly specified herein, can be deployed on a server and accessed by a client remotely calling the online service interface provided by the server, or can be directly deployed and run on a client to access the service.
[0065] Unless otherwise specified, the neural network models referenced or potentially referenced in this application may be deployed on a remote server and invoked remotely on the client, or deployed on a client with the capability to invoke directly. In some embodiments, when running on the client, the corresponding intelligence may be acquired through transfer learning in order to reduce the requirements on the client's hardware resources and avoid excessive consumption of the client's hardware resources.
[0066] Unless otherwise specified, all data involved in this application may be stored remotely on a server or on a local terminal device, as long as it is suitable for use by the technical solution of this application.
[0067] Those skilled in the art will understand that although the various methods in this application are described based on the same concept and thus present commonality among them, they can be performed independently unless otherwise specified. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept; therefore, concepts expressed in the same way, as well as concepts that are appropriately changed for convenience but are expressed differently, should be understood equivalently.
[0068] Unless otherwise expressly stated, the various embodiments disclosed in this application can be combined in a cross-cutting manner to flexibly construct new embodiments, as long as such combination does not depart from the inventive spirit of this application and can meet the needs of the prior art or solve a certain deficiency in the prior art. Those skilled in the art should be aware of such modifications.
[0069] The community unmanned delivery control method of this application can be implemented based on a community unmanned delivery control system. This system includes a distributed intelligent parcel locker cluster, a drone swarm, a hybrid power supply device, a dynamic resource scheduling module, and a drone swarm decision-making layer. The distributed intelligent parcel locker cluster is deployed in multiple strategic locations within the community, with optimized layout based on population density and activity patterns. For example, in high-density residential areas, the distance between parcel lockers is less than 300 meters; in medium- and low-density mixed communities, they are located in densely populated areas; and in suburban villa areas, they are located at community entrances and public service areas. Each parcel locker unit uses a 5G module for high-speed data transmission, its computing unit employs an edge computing chip with AI acceleration capabilities, and its storage unit is configured with different sizes and numbers of compartments based on the community's parcel size distribution.
[0070] In some embodiments, the drone swarm comprises drones of different sizes, with the appropriate model selected based on the delivery task. For example, light drones are used for short-distance, small-item deliveries, while medium or heavy drones are used for long-distance, large-item deliveries. In the drone navigation system, millimeter-wave radar is used for long-range obstacle detection, binocular vision is used to identify packages and target locations, and lidar is used for high-precision positioning. The communication module selects an appropriate frequency band and communication protocol based on the environment, such as using the 5GHz band TDMA protocol in areas with less interference.
[0071] In some embodiments, the hybrid energy supply device is installed on parcel lockers and drones, with solar panels selected for appropriate installation angles and areas based on their location and sunlight conditions. In the energy storage unit, graphene supercapacitors and solid-state lithium batteries are configured in a specific ratio to meet varying power and energy demands. The energy management controller uses the DDPG algorithm for energy dispatching based on real-time energy acquisition and load conditions.
[0072] In some embodiments, the perception layer of the collaborative control platform collects data through various sensors within the community, such as traffic flow monitors that combine geomagnetic sensors and cameras, and environmental monitoring stations that monitor data such as temperature, humidity, and air quality. The analysis layer uses a spatiotemporal graph convolutional network to analyze the collected data, the decision layer generates control strategies based on multi-objective reinforcement learning algorithms, and the execution layer translates the strategies into specific task instructions and issues them to the parcel lockers and drones.
[0073] In some embodiments, the parcel locker adopts a honeycomb modular design, with each module being independently detachable and expandable. Internal compartments are categorized according to parcel type and size; for example, small compartments store document parcels, medium-sized compartments store general merchandise parcels, and large compartments store large items such as home appliances. Special compartments are used to store perishable, fragile, or other special items.
[0074] In some embodiments, within the dynamic resource scheduling module, the macro-strategy network generates a community storage resource allocation strategy based on 7-30 days of historical data, such as residents' shopping habits (increased shopping on weekends) and holiday promotions. The micro-execution network allocates specific storage slots based on real-time data, such as the current occupancy rate of parcel lockers, characteristics of newly arrived packages, and real-time heat maps of the community. For example, when a new package arrives, the system selects the optimal storage location based on characteristics such as package size, weight, and priority, combined with the occupancy status of each parcel locker slot and the expected usage time. Simultaneously, the system implements a "tidal" storage strategy based on residents' daily routines, such as increasing parcel locker capacity in commercial areas during the day and in residential areas at night; and activating an "emergency reservation mechanism" during special weather conditions or large-scale events to reserve space for medical supplies and daily necessities.
[0075] In some embodiments, the drone swarm decision layer employs an improved A* algorithm combined with traffic flow prediction for global path planning. For example, during morning rush hour, it avoids congested areas and selects routes with less traffic. The coordination layer uses a token-passing distributed consensus protocol to avoid conflicts when multiple drones need to cross the same airspace through a virtual time-space reservation system. The execution layer fuses multi-sensor data to achieve centimeter-level positioning accuracy and safe obstacle avoidance. For instance, when an obstacle is detected ahead, the drone determines the type and distance of the obstacle based on sensor data and automatically adjusts its flight direction and altitude.
[0076] Inter-UAV communication employs an adaptive TDMA-CSMA hybrid protocol, dynamically adjusting time slot allocation and contention window size based on network congestion levels. In high-interference environments, it automatically switches to frequency-hopping spread spectrum mode, switching between the 2.4GHz and 5.8GHz bands. Based on an improved contract network protocol and a decentralized task allocation mechanism, the collaborative control platform broadcasts new tasks as they are generated.
[0077] Please see Figure 1 In one embodiment of the community unmanned delivery control method of this application, the method includes:
[0078] Step S10: Respond to the community unmanned delivery control command and obtain the historical time series data corresponding to the community to be delivered. The historical time series data includes delivery feature vectors corresponding to multiple historical time steps. The delivery feature vectors include allocated package data, community population statistics, meteorological data, point of interest spatial data, and e-commerce promotion time data.
[0079] The community unmanned delivery control system in the terminal device can respond to community unmanned delivery control commands and obtain historical time series data corresponding to the community to be delivered. The historical time series data includes delivery feature vectors corresponding to multiple historical time steps. The delivery feature vectors include allocated package data, community population statistics, meteorological data, point-of-interest spatial data, and e-commerce promotion time data. The allocated package data includes the volume and weight of the allocated packages, as well as the express locker number where the allocated packages are placed and the size specifications of its corresponding compartment. The point-of-interest spatial data represents the spatial distribution and type information of commercial areas, residential areas, and public service facilities within and around the community to be delivered.
[0080] Specifically, each delivery feature vector is a dataset covering multiple aspects related to the delivery task. Among these, the allocated package data describes the current status and delivery information of the packages. Specifically, the allocated package data includes the locker number where the package was placed and the size specifications of its corresponding compartment. The locker number indicates which locker the package was placed in; the compartment size specifications indicate the required storage compartment size for the package. This is to ensure that the package fits its size in the locker and avoids situations where the package cannot be stored due to an unsuitable size.
[0081] The community demographic data refers to basic statistical information about the residents of the community, which may include age distribution, gender ratio, income level, family structure, etc. This data helps to understand the community's needs and predict peak delivery periods or the volume of deliveries that may be needed.
[0082] The meteorological data refers to historical weather-related data, such as temperature, humidity, precipitation, and wind speed. This data is crucial for delivery optimization because weather conditions directly impact delivery timeliness, safety, and demand. For example, severe weather conditions may lead to a surge in demand or extended delivery times.
[0083] The spatial data of points of interest (POIs) represents the spatial distribution and type of commercial areas, residential areas, and public service facilities within and around the community to be delivered. This data reflects the potential impact of these locations on the delivery process; for example, commercial areas may have higher parcel demand, while delivery demand in residential areas is typically more dispersed, and public service facilities (such as hospitals and schools) may be high-frequency delivery points. For instance, the system can increase the capacity of parcel lockers in commercial areas during the day and in residential areas at night. Simultaneously, the system has an "emergency reserve mechanism" that automatically reserves space for medical supplies and daily necessities during special weather conditions or large-scale events. Through this spatial data, delivery routes, time, and resource allocation can be optimized.
[0084] E-commerce promotion timing data refers to the timing and frequency of promotional activities conducted by e-commerce platforms. This type of data reflects fluctuations in delivery demand during promotional activities. E-commerce promotions typically trigger surges in order volume, and understanding the timing of promotions helps predict peak delivery demand and prepare resources in advance.
[0085] In some embodiments, prior to the step of obtaining historical time-series data corresponding to the community to be delivered, the following steps are included:
[0086] Step S101: Preprocess the collected raw data, wherein the raw data includes allocated package data, community population statistics, meteorological data, point of interest spatial data, and e-commerce promotion time data;
[0087] Step S102: Unify the original data to a spatial grid of a preset size and a preset time granularity for spatiotemporal alignment, and use Z-score normalization for feature standardization. The preset size includes 300m×300m or 500m×500m, and the preset time granularity includes 1 hour or 2 hours.
[0088] Step S103: Use the spatiotemporal K-nearest neighbor algorithm to fill in missing values and the isolated forest algorithm to detect anomalies. The preprocessed raw data is used to construct the historical time series data corresponding to the community to be delivered, so as to complete the preprocessing of the raw data.
[0089] Step S20: Input the historical time series data into the delivery demand prediction model that has been trained to convergence state, so as to predict the delivery demand heat map for the first preset time period in the future, wherein the delivery demand heat map represents the distribution of express cabinet resources required for the parcels to be allocated.
[0090] After obtaining the historical time series data corresponding to the communities to be delivered, the historical time series data is input into the delivery demand prediction model that has been trained to a converged state to predict the delivery demand heatmap for a first preset duration in the future. The delivery demand heatmap represents the distribution of express locker resources required for the parcels to be allocated. The basic network architecture of the delivery demand prediction model is an LSTM-CNN fusion network, which includes an input layer, an LSTM layer, a spatiotemporal convolutional layer, an attention mechanism, and a fully connected layer. The LSTM layer includes multiple bidirectional LSTM units. The first preset duration includes 24 hours, 48 hours, or 72 hours, etc.
[0091] In some embodiments, the step of training a delivery demand prediction model includes:
[0092] Step S2001: Obtain the training dataset, wherein the training dataset includes multiple training samples and their corresponding supervision labels. The training samples include multiple delivery feature vectors corresponding to multiple historical time steps. The supervision labels represent the distribution of express locker resources occupied by allocated packages in the community to be allocated at historical time steps.
[0093] Step S2002: Input the training dataset into the preset delivery demand prediction model. When the prediction accuracy of the model on the validation set reaches the preset threshold and converges, the training of the delivery demand prediction model is completed.
[0094] In some embodiments, the step of inputting the historical time series data into a delivery demand prediction model trained to a convergent state to predict a delivery demand heatmap for a future first preset duration includes:
[0095] Step S201: Convert the historical time series data of the community to be delivered into a 5-dimensional spatiotemporal feature tensor, wherein the 5-dimensional spatiotemporal feature tensor includes batch size, time dimension, spatial grid dimension and feature channel dimension;
[0096] Step S202: Input the 5-dimensional spatiotemporal feature tensor into the delivery demand prediction model that has been trained to convergence. Sequentially, capture the time series dependencies through bidirectional LSTM units, extract local spatiotemporal patterns through spatiotemporal convolutional layers with 3D convolutional kernels, dynamically adjust the feature importance through a spatial attention weight matrix activated by Sigmoid, and finally generate a heatmap of delivery demand for the first preset duration through a fully connected layer.
[0097] Specifically, historical time-series data is input into a delivery demand prediction model to forecast future delivery demand. This model is a trained and converged LSTM-CNN fusion network, typically trained on historical time-series data corresponding to the communities to be delivered, capable of predicting delivery demand over a certain period. The prediction result is a delivery demand heatmap, representing the intensity of delivery demand in various areas over a future time period. Each area in the heatmap represents the parcel occupancy at a specific location (e.g., a parcel locker) in the future. The heatmap clearly shows which areas will have higher demand and which will have relatively lower demand, thus allowing for optimized resource allocation.
[0098] Furthermore, the LSTM-CNN fusion network is a deep learning architecture that combines Long Short-Term Memory (LSTM) networks and Convolutional Neural Networks (CNNs) to simultaneously utilize the temporal dependencies and spatial features of time-series data. The LSTM layer is a Recurrent Neural Network (RNN) structure used to process time-series data. LSTM can effectively capture long-term and short-term dependencies, avoiding the gradient vanishing problem in traditional RNNs during long-term series learning. Here, the LSTM layer is mainly used to analyze and process the temporal dependencies of historical time-series data. In LSTM, data at each time step is processed sequentially, while bidirectional LSTM units process the data both forward and backward. This allows the model to consider past and future information when making time-series predictions, improving prediction accuracy. Spatiotemporal convolutional layers are mainly used to extract local spatial features from the data. In delivery demand prediction, this helps capture the spatial changes in delivery demand across different regions and at different times. Through convolutional operations, CNNs can identify patterns in spatial distribution, such as the volatility and concentration trends of delivery demand in a particular region. Attention mechanisms refer to models that "focus" on specific parts of the input data rather than treating all data evenly. In delivery demand forecasting, attention mechanisms help models identify which times or areas have more critical delivery demand and should be given priority consideration. The fully connected layer is the last layer in a neural network, used to map the preceding high-dimensional features to the final output. In this model, the fully connected layer summarizes the extracted spatiotemporal features and outputs the final prediction result, i.e., a heatmap of delivery demand over a future period.
[0099] Furthermore, the LSTM-CNN fusion network is trained using historical data to learn its temporal dependencies and spatial patterns. Once the model has converged, it can make efficient predictions based on new input data. When historical time-series data corresponding to the communities to be delivered is input into the trained model, the model outputs a heatmap of future delivery demand. This heatmap shows the occupancy of parcel locker resources over a future period, helping the system to allocate and optimize resources.
[0100] Through this prediction process, the system can automatically infer the future distribution of delivery demand based on historical time-series data corresponding to the communities to be delivered, helping the system to allocate, schedule, and optimize resources. For example, if the demand forecast for some areas is too high, the system can reserve more parcel locker resources in advance; conversely, if the demand for some areas is low, resources can be reduced or adjusted.
[0101] Step S30: Based on the delivery demand heat map, the space utilization rate of each express locker in the community to be delivered and the data of parcels to be allocated are obtained every second preset time interval. A preset dynamic resource scheduling model is used to determine the reward function value according to the space utilization rate of the express locker and the data of parcels to be allocated. The target express locker number to which the parcel to be allocated is assigned and the size and specifications of its corresponding compartment are determined according to the reward function value. The second preset time interval is less than the first preset time interval. The data of parcels to be allocated includes the volume, weight and priority of the parcels to be allocated.
[0102] The historical time series data is input into a delivery demand prediction model that has been trained to convergence. After predicting a delivery demand heatmap for the first preset time period, based on the delivery demand heatmap, the space utilization rate of each express locker in the community to be delivered and the data of parcels to be allocated are obtained every second preset time period. A preset dynamic resource scheduling model is used to determine the reward function value based on the express locker space utilization rate and the data of parcels to be allocated. Based on the reward function value, the target express locker number to which the parcel to be allocated is assigned and the size and specifications of its corresponding compartment are determined. The second preset time period is shorter than the first preset time period. The data of parcels to be allocated includes the volume, weight and priority of the parcels to be allocated. The second preset time period includes 4 to 6 hours, etc.
[0103] Specifically, the delivery demand heatmap is a prediction of the distribution of delivery resource demand over a relatively long period (a first preset duration, such as 24 hours, 48 hours, or 72 hours) based on historical data. However, in actual delivery, the utilization rate of parcel locker space and the number of parcels awaiting distribution are dynamic. Therefore, a shorter second preset duration (such as 4 to 6 hours) is needed to trigger a real-time status collection and scheduling closed loop to ensure that resource allocation matches the actual dynamic scenario and avoid the accumulation of discrepancies between prediction and reality.
[0104] Through IoT sensing and cabinet management system, the space utilization rate of each express cabinet in the community to be delivered can be obtained according to the second preset time period. The space utilization rate of the express cabinet represents the ratio of the volume of the used compartments to the total volume of the compartments.
[0105] The collected data on the space utilization rate of express delivery lockers and the data on parcels to be allocated (including volume, weight, and priority) are used as the state input of the dynamic resource scheduling model to construct a mapping relationship between physical resources on the supply side and parcel demands on the demand side, providing basic data for subsequent reward function calculation and strategy optimization.
[0106] The dynamic resource scheduling model includes a meta-learning model. This application uses a meta-learning model as an example, but this does not constitute any limitation on this application. Meta-learning, also known as "learning of learning" or "learning how to learn," is a method that enables machine learning systems to learn how to acquire knowledge from past experiences and apply this knowledge to new tasks. In traditional machine learning, models are trained on a large amount of task-specific data. The goal of meta-learning is to "learn" how to quickly adapt to new tasks, especially when the number of tasks is small or labeled data is scarce. The reward function corresponding to the dynamic resource scheduling model is the core optimization guide. The steps to construct the reward function of the dynamic resource scheduling model include:
[0107] Step S301: Obtain the first weight coefficient corresponding to the space utilization rate of the express lockers in the community to be allocated, the second weight coefficient corresponding to the average pickup distance of users, the third weight coefficient corresponding to the system energy consumption, and the fourth weight coefficient corresponding to the flexibility value for responding to sudden delivery needs.
[0108] Step S302: Calculate and determine the first product between the space utilization rate of the express cabinet and the first weight coefficient, the second product between the average pickup distance of the user and the second weight coefficient, calculate and determine the third product between the system energy consumption and the third weight coefficient, and calculate and determine the fourth product between the flexibility value for responding to sudden delivery needs and the fourth weight coefficient.
[0109] Step S303: Calculate and determine the first difference between the first product and the second product, calculate and determine the second difference between the first difference and the third product, and construct the dynamic resource scheduling model based on the first sum between the second difference and the fourth product.
[0110] Specifically, the dynamic resource scheduling model is a meta-learning model. It determines the target express locker number and corresponding compartment size of the package to be allocated based on the reward function value. The core is to dynamically optimize the weight coefficients of the reward function and the decision-making strategy through the meta-learning ability to "quickly adapt to new scenarios". The specific logic is as follows:
[0111] The meta-learning model is first pre-trained in multiple heterogeneous community scenarios (such as high-density residential areas, commercial areas, and mixed communities) to learn the mapping relationship between "different scenario features and reward function weight coefficients (first to fourth weight coefficients)" through a large amount of historical data. For example, in a commercial area scenario, the system learns that "the utilization rate of express delivery locker space (first weight coefficient) should be higher than the average user pickup distance (second weight coefficient)", while in an aging community, "the average user pickup distance (second weight coefficient) needs to be given higher priority", forming general meta-knowledge (i.e., "how to adjust weights according to the scenario").
[0112] For the current community awaiting delivery, the meta-learning model, based on real-time collected community features (such as population structure, density of parcel locker distribution, and frequency of historical sudden demand), calls upon pre-trained meta-knowledge to quickly adjust the first to fourth weight coefficients in step S301. For example, if it detects that the current community has recently experienced frequent e-commerce promotions (more sudden demand), it will increase the fourth weight coefficient (the flexibility value weight for responding to sudden demand), making the reward function more focused on adaptability to sudden scenarios.
[0113] Based on the adjusted weighting coefficients, the reward function value is calculated according to steps S302-S303: The first difference between the first product and the second product is determined by calculation; the second difference between the first difference and the third product is determined by calculation; and the first sum between the second difference and the fourth product is used to quantify the merits of the action of "allocating packages to be assigned to different parcel lockers and compartments". For example, in a certain scheme, if "the parcel locker space utilization rate is high (large first product), the user pickup distance is short (small second product, large first difference), the system energy consumption is low (small third product, large second difference), and there are many reserved emergency compartments (large fourth product, large first sum)," then its reward function value will be higher.
[0114] Furthermore, the meta-learning model, through a "small-scale trial-and-error - rapid iteration" mechanism, tests different "package-delivery locker-compartment" allocation combinations in the current community scenario. Based on the reward function value corresponding to each combination, it selects the solution with the highest reward value and outputs the target delivery locker number and compartment size specifications corresponding to that solution. At the same time, the current decision results are fed back into the meta-knowledge to further optimize the "scenario-weight" mapping relationship, improving the adaptation speed and accuracy of subsequent decisions.
[0115] In short, the meta-learning model learns general rules first and then quickly adapts to specific scenarios, ensuring that the reward function is always aligned with the core needs of the current community. Finally, through quantitative comparison of reward function values, it outputs the optimal target express cabinet number and compartment size specification allocation result for the packages to be assigned.
[0116] In some embodiments, the expression for the reward function of the dynamic resource scheduling model is:
[0117] R = α·η space -β·D user -γ·E consumption +δ·F flexibility ,
[0118] Where α represents the first weighting coefficient corresponding to the space utilization rate of the parcel locker, β represents the second weighting coefficient corresponding to the average pickup distance of users, γ represents the third weighting coefficient corresponding to the system energy consumption, and δ represents the fourth weighting coefficient corresponding to the flexibility value for responding to sudden delivery demands; η space The space utilization rate is represented by the ratio of the volume of used grid compartments to the total volume of grid compartments; D user E represents the average pickup distance for users, calculated as the average Euclidean distance from the user's residence to the designated parcel locker; consumption This indicates system energy consumption, including energy consumption for cooling / heating, communication, and control; F flexibility This represents the system's flexibility in responding to sudden delivery demands; it is the ratio between the available compartment volume of each parcel locker and the predicted demand compartment volume.
[0119] As can be seen from the above embodiments, based on the reward function of the above dynamic resource scheduling model, the advantages and disadvantages of the package allocation strategy are quantified, guiding the system to find a balance between "improving the utilization rate of express lockers, shortening the distance for users to pick up packages, reducing energy consumption, and enhancing adaptability to sudden scenarios", and finally determining the target express locker number and corresponding compartment specifications of the packages to be allocated.
[0120] Step S40: In response to the drone delivery command, a delivery path planning algorithm is used to perform global path planning based on the target parcel locker number assigned to the package to be delivered. Based on the improved contract network protocol and a preset bid value calculation model, during task allocation, the remaining battery power of the drone, the distance to the target parcel locker number, the estimated delivery time, and the deviation of the delivery of the package to be delivered from the drone's current task plan are comprehensively evaluated to generate a bid value to match the optimal drone for execution, so as to complete the unmanned delivery control in the community.
[0121] Based on the delivery demand heatmap, the space utilization rate of each express locker and the data of unassigned packages in the community to be delivered are acquired every second preset time interval. A preset dynamic resource scheduling model is used to determine the reward function value based on the space utilization rate of the express lockers and the data of unassigned packages. After determining the target express locker number and its corresponding compartment size specifications for the unassigned package based on the reward function value, the drone delivery command is responded to. A delivery path planning algorithm is used to perform global path planning based on the target express locker number for the unassigned package. Based on the improved contract network protocol and a preset bid value calculation model, during task allocation, the remaining power of the drone, the distance to the target express locker number, the estimated delivery time, and the deviation of the delivery of the unassigned package from the drone's current task plan are comprehensively evaluated to generate a bid value to match the optimal execution drone, so as to complete the unmanned delivery control of the community. The delivery path planning algorithm includes an improved A* algorithm.
[0122] In some embodiments, the improved A* algorithm of this application takes the target express locker number as the endpoint and comprehensively considers multiple constraints of drone flight (such as no-fly zones, obstacles, air traffic flow, and remaining battery threshold). It calculates the total path cost (including distance cost, energy cost, time cost, and safety redundancy cost) by constructing a "cost function". For example, in path search, it not only pursues the shortest physical distance but also avoids dangerous areas such as high-voltage power lines (increasing the weight of safety costs), prioritizes downwind routes (reducing energy costs), and reserves the remaining range redundancy corresponding to emergency battery power (ensuring return capability).
[0123] By iteratively searching for the optimal path nodes, the algorithm generates a global path from the drone's current location to the target parcel locker. This ensures path feasibility (no collisions, no no-fly zones) and achieves multi-objective optimization of "shorter distance, lower energy consumption, and better time efficiency," providing path guidance for drones to accurately, efficiently, and safely deliver parcels.
[0124] In some embodiments, the step of constructing a bid value calculation model includes:
[0125] Step S401: Obtain the fifth weighting coefficient corresponding to the remaining battery power of the drone, the sixth weighting coefficient corresponding to the distance to the target parcel locker number, the seventh weighting coefficient corresponding to the estimated delivery time, and the eighth weighting coefficient corresponding to the deviation of the delivery of the parcel from the drone's current task plan.
[0126] Step S402: Calculate and determine the fifth product between the remaining battery power of the drone and the fifth weighting coefficient; calculate and determine the first ratio between the sixth weighting coefficient and the distance to the target express cabinet number; calculate and determine the sixth product between the estimated delivery time and the seventh weighting coefficient; calculate and determine the seventh product between the deviation of the delivery of the parcel to be assigned from the current task plan of the drone and the eighth weighting coefficient.
[0127] Step S403: Construct the bid value calculation model based on the second sum between the fifth product, the first ratio, and the sixth product, and based on the third difference between the second sum and the seventh product.
[0128] Specifically, the expression for the bid value calculation model is as follows:
[0129] Bid i =w1·E remain +w2 / D task +w3·T completion -w4·C deviation ,
[0130] Where w1 represents the fifth weighting coefficient corresponding to the drone's remaining battery power; w2 represents the sixth weighting coefficient corresponding to the distance between the drone and the target parcel locker number; w3 represents the seventh weighting coefficient corresponding to the estimated delivery time; w4 represents the eighth weighting coefficient corresponding to the deviation of the delivery of parcels from the drone's current task plan; E remain Indicates the remaining battery power of the drone; D task Indicates the distance from the drone to the target parcel locker number; T completion Indicates the estimated delivery time; C deviation This indicates the degree of deviation of the delivery of packages to be assigned from the drone's current mission plan.
[0131] As shown in the above bid value calculation model, the system aggregates the bid values of all participating drones and selects the drone with the highest bid value as the optimal execution drone. If there are drones with the same bid value, further screening can be conducted through secondary evaluation (such as auxiliary indicators like historical task completion rate and equipment health). After matching, a delivery instruction containing the target parcel locker number and compartment specifications of the package to be delivered is issued to the selected drone, driving it to execute the task. Through this process, the bid value calculation model transforms abstract requirements such as "battery power, distance, timeliness, and plan stability" into quantifiable comparative indicators, achieving precise matching between drones and delivery tasks and improving the overall efficiency and reliability of unmanned delivery in the community.
[0132] In a further embodiment, after generating a bid value to match the optimal execution steps of the drone during task allocation based on a preset bid value calculation model, the following steps are included: (The model involves comprehensively evaluating the drone's remaining battery power, distance to the target parcel locker number, estimated delivery time, and deviation of the delivery of the unallocated package from the drone's current task plan.)
[0133] Step S4001: Construct the energy state space of the energy sharing network, wherein the state space includes the state of charge of the supercapacitor, the state of charge of the solid-state lithium battery, the load power, the ambient temperature, and the solar energy prediction value for a preset future duration.
[0134] Step S4002: Construct the energy action space of the energy sharing network, wherein the energy action space includes the supercapacitor output power, the lithium battery output power and the energy sharing power;
[0135] Step S4003: Based on the deep deterministic policy gradient algorithm, energy scheduling is performed according to the energy state space, with the optimization objective of maximizing a preset energy management reward function, to output the energy action space, wherein the expression of the energy management reward function is:
[0136] R = 0.7·(1-|SOC) bat -0.5|)+0.2·η system -0.1·P share ,
[0137] Where, η system P represents the overall energy efficiency of the system. share State of Charge (SOC) represents the amount of energy transferred between the drone and the parcel locker, or between different energy storage units. bat This indicates the state of charge of a solid-state lithium battery.
[0138] Specifically, the energy scheduling process based on the Deep Deterministic Policy Gradient (DDPG) algorithm involves integrating deep neural networks with the deterministic policy gradient method to continuously optimize the strategy in a continuous energy state space and action space to maximize the energy management reward function. The energy state space (S) serves as the input to the DDPG algorithm and contains key state parameters related to the system's current energy, such as: supercapacitor state of charge (SOC_cap), solid-state lithium battery state of charge (SOC_bat), real-time load power (P_load), ambient temperature (T_ambient), and solar energy prediction for the next hour (E_solar_pred). These parameters collectively reflect the system's current energy reserves, supply and demand relationship, and environmental impact, providing a "current situation basis" for scheduling decisions. The energy state space (S) is represented as S = [SOC_cap, SOC_bat, P_load, T_ambient, E_solar_pred].
[0139] The energy action space (A), as the output of the DDPG algorithm, is represented as A = [P_cap, P_bat, P_share]. It contains the continuous energy scheduling actions that the system can execute, such as: supercapacitor output power (P_cap)), lithium battery output power (P_bat)), and energy sharing power (P_share). These actions directly determine the distribution and flow of energy among different energy storage units and devices, and serve as the "operational carrier" for achieving scheduling.
[0140] The energy management reward function (R) serves as the optimization objective, quantifying the merits and demerits of the current action. Through positive rewards (such as maintaining the high-efficiency state of charge of lithium batteries and improving system energy efficiency) and negative penalties (such as losses caused by excessive energy sharing), the energy management reward function guides the algorithm towards optimization in the direction of "energy saving, equipment protection, and efficient collaboration".
[0141] As demonstrated by the above embodiments, the Deep Deterministic Policy Gradient (DDPG) algorithm can output a near-optimal energy action space (A) based on the real-time energy state space (S). For example, during peak load periods and when solar energy is abundant, it prioritizes the use of solar power and supercapacitors to reduce lithium battery losses; during low load periods, it balances the state of charge of each energy storage unit through energy sharing. Ultimately, these actions maximize the reward function (R) to achieve the goals of "high efficiency, low consumption, and device friendliness" in energy scheduling. The DDPG algorithm learns the optimal scheduling strategy autonomously in the continuous energy state and action space through a closed loop of "state awareness - action output - reward feedback - network iteration," enabling the system to dynamically adapt to the complex changes in energy supply and demand in unmanned delivery scenarios in communities.
[0142] As can be seen from the above embodiments, compared with the prior art, the present application addresses the problems commonly faced by existing community unmanned delivery control methods, such as low efficiency, uneven resource allocation, frequent coordination conflicts, insufficient energy management, and weak system robustness. The present application has, but is not limited to, the following beneficial effects:
[0143] Firstly, the community unmanned delivery control method of this application can greatly improve delivery efficiency, with a 172.7% increase in last-mile delivery efficiency. Compared with the traditional model, this application achieves a significant improvement in delivery efficiency. By rationally allocating the space of the express cabinet, optimizing the drone path planning and task allocation, the delivery time is greatly shortened, improving the response speed and timeliness of delivery.
[0144] Secondly, the community unmanned delivery control method proposed in this application improves the space utilization rate of express delivery lockers by 93.6%. It adopts dual-timescale reinforcement learning to optimize the storage space of express delivery lockers, making the space use of express delivery lockers more reasonable and efficient, avoiding the problems of resource waste and insufficient space, and improving the storage capacity of each express delivery locker.
[0145] Third, the community unmanned delivery control method of this application can reduce resource conflicts and energy consumption, and reduce the drone task conflict rate by 90.6%. Through the improved A* algorithm and contract network protocol, the system effectively resolves conflicts in multi-drone collaboration, ensures smooth multi-drone task allocation, and improves the overall delivery collaboration efficiency.
[0146] Fourth, the community unmanned delivery control method of this application reduces energy consumption by 44.3%. The application of the deep deterministic policy gradient algorithm in energy scheduling effectively optimizes energy use. Through intelligent scheduling and control, unnecessary energy consumption is reduced, resulting in a significant improvement in the system's energy efficiency compared to traditional methods.
[0147] Fourth, the community unmanned delivery control method proposed in this application enhances the system's robustness and sustainability. It has a 96-120 hour off-grid operation capability, and the hybrid energy supply device integrates multiple energy harvesting and storage technologies, ensuring the system can operate continuously and stably for extended periods without grid support. This feature greatly enhances the system's robustness, especially in extreme situations such as power outages, guaranteeing the reliability of the unmanned delivery system.
[0148] Fifth, the community unmanned delivery control method of this application integrates an innovative architecture of distributed intelligent express cabinet clusters, drone swarms, and a collaborative control platform, enabling the system to flexibly respond to various delivery needs and complex environmental changes, exhibiting stronger adaptability and stability. Accurate delivery demand prediction is achieved through an LSTM-CNN fusion network. Utilizing historical data and multi-dimensional features (such as allocated package data, demographic data, and meteorological data), future delivery can be effectively predicted. Based on the prediction results and dynamic express cabinet space utilization, the system employs a dynamic resource scheduling model to intelligently allocate packages, ensuring maximum resource utilization while meeting various requirements such as package volume, weight, and priority.
[0149] In summary, this application demonstrates significant advantages in improving last-mile delivery efficiency, resource utilization, collaborative operations, energy management, and system robustness. It addresses bottlenecks in traditional delivery models by optimizing delivery processes, increasing resource utilization, and reducing energy consumption, thereby significantly lowering overall operating costs. This more efficient delivery model not only reduces enterprise costs but also enhances the consumer experience, driving socio-economic benefits. Intelligent energy scheduling and energy consumption optimization further reduce carbon emissions, aligning with green and sustainable development requirements, contributing to environmental protection, and possessing broad application prospects and substantial socio-economic benefits.
[0150] Please see Figure 2A community unmanned delivery control device, provided to meet one of the purposes of this application, includes a data acquisition module 1100, a delivery heatmap prediction module 1200, a dynamic resource scheduling module 1300, and an unmanned delivery control module 1400. The data acquisition module 1100 is configured to respond to a community unmanned delivery control command and acquire historical time-series data corresponding to the community to be delivered. The historical time-series data includes delivery feature vectors corresponding to multiple historical time steps, including allocated parcel data, community population statistics, meteorological data, point-of-interest spatial data, and e-commerce promotion time data. The delivery heatmap prediction module 1200 is configured to input the historical time-series data into a delivery demand prediction model trained to convergence to predict a delivery demand heatmap for a first preset time period. The delivery demand heatmap represents the distribution of express locker resources required for allocated parcels. The dynamic resource scheduling module 1300 is configured to, based on the delivery demand heatmap, acquire the space utilization rate of each express locker in the community to be delivered and the data of parcels to be allocated every second preset time period, and use a preset dynamic resource scheduling method. The scheduling model determines a reward function value based on the space utilization rate of the parcel locker and the data of parcels to be assigned. Based on the reward function value, it determines the target parcel locker number to which the parcel to be assigned is located and the corresponding compartment size. The second preset time is less than the first preset time. The data of parcels to be assigned includes the volume, weight, and priority of the parcels. The unmanned delivery control module 1400 is configured to respond to drone delivery commands. It employs a delivery path planning algorithm to perform global path planning based on the target parcel locker number assigned to the parcel. Based on an improved contract network protocol and a preset bid value calculation model, during task allocation, it comprehensively evaluates the drone's remaining battery power, the distance to the target parcel locker number, the estimated delivery time, and the deviation of the parcel delivery from the drone's current task plan. This generates a bid value to match the optimal drone for execution, thereby completing the unmanned delivery control in the community.
[0151] Based on any embodiment of this application, please refer to Figure 3 Another embodiment of this application also provides an electronic device, which can be implemented by a computer device, such as... Figure 3The diagram shows the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. The computer-readable storage medium stores an operating system, a database, and computer-readable instructions. The database may store a sequence of control information. When the computer-readable instructions are executed by the processor, the processor can implement a community unmanned delivery control method. The processor of the computer device provides computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute the community unmanned delivery control method of this application. The network interface of the computer device is used for communication with a terminal. Those skilled in the art will understand that… Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0152] In this embodiment, the processor is used to execute... Figure 2 The specific functions of each module are defined within the device, and the memory stores the program code and various data required to execute these modules or sub-modules. The network interface is used for data transmission between the user terminal and the server. In this embodiment, the memory stores the program code and data required to execute all modules in the community unmanned delivery control device of this application, and the server can call the server's program code and data to execute the functions of all modules.
[0153] This application also provides a storage medium storing computer-readable instructions, which, when executed by one or more processors, cause the one or more processors to perform the steps of the community unmanned delivery control method described in any embodiment of this application.
[0154] This application also provides a computer program product, including a computer program / instructions that, when executed by one or more processors, implement the steps of the community unmanned delivery control method described in any embodiment of this application.
[0155] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments of this application can be implemented by a computer program instructing related hardware. This computer program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. The aforementioned storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0156] The above description is only a partial embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for controlling unmanned delivery in a community, characterized in that, include: In response to the community unmanned delivery control command, the system obtains the historical time series data corresponding to the community to be delivered. The historical time series data includes delivery feature vectors corresponding to multiple historical time steps. The delivery feature vectors include allocated package data, community population statistics, meteorological data, point of interest spatial data, and e-commerce promotion time data. The historical time series data is input into the delivery demand prediction model that has been trained to convergence, so as to predict the delivery demand heat map for the first preset time period in the future, wherein the delivery demand heat map represents the distribution of express cabinet resources required for the parcels to be allocated. Based on the delivery demand heatmap, the space utilization rate of each express locker in the community to be delivered and the data of parcels to be allocated are obtained every second preset time interval. A preset dynamic resource scheduling model is used to determine the reward function value based on the space utilization rate of the express locker and the data of parcels to be allocated. Based on the reward function value, the target express locker number to which the parcel to be allocated is assigned and the size and specifications of its corresponding compartment are determined. The second preset time interval is less than the first preset time interval. The data of parcels to be allocated includes the volume, weight and priority of the parcels to be allocated. In response to drone delivery instructions, a delivery path planning algorithm is used to perform global path planning based on the target parcel locker number assigned to the package to be delivered. Based on an improved contract network protocol and a preset bid value calculation model, during task allocation, the remaining battery power of the drone, the distance to the target parcel locker number, the estimated delivery time, and the deviation of the delivery of the package to be delivered from the drone's current task plan are comprehensively evaluated to generate a bid value to match the optimal drone for execution, thereby completing the unmanned delivery control in the community.
2. The community unmanned delivery control method according to claim 1, characterized in that, Before obtaining the historical time-series data corresponding to the communities to be delivered, the following steps are included: The collected raw data is preprocessed, including allocated package data, community population statistics, meteorological data, point-of-interest spatial data, and e-commerce promotion time data. The original data is unified to a spatial grid of a preset size and a preset time granularity for spatiotemporal alignment, and Z-score normalization is used for feature standardization. The preset size includes 300m×300m or 500m×500m, and the preset time granularity includes 1 hour or 2 hours. The spatiotemporal K-nearest neighbor algorithm was used for missing value imputation, and the isolated forest algorithm was used for anomaly detection. The preprocessed raw data was used to construct the historical time series data corresponding to the community to be delivered, thus completing the preprocessing of the raw data.
3. The community unmanned delivery control method according to claim 1, characterized in that, The steps include: The historical time series data of the communities to be delivered is converted into a 5-dimensional spatiotemporal feature tensor, wherein the 5-dimensional spatiotemporal feature tensor includes batch size, time dimension, spatial grid dimension and feature channel dimension; The 5-dimensional spatiotemporal feature tensor is input into the delivery demand prediction model that has been trained to convergence. The time series dependencies are captured sequentially through bidirectional LSTM units, local spatiotemporal patterns are extracted through spatiotemporal convolutional layers with 3D convolutional kernels, feature importance is dynamically adjusted through a spatial attention weight matrix activated by Sigmoid, and finally a heatmap of delivery demand for the first preset duration is generated through a fully connected layer.
4. The community unmanned delivery control method according to claim 1, characterized in that, The steps for constructing the reward function of a dynamic resource scheduling model include: The system obtains a first weighting coefficient corresponding to the space utilization rate of the express lockers in the community to be allocated, a second weighting coefficient corresponding to the average pickup distance of users, a third weighting coefficient corresponding to the system energy consumption, and a fourth weighting coefficient corresponding to the flexibility value in response to sudden delivery needs. The system calculates and determines the first product between the space utilization rate of the express locker and the first weighting coefficient, the second product between the average pickup distance of the user and the second weighting coefficient, the third product between the system energy consumption and the third weighting coefficient, and the fourth product between the flexibility value for responding to sudden delivery needs and the fourth weighting coefficient. A first difference between the first product and the second product is calculated and determined, a second difference between the first difference and the third product is calculated and determined, and the dynamic resource scheduling model is constructed based on the first sum between the second difference and the fourth product.
5. The community unmanned delivery control method according to claim 1, characterized in that, The steps for constructing a bid value calculation model include: The fifth weighting coefficient corresponding to the remaining battery power of the drone, the sixth weighting coefficient corresponding to the distance to the target parcel locker number, the seventh weighting coefficient corresponding to the estimated delivery time, and the eighth weighting coefficient corresponding to the deviation of the delivery of the parcel from the drone's current task plan are obtained. The fifth product between the remaining battery power of the drone and the fifth weighting coefficient is calculated; the first ratio between the sixth weighting coefficient and the distance to the target express cabinet number is calculated; the sixth product between the estimated delivery time and the seventh weighting coefficient is calculated; and the seventh product between the deviation of the delivery of the parcel to be assigned from the current task plan of the drone and the eighth weighting coefficient is calculated. The bid value calculation model is constructed based on the second sum between the fifth product, the first ratio, and the sixth product, and based on the third difference between the second sum and the seventh product.
6. The community unmanned delivery control method according to claim 1, characterized in that, Based on a pre-defined bid value calculation model, during task allocation, the remaining battery power of the drone, the distance to the target parcel locker number, the estimated delivery time, and the deviation of the delivery of the undelivered package from the drone's current task plan are comprehensively evaluated to generate a bid value to match the optimal steps for drone execution. This includes: Construct an energy state space for an energy sharing network, wherein the state space includes the state of charge of supercapacitors, the state of charge of solid-state lithium batteries, load power, ambient temperature, and solar energy prediction values for a preset future duration. The energy action space for constructing an energy sharing network includes supercapacitor output power, lithium battery output power, and energy sharing power; Based on the deep deterministic policy gradient algorithm, energy scheduling is performed according to the energy state space, with the optimization objective of maximizing a preset energy management reward function, to output the energy action space. The expression for the energy management reward function is: R=0.7.(1-|SOC bat -0.5|)+0.2·η system -0.1·P share , Where, η system P represents the overall energy efficiency of the system. share State of Charge (SOC) represents the amount of energy transferred between the drone and the parcel locker, or between different energy storage units. bat This indicates the state of charge of a solid-state lithium battery.
7. The community unmanned delivery control method according to any one of claims 1 to 6, characterized in that, The allocated package data includes the volume and weight of the allocated package, as well as the express locker number where the allocated package was placed and the size of its corresponding compartment. The point of interest spatial data represents the spatial distribution and type information of the community to be delivered and the surrounding commercial areas, residential areas and public service facilities. The basic network architecture of the delivery demand prediction model is an LSTM-CNN fusion network, which includes an input layer, an LSTM layer, a spatiotemporal convolutional layer, an attention mechanism, and a fully connected layer. The LSTM layer includes multiple bidirectional LSTM units. The delivery route planning algorithm includes an improved A* algorithm.
8. A community unmanned delivery control device, characterized in that, include: The data acquisition module is configured to respond to community unmanned delivery control commands and acquire historical time series data corresponding to the community to be delivered. The historical time series data includes delivery feature vectors corresponding to multiple historical time steps. The delivery feature vectors include allocated package data, community population statistics, meteorological data, point of interest spatial data, and e-commerce promotion time data. The delivery heat map prediction module is configured to input the historical time series data into a delivery demand prediction model that has been trained to a convergent state, so as to predict the delivery demand heat map for the first preset time period in the future, wherein the delivery demand heat map represents the distribution of express cabinet resources required for the parcels to be allocated; The dynamic resource scheduling module is configured to acquire the space utilization rate of each express locker and the data of parcels to be allocated in the community to be delivered every second preset time interval based on the delivery demand heat map. The module uses a preset dynamic resource scheduling model to determine the reward function value based on the express locker space utilization rate and the data of parcels to be allocated. Based on the reward function value, the module determines the target express locker number to which the parcels to be allocated are assigned and the size specifications of their corresponding compartments. The second preset time interval is shorter than the first preset time interval. The data of parcels to be allocated includes the volume, weight and priority of the parcels to be allocated. The unmanned delivery control module is configured to respond to drone delivery commands. It employs a delivery path planning algorithm to perform global path planning based on the target parcel locker number assigned to the package to be delivered. Based on an improved contract network protocol and a preset bid value calculation model, during task allocation, it comprehensively evaluates the drone's remaining battery power, distance to the target parcel locker number, estimated delivery time, and deviation of the package to be delivered from the drone's current task plan to generate a bid value to match the optimal drone for execution, thereby completing the unmanned delivery control in the community.
9. An electronic device comprising a central processing unit and a memory, characterized in that, The central processing unit is used to invoke and run a computer program stored in the memory to perform the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores, in the form of computer-readable instructions, a computer program implemented according to any one of claims 1 to 7, which, when invoked by a computer, executes the steps included in the corresponding method.
Citation Information
Cited By
Communication linkage scheduling method of logistics unmanned aerial vehicle and intelligent turnover cabinet
CN122372926A
Communication linkage scheduling method of logistics unmanned aerial vehicle and intelligent turnover cabinet
CN122372926B