Efficient medical logistics order scheduling and distribution method and system

Through multimodal drug feature extraction, adaptive space-time distribution network and double-layer heterogeneous graph neural network, the problems of errors in drug storage requirements and low distribution efficiency in drug flow are solved, and efficient and safe drug flow order scheduling and distribution are achieved.

CN120387760AInactive Publication Date: 2025-07-29ZHONGJIAN YUNKANG (GUANGZHOU) LOGISTICS SUPPLY CHAIN CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510460722.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the existing medical flow order scheduling and distribution, errors are prone to judgments on drug storage requirements and distribution restrictions. The traditional distribution area division method cannot be adjusted according to environmental factors, resulting in hidden dangers in drug quality and safety control, low distribution efficiency and increased operating costs.

Method used

Build a multi-modal drug feature extraction module, integrate drug images, texts and storage condition information, and generate drug feature vectors; establish an adaptive space-time distribution network for dynamic grid division; build a two-layer heterogeneous graph neural network for path planning, set up a reinforcement learning scheduling system, and update the scheduling strategy through federated learning.

Benefits of technology

It improves the accuracy of identification of drug storage requirements and distribution restrictions, improves the flexibility of distribution resource allocation and the rationality of path planning, and achieves the balance and optimization of drug quality maintenance and delivery timeliness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387760A_ABST
    Figure CN120387760A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient medicine logistics order scheduling and distribution method and system, and relates to the technical field of medicine logistics, and the method comprises the steps: constructing a multi-modal medicine feature extraction module, and outputting an order constraint matrix; establishing an adaptive space-time distribution network according to the order constraint matrix, and performing dynamic grid division on a distribution area to generate a multi-level distribution area model; constructing a double-layer heterogeneous graph neural network, and generating a distribution path planning candidate scheme set considering drug characteristics; defining a state space and an action space, constructing a dual-target reward function, and carrying out iterative training to obtain a distribution scheduling strategy model; and aggregating multi-region distribution empirical data by adopting a federated learning framework, and periodically updating the scheduling strategy model based on a differential privacy mechanism. According to the method, the multi-modal drug feature extraction module is constructed, and the image, text and storage condition information of the drug are fused, so that the drug storage requirement and the distribution limitation condition can be identified more accurately, and the identification accuracy of the distribution constraint is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of pharmaceutical logistics, in particular to an efficient pharmaceutical logistics order scheduling and distribution method and system. Background Art

[0002] Currently, the scheduling and distribution of pharmaceutical logistics orders mainly rely on manual operations, and it is easy to make mistakes in judging the storage requirements of drugs and distribution restrictions. The extraction of features such as the images, instruction manual texts, and storage conditions of drugs is not comprehensive enough to meet the differentiated distribution needs of different drugs. In addition, the traditional distribution area division method is relatively fixed and cannot be adjusted in a timely manner according to environmental factors such as road conditions and temperature. The existing distribution route planning also fails to fully consider the requirements of drug timeliness and storage condition constraints, and the scheduling strategy often focuses too much on a single goal and is difficult to balance drug quality and distribution efficiency.

[0003] These problems pose multiple challenges to the actual operation of pharmaceutical logistics. On the one hand, there are potential hazards in the control of drug quality and safety, and some temperature-sensitive drugs may have their efficacy affected due to improper storage and transportation conditions. On the other hand, the distribution efficiency is not ideal enough to meet the timely distribution requirements of urgently needed drugs. At the same time, the utilization efficiency of distribution resources is not high, resulting in an increase in operating costs. In addition, due to the lack of effective accumulation and utilization of distribution experience data, it is difficult to continuously improve the quality of distribution services. Summary of the Invention

[0004] In view of the problems existing in the existing pharmaceutical logistics order scheduling and distribution, the present invention is proposed.

[0005] Therefore, the problem to be solved by the present invention is how to achieve efficient scheduling and safe distribution of pharmaceutical logistics orders.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] In a first aspect, an embodiment of the present invention provides an efficient pharmaceutical logistics order scheduling and distribution method, which includes constructing a multimodal drug feature extraction module. The multimodal drug feature extraction module generates a drug feature vector by fusing multimodal drug information, identifies drug storage requirements and distribution restriction conditions based on the drug feature vector, and outputs an order constraint matrix; establishes an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input and the environmental features as the auxiliary input, dynamically divides the distribution area into grids, and generates a multi-level distribution area model; constructs a double-layer heterogeneous graph neural network, maps order nodes and distribution area nodes to a graph data structure based on the multi-level distribution area model, performs graph convolutional operations to extract the correlation features between nodes, sets drug timeliness weights and storage condition weights in the graph convolutional layer, and generates a candidate set of distribution path planning considering drug characteristics; sets up a reinforcement learning scheduling system, uses the temperature field state features and road network state features as the state space, takes the candidate set of distribution path planning as the action space, constructs a double-objective reward function, and iteratively trains through the policy gradient algorithm to obtain a distribution scheduling policy model; collects distribution environment data on the distribution vehicle equipped with an edge computing unit, aggregates multi-region distribution experience data using a federated learning framework, and periodically updates the scheduling policy model based on the differential privacy mechanism.

[0008] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method of the present invention, the construction of the multimodal drug feature extraction module includes the following steps: collecting drug image data, extracting drug appearance features through a hierarchical attention visual network of a drug packaging identifier to generate a drug image feature matrix; performing word segmentation on the drug instruction manual text, and using a domain-adaptive BERT language model integrating a drug knowledge graph to extract drug key attribute information to generate a drug text feature matrix; establishing a drug storage condition quantization module, performing a linear interval mapping on drug storage parameters to generate a drug storage feature matrix; constructing a multimodal feature fusion network, inputting the drug image feature matrix, the drug text feature matrix, and the drug storage feature matrix into a feature fusion layer, integrating different modal features through a feature weighted fusion mechanism, and outputting a drug feature vector; designing a bidirectional attention network, taking the drug feature vector as the query input, taking a preset distribution condition library as the key-value input, using a multi-head attention structure to extract drug distribution requirements, and setting up an independent timeliness evaluation module to generate drug distribution restriction features; establishing a constraint matrix generation module, combining the drug distribution restriction features with the order basic information, and forming an order constraint matrix through a gradient constraint propagation algorithm.

[0009] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method described in the present invention, wherein: the environmental features include road condition sensing data, temperature monitoring data, and population density data of the distribution area; establishing an adaptive spatio-temporal allocation network according to the order constraint matrix includes the following steps: constructing a spatio-temporal constraint analysis module to analyze the order constraint matrix and form a spatio-temporal constraint intensity distribution map of the distribution area; designing a multi-source road condition perception network to collect road condition sensing data through a distributed sensor array, and combining the spatio-temporal distribution of medical institutions to establish the road network state characteristics; constructing a temperature field dynamic modeling network to model the regional temperature distribution using a temperature interpolation network based on key monitoring points, incorporating the temperature threshold constraints for drug storage and transportation, and generating the temperature field state characteristics; designing a population activity predictor to establish a regional activity prediction model based on the population density data of the distribution area and output the population activity state characteristics; constructing a constraint-driven adaptive feature extractor to perform non-linear feature transformation on the road network state characteristics, temperature field state characteristics, and population activity state characteristics to generate a time-varying environmental feature field; superimposing the constraint intensity distribution map and the time-varying environmental feature field, and determining the grid deformation direction by solving the variational equation for minimizing the distribution cost to obtain the adaptive grid division result; performing hierarchical processing on the adaptive grid division result, and performing hierarchical merging according to the regional adjacency relationship and environmental feature similarity to output a multi-level distribution area model.

[0010] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method described in the present invention, wherein: constructing a two-layer heterogeneous graph neural network includes the following steps: decomposing the order constraint matrix, extracting key constraint parameters, and constructing a multi-dimensional feature vector of the order node; parsing the multi-level distribution area model to obtain the spatio-temporal features and environmental attributes of the regional node, and constructing a multi-dimensional feature vector of the regional node; designing a graph structure generator for drug constraint perception to calculate the node connection strength according to the constraint matching degree between the order node and the regional node, and establishing a two-layer heterogeneous graph with drug distribution constraints; constructing a time-effectiveness adaptive edge weight calculation module to convert the drug distribution time-effectiveness requirement into an edge weight coefficient; constructing a storage condition constraint factor calculation module to generate a constraint factor matrix according to the drug storage condition requirements and regulate the feature transfer in the graph convolution operation; adopting a graph attention mechanism to set different weights for the feature aggregation channels based on the constraint factor matrix and update the node representation; designing a path evaluation module based on the node representation to comprehensively consider the time-effectiveness matching degree, storage condition satisfaction degree, and path feasibility, and generating a candidate set of distribution path planning schemes.

[0011] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method of the present invention, the steps for setting up the reinforcement learning scheduling system are as follows: Integrate the temperature field state characteristics and the road network state characteristics to generate a distribution environment state matrix; Map the distribution path planning candidate solution set into an action feature vector, and construct an action description matrix including path attributes and pharmaceutical distribution requirements; Design a dual-objective evaluation network architecture based on the distribution environment state matrix and the action description matrix, including a quality value network for evaluating the temperature control quality of pharmaceuticals and an efficiency value network for evaluating the distribution timeliness. The evaluation results of the two networks jointly guide the policy network to generate distribution path planning decisions; Design a dual-objective reward function based on pharmaceutical characteristics, including a quality maintenance reward function based on temperature deviation and light exposure, and a timeliness performance reward function based on the distribution on-time rate; Construct a policy gradient algorithm framework, iteratively optimize the parameters of the quality value network and the efficiency value network based on the dual-objective reward function, and guide the update of the policy network parameters to train a distribution scheduling policy model.

[0012] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method of the present invention, the evaluation process of the quality value network is as follows: Extract the environment state characteristics from the distribution environment state matrix through a convolutional neural network, and extract the path planning characteristics from the action description matrix through a graph convolutional network; Design an environment-path interaction module to perform attention-weighted fusion of the environment state characteristics and the path planning characteristics to generate pharmaceutical quality impact characteristics; Construct a quality impact evaluation module to process the pharmaceutical quality impact characteristics, where the quality impact evaluation module includes a difference calculation unit, a spatio-temporal mapping unit, and an accumulation effect unit; Among them, the difference calculation unit detects the mutation degree of environmental parameters, the spatio-temporal mapping unit evaluates the environmental exposure duration of pharmaceuticals in different path segments, and the accumulation effect unit calculates the long-term impact of environmental fluctuations based on the environmental exposure duration; Weight the output result of the quality impact evaluation module based on the quality sensitivity of the pharmaceutical type to obtain a quality evaluation score.

[0013] As a preferred solution of the efficient pharmaceutical logistics order scheduling and distribution method of the present invention, the specific formula of the dual-objective reward function is as follows:

[0014] R = λ q ·R q +λ t ·R t

[0015] Where λ q 、λ t are the weight coefficients of the quality maintenance reward function and the timeliness performance reward function, R q is the quality maintenance reward function, R t is the timeliness performance reward function,

[0016] R q = Qv -η1·∫|D(T t )|dt - η2·L

[0017]

[0018] where Q v is the quality evaluation score output by the quality value network, T t is the actual temperature at time t during transportation, L is the cumulative value of light exposure, η1 and η2 are the penalty coefficients for temperature deviation integration and light exposure, D(T t ) is the temperature deviation function, E v is the efficiency evaluation score output by the efficiency value network, t actual is the actual delivery time, t required is the required delivery time, Δt is the delivery delay time, and γ1 and γ2 are the penalty coefficients for time ratio and delay time.

[0019] In a second aspect, an embodiment of the present invention provides an efficient pharmaceutical logistics order scheduling and delivery system, which includes a feature extraction module for constructing a multimodal drug feature extraction module. The multimodal drug feature extraction module generates a drug feature vector by fusing multimodal drug information, identifies drug storage requirements and delivery restriction conditions based on the drug feature vector, and outputs an order constraint matrix; a regional allocation module for establishing an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input and the environmental features as the auxiliary input, dynamically divides the delivery area into grids, and generates a multi-level delivery area model; a path planning module for constructing a double-layer heterogeneous graph neural network, mapping order nodes and delivery area nodes to a graph data structure based on the multi-level delivery area model, performing graph convolution operations to extract the correlation features between nodes, setting the drug timeliness weight and storage condition weight in the graph convolution layer, and generating a candidate set of delivery path planning considering drug characteristics; a scheduling strategy module for setting up a reinforcement learning scheduling system, using the temperature field state feature and road network state feature as the state space, taking the candidate set of delivery path planning as the action space, constructing a reward function based on the drug quality maintenance index and delivery timeliness rate, and obtaining a delivery scheduling strategy model through iterative training of the policy gradient algorithm; a model update module for collecting delivery environment data on the delivery vehicle equipped with an edge computing unit, aggregating multi-region delivery experience data using a federated learning framework, and performing periodic updates on the scheduling strategy model based on the differential privacy mechanism.

[0020] The beneficial effects of the present invention are as follows: By constructing a multi-modal drug feature extraction module and integrating the image, text, and storage condition information of drugs, the present invention can more accurately identify the drug storage requirements and distribution restriction conditions, improving the accuracy of identifying distribution constraints; The designed adaptive spatio-temporal allocation network effectively improves the flexibility of distribution resource allocation through a dynamic grid division method; The double-layer heterogeneous graph neural network is used for order-region matching, and drug timeliness weights and storage condition weights are introduced in the graph convolution layer, improving the rationality of path planning; The reinforcement learning scheduling system based on the double-objective evaluation network achieves the balanced optimization of drug quality maintenance and distribution timeliness, enabling the distribution plan to better meet the special needs of pharmaceutical logistics. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a framework flowchart of an efficient pharmaceutical logistics order scheduling and distribution method.

[0023] Figure 2 It is a flowchart for constructing a multi-modal drug feature extraction module of an efficient pharmaceutical logistics order scheduling and distribution method.

[0024] Figure 3 It is a flowchart for establishing an adaptive spatio-temporal allocation network of an efficient pharmaceutical logistics order scheduling and distribution method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the above-mentioned objects, features, and advantages of the present invention more obvious and understandable, the following will provide a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.

[0026] In the following description, many specific details are set forth to fully understand the present invention. However, the present invention can also be implemented in other ways different from the description herein. Those skilled in the art can make similar generalizations without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.

[0027] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that mutually excludes other embodiments.

[0028] Embodiment 1, refer toFigures 1 to 3 , which is the first embodiment of the present invention. This embodiment provides an efficient pharmaceutical logistics order scheduling and distribution method. The framework flowchart is as shown in Figure 1 , including

[0029] S1: Construct a multi-modal drug feature extraction module. The multi-modal drug feature extraction module generates a drug feature vector by fusing drug image information, instruction text information, and storage condition information, and applies a bidirectional attention mechanism to identify drug storage requirements and distribution restriction conditions, and outputs an order constraint matrix.

[0030] Specifically, the flowchart for constructing the multi-modal drug feature extraction module is as shown in Figure 2 , including the following steps:

[0031] S1.1: Collect drug image data, and extract the drug appearance features through the hierarchical attention visual network of the drug packaging identifier to generate a drug image feature matrix.

[0032] Among them, the hierarchical attention visual network of the fused drug packaging identifier includes a drug feature extraction layer and a packaging integrity recognition layer. The drug feature extraction layer is provided with a multi-scale feature pyramid structure, and the packaging integrity recognition layer uses an object detection algorithm to locate packaging defects, and integrates the drug appearance features through a feature fusion network to output a drug image feature matrix.

[0033] In this embodiment, the drug appearance features include at least one of the following: drug packaging integrity feature, volume size feature, storage state feature, packaging material feature, and appearance color feature.

[0034] S1.2: Perform word segmentation on the drug instruction text, and use the domain-adaptive BERT language model integrating the drug knowledge graph to extract the key attribute information of the drug to generate a drug text feature matrix.

[0035] Among them, the domain-adaptive BERT language model integrating the drug knowledge graph includes a drug professional dictionary matching layer, a drug knowledge graph enhancement layer, and a drug attribute extraction layer. The drug professional dictionary matching layer identifies medical professional terms in the drug instruction; the drug knowledge graph enhancement layer integrates drug entity relationship information and attribute information, extracts features of key information such as storage conditions and usage restrictions, and generates a drug text feature matrix.

[0036] In one embodiment, the drug key attribute information includes the specified storage temperature range of the drug, light avoidance requirement, moisture-proof requirement, drug validity period, and packaging specification information.

[0037] S1.3: Establish a drug storage condition quantification module, perform a linear interval mapping on the drug storage parameters, and generate a drug storage feature matrix.

[0038] Among them, the drug storage parameters are selected from at least one of the temperature range, humidity range, light requirement, pressure requirement, and shockproof requirement.

[0039] S1.4: Construct a multi-modal feature fusion network, input the drug image feature matrix, drug text feature matrix, and drug storage feature matrix into the feature fusion layer, integrate different modal features through a feature weighted fusion mechanism, and output the drug feature vector.

[0040] S1.5: Design a bidirectional attention network, use the drug feature vector as the query input, use the preset delivery condition library as the key-value input, adopt a multi-head attention structure to extract the drug delivery requirements, and set up an independent timeliness evaluation module to generate drug delivery restriction features.

[0041] Among them, the multi-head attention structure encodes the features of delivery restriction conditions such as the drug storage temperature range and transportation timeliness; the timeliness evaluation module calculates the weight coefficients respectively based on the drug expiration date and delivery time, dynamically weights the delivery condition features, and generates drug delivery restriction features. In addition, the preset delivery condition library is stored in a relational database, including a delivery basic condition table, a historical delivery record table, and a delivery evaluation table, establishes a data relationship through foreign key association, and sets up a regular update mechanism to maintain data timeliness.

[0042] In one embodiment, the drug delivery requirements include delivery timeliness requirements, temperature control requirements, light avoidance requirements, transportation attitude requirements, and shockproof requirements.

[0043] S1.6: Establish a constraint matrix generation module, combine the drug delivery restriction features with the order basic information, and form an order constraint matrix through the gradient constraint propagation algorithm.

[0044] In one embodiment, the order basic information includes delivery address information, delivery time window information (such as required delivery time), order priority information, consignee contact information, and special delivery instruction information.

[0045] Preferably, the multi-modal drug feature extraction module constructed by the present invention can automatically and accurately identify the drug packaging integrity and storage requirements by fusing visual, text, and storage condition information, dynamically adjust the delivery priority in combination with the timeliness weight, effectively solve the problems of drug quality and safety control and differentiated delivery requirements in pharmaceutical logistics, significantly reduce the risk of manual judgment errors, and improve the delivery efficiency.

[0046] S2: Establish an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input, takes road condition sensing data, temperature monitoring data, and delivery area population density data as environmental feature inputs, dynamically divides the delivery area into grids, and generates a multi-level delivery area model.

[0047] It should be noted that based on the particularity of pharmaceutical logistics, this solution focuses on addressing the following key requirements: First, a dynamic temperature field modeling network is constructed. The temperature threshold constraints for drug storage and transportation are transformed into limiting conditions for temperature monitoring points. Based on the temperature fluctuation tolerance of cold-chain drugs, constrained temperature interpolation calculations are performed to achieve precise monitoring and prediction of the temperature field in the distribution area. Second, the deformation direction of the grid is determined by solving the variational equation for minimizing the distribution cost. Multiple influencing factors such as the regularity constraint of the grid shape, the goal of minimizing the distribution cost, and the environmental feature field are unified into a mathematical framework, breaking through the limitations of traditional static grid division and realizing the overall coordinated deformation and adaptive optimization of the grid.

[0048] Specifically, the flowchart for establishing the adaptive spatio-temporal allocation network is as Figure 3 shown, including the following steps:

[0049] S2.1: Construct a spatio-temporal constraint analysis module to analyze the order constraint matrix and form a spatio-temporal constraint intensity distribution map of the distribution area.

[0050] Specifically, through feature mapping, the time requirements and spatial coordinates in the order constraint matrix are converted into spatio-temporal basic features; based on the spatio-temporal basic features, numerical mappings of the temperature sensitivity, light sensitivity, and humidity sensitivity in combination with the drug storage conditions are performed, and the combined storage requirements of different types of drugs are integrated to generate an order constraint feature matrix; the distribution area is divided into grids according to the order constraint feature matrix, the cumulative constraint intensity of each grid cell is calculated to form a constraint intensity matrix; taking the constraint intensity matrix as the input, a multi-scale constraint intensity tensor is generated by introducing a drug timeliness decay factor in the feature extraction network; spatial smoothing processing based on the order density and the distribution of regional medical resources is performed on the multi-scale constraint intensity tensor, and the spatio-temporal constraint intensity distribution map is output.

[0051] S2.2: Design a multi-source road condition perception network to collect road condition sensing data through a distributed sensor array and establish a road network state feature in combination with the spatio-temporal distribution of medical institutions.

[0052] S2.3: Construct a dynamic temperature field modeling network, use a temperature interpolation network based on key monitoring points to model the regional temperature distribution, incorporate the temperature threshold constraints for drug storage and transportation, and generate a temperature field state feature.

[0053] First, collect distributed temperature sensing data within the distribution area, construct a regional temperature monitoring map based on the sensor location information, and initialize the nodes in the regional temperature monitoring map to the temperature values at the corresponding locations. Secondly, establish a temperature interpolation model based on distance weights, and construct a temperature distribution estimation matrix in combination with the geographical environment characteristics. Subsequently, based on the temperature fluctuation tolerance of cold-chain drugs, convert the drug storage and transportation temperature threshold constraints into the limiting conditions of temperature monitoring points, perform temperature interpolation calculation with constraints on the monitoring area, and generate the regional temperature distribution characteristics. Furthermore, based on the regional temperature distribution characteristics, combine the real-time temperature monitoring data and the drug temperature sensitivity to generate the temperature field state characteristics representing the dynamic changes of the regional temperature.

[0054] S2.4: Design a population activity predictor, establish a regional activity prediction model based on the population density data in the distribution area, and output the population activity state characteristics.

[0055] S2.5: Construct a constraint-driven adaptive feature extractor, perform non-linear feature transformation on the road network state characteristics, temperature field state characteristics and population activity state characteristics to generate a time-varying environmental feature field.

[0056] Among them, the non-linear feature transformation adopts a multi-layer perceptron structure, maps multi-source features to a unified feature space through a learnable weight matrix, and uses the ReLU activation function to introduce non-linear transformation to achieve the adaptive fusion of features.

[0057] S2.6: Superimpose the constraint intensity distribution map and the time-varying environmental feature field, determine the grid deformation direction by solving the variational equation that minimizes the distribution cost, and obtain the adaptive grid division result.

[0058] Specifically, when determining the grid deformation direction, first establish a distribution cost function with the safety of drug distribution as the core. This function includes: a temperature field feature weight term based on the drug temperature sensitivity, a road condition state weight term based on the drug timeliness, a population activity feature weight term based on the order density, and incorporates the urgency of the drug distribution time window. Then, achieve the adaptive fusion of the constraint intensity distribution map and the environmental feature field through dynamic weight adjustment, construct a variational functional based on the grid node displacement, and take the grid shape regularity and the minimum distribution cost as the optimization objectives. Then, solve the Euler equation of the variational functional to obtain the displacement field of the grid nodes. Finally, determine the grid deformation direction according to the coupling relationship between the constraint intensity gradient and the environmental feature field, and adjust the grid deformation rate according to the drug priority to generate the adaptive grid division result.

[0059] It should be noted that the superposition process adopts a differential fusion method based on the characteristics of drugs. For temperature-sensitive drugs, the characteristics of the temperature field are mainly considered; for high-timeliness drugs, the road conditions are mainly concerned; and in areas with concentrated order density, the characteristics of population activities need to be combined, so as to adapt to the dynamic changes of the environment while ensuring the safety of drug distribution. Therefore, the superposition process needs to dynamically adjust the weights of various environmental characteristics according to the characteristics of different drugs.

[0060] S2.7: Perform hierarchical processing on the adaptive grid division results, and perform hierarchical merging according to the regional adjacency relationship and environmental feature similarity, and output a multi-level distribution area model.

[0061] Through the above design, this solution has significant advantages in the pharmaceutical logistics scenario: in terms of ensuring drug safety, it significantly reduces the temperature deviation rate and effectively improves the drug integrity rate; in terms of distribution efficiency, it greatly shortens the response time for emergency orders and significantly improves the distribution on-time rate; in terms of resource utilization, the utilization rate of distribution vehicles is significantly improved and the operating cost is effectively reduced. This solution realizes the overall consideration of safety, timeliness and economy in pharmaceutical logistics and provides reliable technical support for intelligent pharmaceutical logistics.

[0062] S3: Construct a two-layer heterogeneous graph neural network, map the order nodes and distribution area nodes to the graph data structure based on the multi-level distribution area model, perform graph convolution operations to extract the correlation features between nodes, and set the drug timeliness weight and storage condition weight in the graph convolution layer to generate a candidate set of distribution path plans considering drug characteristics.

[0063] It should be noted that based on the timeliness requirements of urgently needed drugs and the safety requirements of drugs with special storage and transportation conditions in pharmaceutical logistics, the drug timeliness weight and storage condition weight are introduced in the graph convolution layer to realize the intelligent trade-off of drug distribution timeliness and safety.

[0064] Specifically, it includes the following steps:

[0065] S3.1: Decompose the order constraint matrix, extract key constraint parameters, and construct a multi-dimensional feature vector of order nodes.

[0066] In one embodiment, the key constraint parameters include but are not limited to drug timeliness level, storage temperature range, light protection requirements, etc.

[0067] S3.2: Analyze the multi-level distribution area model, obtain the spatio-temporal features and environmental attributes of the area nodes, and construct a multi-dimensional feature vector of the area nodes.

[0068] S3.3: Design a graph structure generator that perceives drug constraints, calculate the node connection strength according to the constraint matching degree between the order nodes and the area nodes, and establish a two-layer heterogeneous graph with drug distribution constraints.

[0069] Specifically, based on the multi-dimensional feature vectors of order nodes and area nodes, an order node set and an area node set are constructed respectively; the storage temperature range of the drug is extracted from the order nodes The drug shelf-life level and the light protection requirement level; the environmental temperature range is extracted from the area nodes The regional distribution shelf-life capacity and the light protection capacity; calculate the temperature matching degree M T , and judge the overlapping degree between the drug temperature range of the order node and the environmental temperature range of the area node. The specific formula is as follows:

[0070]

[0071] Among them, δ is a small constant to prevent division by zero, are the upper and lower limits of the drug storage temperature range, are the upper and lower limits of the environmental temperature range. When the intervals completely overlap, M T = 1, when there is no overlap, M T = 0, and the value range of M T is [0, 1].

[0072] Furthermore, calculate the shelf-life matching degree M E , and evaluate whether the regional distribution shelf-life capacity meets the drug shelf-life level requirements. The specific formula is as follows:

[0073]

[0074] Among them, and are both pre-quantified grade parameters (integers), representing the regional distribution shelf-life capacity and the drug shelf-life level requirements respectively. The larger the value, the higher the requirement / capacity.

[0075] Furthermore, calculate the light protection matching degree M L , and compare the regional light protection capacity with the drug light protection requirement level. The specific formula is as follows:

[0076]

[0077] Among them, and are both pre-quantified grade parameters (integers), representing the regional light protection capacity and the drug light protection requirement level respectively. The larger the value, the higher the requirement / capacity.

[0078] Even further, based on the temperature matching degree, the shelf-life matching degree, and the light protection matching degree, calculate the comprehensive constraint matching degree. The specific formula is as follows:

[0079] M comp = ω T·M T +ω E ·M E +ω L ·M L

[0080] Among them, M T , M E , M L are the temperature matching degree, aging matching degree, and light protection matching degree respectively, and ω T , ω E , ω L are the corresponding weight coefficients, and satisfy ω T +ω E +ω L = 1.

[0081] Furthermore, set the connection strength threshold, and map the comprehensive constraint matching degree to the node connection strength. The specific formula is as follows:

[0082]

[0083] Among them, θ is the connection strength threshold, M comp is the comprehensive constraint matching degree, and the value range of S conn is [0, 1].

[0084] Furthermore, according to the calculated node connection strength S conn , establish a weighted edge from the order node to the regional node, where the edge weight is equal to the connection strength S conn , so as to form a two-layer heterogeneous graph with drug delivery constraint information.

[0085] It should be noted that choosing the storage temperature range, aging level, light protection requirement level, environmental temperature range, regional delivery aging ability, and light protection ability as the core parameters for calculating the constraint matching degree is based on the particularity and safety requirements of pharmaceutical logistics. These parameters correspond one by one to form a matching evaluation in three dimensions: temperature, aging, and light protection. It not only covers the key constraint requirements of drug delivery, but also facilitates quantitative calculation and actual operation, and can effectively support the accurate matching and scheduling optimization between orders and regions.

[0086] S3.4: Construct an edge weight calculation module with time - effectiveness self - adaptation, and convert the drug delivery time - effectiveness requirement into an edge weight coefficient.

[0087] Among them, a greater weight is given to the delivery path with high time - effectiveness requirements.

[0088] S3.5: Construct a storage condition constraint factor calculation module, generate a constraint factor matrix according to the drug storage condition requirements, and regulate the feature transfer in graph convolution operations.

[0089] Specifically, a storage condition feature matrix is generated based on the drug storage condition rule set; an environmental adaptability matrix is constructed by calculating the matching degree of the distribution area environmental parameters; a storage risk estimator is set to generate a risk coefficient matrix; the environmental adaptability matrix and the risk coefficient matrix are combined to generate the feature transfer intensity between nodes; the feature transfer intensity is normalized to form a constraint factor matrix; the constraint factor matrix is integrated into the graph convolutional layer to achieve selective feature transfer; the constraint factor matrix is dynamically adjusted through a constraint feedback mechanism.

[0090] S3.6: Adopt a graph attention mechanism to set different weights for the feature aggregation channels based on the constraint factor matrix and update the node representations.

[0091] Specifically, a channel attention calculation module is constructed to generate a channel weight matrix; the node feature vectors are divided into multiple constraint-related sub-feature groups; a constraint adaptive attention calculation unit is used to generate the feature transfer weights between nodes; the weighted combination of features is achieved through a hierarchical feature aggregation network and cross-channel feature fusion; the node representations are optimized by using residual connections and layer normalization.

[0092] S3.7: Design a path evaluation module based on the node representations, comprehensively consider the timeliness matching degree, storage condition satisfaction degree, and path feasibility, and generate a set of candidate solutions for the distribution path planning.

[0093] S4: Set up a reinforcement learning scheduling system, use the temperature field state features and road network state features as the state space, use the set of candidate solutions for the distribution path planning as the action space, construct a two-objective reward function, and iteratively train through the policy gradient algorithm to obtain a distribution scheduling policy model.

[0094] Specifically, it includes the following steps:

[0095] S4.1: Integrate the temperature field state features and the road network state features to generate a distribution environment state matrix.

[0096] S4.2: Map the set of candidate solutions for the distribution path planning to action feature vectors, and construct an action description matrix including path attributes and drug distribution requirements.

[0097] S4.3: Design a two-objective evaluation network architecture based on the distribution environment state matrix and the action description matrix, and the evaluation results of the two networks jointly guide the policy network to generate distribution path planning decisions.

[0098] Among them, the two-objective evaluation network includes a quality value network for evaluating the temperature control quality of drugs and an efficiency value network for evaluating the distribution timeliness. Both the quality value network and the efficiency value network receive the distribution environment state matrix and the action description matrix as inputs, and achieve the coupled evaluation of the environmental state and the path planning through a feature interaction mechanism.

[0099] Specifically, the feature extraction process of the dual-objective evaluation network includes: extracting environmental state features from the distribution environment state matrix through a convolutional neural network, and extracting path planning features from the action description matrix through a graph convolutional network. Based on the extracted features, the quality value network and the efficiency value network respectively perform the following evaluation processes: The quality value network generates drug quality impact features by designing an environment-path interaction module to perform attention-weighted fusion of the environmental state features and the path planning features. This feature characterizes the sequence of environmental impacts that the drug may suffer during transportation along the planned path. Then, a quality shock evaluation module is constructed to process the drug quality impact features. The quality shock evaluation module includes a differential calculation unit, a spatio-temporal mapping unit, and an accumulation effect unit. The differential calculation unit detects the mutation degree of environmental parameters and outputs the environmental parameter mutation value. The spatio-temporal mapping unit evaluates the environmental exposure duration of the drug in different path segments and outputs the environmental exposure duration value. The accumulation effect unit calculates the long-term impact of environmental fluctuations based on the environmental exposure duration and outputs the environmental fluctuation impact value. Finally, the environmental parameter mutation value of the differential calculation unit, the environmental exposure duration value of the spatio-temporal mapping unit, and the environmental fluctuation impact value of the accumulation effect unit are weighted based on the quality sensitivity of the drug type to obtain the quality evaluation score. This design avoids the limitation of traditional methods that only consider the environmental state at a single moment. Through the fine characterization of drug quality impact features, it realizes the dynamic evaluation and early warning of drug quality risks during the entire transportation process, and is particularly suitable for the cold chain transportation quality assurance of biologics and vaccines with high temperature sensitivity.

[0100] Furthermore, the efficiency value network generates a dynamic traffic condition representation by designing an environment-path fusion layer to combine the environmental state features and the path planning features. This representation contains the traffic state of the road network topology under the current environmental conditions. Then, the timeliness score and the reliability score are calculated respectively through a dual-branch evaluation structure. The timeliness score is determined based on the drug delivery time limit level and the dynamic traffic conditions, and the reliability score is determined based on the historical delivery data and the current environmental state. Finally, the scoring results of the two branches are weighted and fused to obtain the efficiency evaluation score. This design overcomes the problem of insufficient response of existing distribution systems to environmental changes. Through the dynamic traffic condition representation, it realizes the accurate characterization of actual traffic conditions and provides a reliable path planning basis for orders with high timeliness requirements such as emergency drugs.

[0101] It should be noted that the quality value network focuses on drug quality impact features, emphasizing the impact of the environment on the drug itself. The efficiency value network focuses on the dynamic traffic condition representation, emphasizing the impact of the environment on the traffic state of the road network.

[0102] In addition, the present invention designs an evaluation guidance mechanism to convert the evaluation results of the dual-network into the basis for action selection of the policy network. The execution process of the evaluation guidance mechanism includes: First, normalize the quality evaluation score and the efficiency evaluation score; then, set action selection constraint conditions according to the evaluation scores; finally, input the constraint conditions into the policy network to guide the generation of path planning decisions. This mechanism realizes the precise guidance of the evaluation results for path planning. Through the dynamic adjustment of multi-dimensional constraint conditions, it ensures that the generated distribution plan can meet the requirements of both drug quality maintenance and distribution efficiency, effectively solving the problem that it is difficult for the evaluation results in traditional methods to effectively guide the generation of decisions.

[0103] Preferably, the above-mentioned dual-objective evaluation network architecture realizes the deep coupling of the environmental state and path planning through the feature interaction mechanism, overcoming the problem of the disconnection between environmental evaluation and path evaluation in traditional methods. The quality value network accurately evaluates the drug quality risk through path-environment interaction features, and the efficiency value network accurately predicts the distribution efficiency through environment-path fusion. The evaluation guidance mechanism ensures that the evaluation results can effectively guide path planning decisions. This design comprehensively considers the interactive effects of environmental state and path selection in pharmaceutical cold chain logistics, improving the reliability and efficiency of the distribution plan.

[0104] S4.4: Design a dual-objective reward function based on drug characteristics, including a quality maintenance reward function based on the evaluation score of the quality value network and a timeliness performance reward function based on the evaluation score of the efficiency value network.

[0105] In pharmaceutical cold chain logistics, drug distribution faces two core challenges: on the one hand, drugs are highly sensitive to environmental conditions such as temperature, and temperature fluctuations can lead to a decline in drug efficacy or even invalidation, so it is necessary to strictly control the distribution environment throughout the process; on the other hand, some drugs such as vaccines and emergency drugs have strict distribution timeliness requirements, and delays may endanger the lives of patients. Based on these two special needs, the present invention designs an innovative dual-objective reward function: in terms of quality maintenance, evaluate the cumulative effect of temperature fluctuations through the temperature deviation integral term and consider the influence of environmental factors such as light; in terms of distribution timeliness, construct a penalty term using the ratio of the actual distribution time to the required time to adapt to the differentiated timeliness requirements of different distribution tasks.

[0106] Specifically, the specific formula of the dual-objective reward function is as follows:

[0107] R = λ q ·R q + λ t ·R t

[0108] Wherein, R q is the quality maintenance reward function, R tis the aging performance reward function, λ q , λ t are the weight coefficients of the quality retention reward function and the aging performance reward function, satisfying λ q +λ t = 1.

[0109] The calculation formula of the quality retention reward function R q is as follows:

[0110] R q = Q v -η1·∫|D(T t )|dt - η2·L

[0111] where Q v is the quality evaluation score output by the quality value network, T t is the actual temperature at time t during transportation, L is the cumulative value of light exposure, η1 and η2 are the penalty coefficients for temperature deviation integration and light exposure, and D(T t ) is the temperature deviation function, and the specific formula is as follows:

[0112]

[0113] where T min and T max are respectively the lower and upper limits of the specified storage temperature of the drug.

[0114] In addition, the calculation formula of the aging performance reward function R t is as follows:

[0115]

[0116] where E v is the efficiency evaluation score output by the efficiency value network, t actual is the actual delivery time, t required is the required delivery time, Δt is the delivery delay time, and γ1 and γ2 are the penalty coefficients for the time ratio and the delay time.

[0117] S4.5: Construct a policy gradient algorithm framework, iteratively optimize the parameters of the quality value network and the efficiency value network based on the dual-objective reward function, and guide the update of the policy network parameters to train a delivery scheduling policy model.

[0118] S5: Install an edge computing unit on the delivery vehicle to collect delivery environment data, use the federated learning framework to aggregate multi-region delivery experience data, and perform periodic updates on the scheduling policy model based on the differential privacy mechanism.

[0119] Specifically, a pharmaceutical quality monitoring module is deployed in the edge computing unit of the distribution vehicle to collect the set of drug environmental states; an edge training module oriented to drug quality is constructed, which combines the set of drug environmental states with the distribution scheduling strategy model and performs local training based on the drug quality maintenance index to generate the local model parameter update amount; a federated learning aggregation architecture based on drug attributes is designed to divide the distributed drugs into sub-categories according to their attribute characteristics, aggregate the local model parameter update amounts of the same category to generate the category model gradient; a pharmaceutical data privacy protection module is constructed to add noise processing to the category model gradient and set the gradient clipping threshold; a model fusion mechanism constrained by drug quality is designed to perform weighted aggregation on the model gradients of multiple drug categories to generate the global model update amount; a drug distribution efficiency evaluation module is established to evaluate the updated distribution scheduling strategy model based on drug quality indicators and distribution efficiency indicators, and dynamically adjust the category weights; a cycle update mechanism oriented to drug quality is set to apply the global model update amount to the distribution scheduling strategy model, complete the model parameter update and distribution.

[0120] Furthermore, this embodiment also provides an efficient pharmaceutical logistics order scheduling and distribution system, including a feature extraction module for constructing a multi-modal drug feature extraction module. The multi-modal drug feature extraction module generates a drug feature vector by fusing multi-modal drug information, identifies the drug storage requirements and distribution restriction conditions based on the drug feature vector, and outputs an order constraint matrix; a region allocation module for establishing an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input and the environmental features as the auxiliary input, dynamically divides the distribution area into grids, and generates a multi-level distribution area model; a path planning module for constructing a two-layer heterogeneous graph neural network, mapping the order nodes and distribution area nodes to the graph data structure based on the multi-level distribution area model, performing graph convolution operations to extract the correlation features between nodes, and setting the drug timeliness weight and storage condition weight in the graph convolution layer to generate a set of candidate distribution path planning schemes considering drug characteristics; a scheduling strategy module for setting up a reinforcement learning scheduling system, using the temperature field state feature and road network state feature as the state space, and the set of candidate distribution path planning schemes as the action space, constructing a reward function based on the drug quality maintenance index and distribution timeliness rate, and iteratively training through the policy gradient algorithm to obtain the distribution scheduling strategy model; a model update module for collecting distribution environment data on the distribution vehicle equipped with an edge computing unit, aggregating the distribution experience data of multiple regions using the federated learning framework, and performing periodic updates on the scheduling strategy model based on the differential privacy mechanism.

[0121] In summary, by constructing a multi-modal drug feature extraction module and integrating the image, text, and storage condition information of drugs, the present invention can more accurately identify the drug storage requirements and distribution restriction conditions, improving the recognition accuracy of distribution constraints. The designed adaptive spatio-temporal allocation network effectively enhances the flexibility of distribution resource allocation through a dynamic grid division method. The double-layer heterogeneous graph neural network is used for order-region matching, and the drug timeliness weight and storage condition weight are introduced in the graph convolutional layer, improving the rationality of path planning. The reinforcement learning scheduling system based on the double-objective evaluation network achieves the balanced optimization of drug quality maintenance and distribution timeliness, enabling the distribution plan to better meet the special needs of pharmaceutical logistics.

[0122] Example 2, referring to Figures 1 to 3 , which is the second embodiment of the present invention. This embodiment provides an efficient pharmaceutical logistics order scheduling and distribution method. To verify the beneficial effects of the present invention, scientific demonstration is carried out through economic benefit calculation and simulation experiments.

[0123] To verify the effectiveness of the efficient pharmaceutical logistics order scheduling and distribution method proposed by the present invention, a 6-month simulation experiment was carried out in a provincial pharmaceutical logistics distribution center. The distribution center covers an area of about 3,000 square kilometers, with an average daily order volume of about 2,000 orders. The types of drugs distributed include normal-temperature drugs, cold-chain drugs, and special control drugs, etc. The experimental environment adopts a distributed edge computing architecture, equipped with a temperature sensor network and a road condition monitoring system. Edge computing units are deployed on 200 distribution vehicles for data collection and processing. First, a multi-modal drug feature extraction module is constructed, collecting 50,000 drug package images and corresponding instruction manual texts. ResNet-50 is used as the backbone network to construct a hierarchical attention visual network for feature extraction, with the batch size set to 32 and the learning rate to 0.001. After training, in the order constraint matrix output by this module, the drug feature extraction accuracy reaches 98.2%, and the packaging integrity recognition accuracy is 99.1%.

[0124] Based on the generated order constraint matrix, an adaptive spatio-temporal allocation network is further constructed. This network utilizes the real-time data of 2,000 temperature monitoring points and 500 road condition monitoring points arranged in the distribution area. The interpolation accuracy of the temperature field dynamic modeling network is set to 0.1°C, and the initial grid size is 500 m × 500 m. Combining the drug storage requirements in the order constraint matrix, grid adaptive deformation is achieved through variational equation solving, and the minimum grid size can reach 100 m × 100 m. The experimental results show that this network controls the temperature deviation rate within ±0.3°C, and the adaptability index of grid division in the generated multi-level distribution area model is improved by 42.5% compared with the traditional fixed grid.

[0125] On this basis, a two-layer heterogeneous graph neural network is constructed to process order scheduling. This network is based on a multi-level distribution area model, maps order nodes and distribution area nodes to a graph structure, the dimension of the node feature vector is 256, and an 8-layer graph convolution structure is adopted. The network is trained using the 6-month distribution data (about 1 million records) accumulated by the first two modules, trained for 200 rounds through the Adam optimizer, and the initial learning rate is 0.0001. In the experiment, based on the order constraint matrix, the temperature matching degree threshold is set to 0.9 and the timeliness matching degree threshold is set to 0.85. The satisfaction degree of the generated distribution path planning candidate solution set for the drug storage conditions reaches 96.8%.

[0126] Subsequently, a reinforcement learning scheduling system is constructed, with the temperature field state feature and the road network state feature as the state space, and the generated distribution path planning candidate solution set as the action space. The system adopts a dual-objective evaluation network architecture. Both the quality value network and the efficiency value network adopt a 3-layer fully connected structure, with the quality maintenance weight coefficient of 0.6 and the timeliness performance weight coefficient of 0.4. After 1000 rounds of iterative training of the policy gradient algorithm, the system realizes a temperature qualification rate of 99.5% for cold chain drugs based on the features and constraint conditions provided by the previous module. In terms of order response speed, the scheduling time interval for ordinary orders is 20 - 45 seconds, and the average scheduling time is 30 seconds; the scheduling time interval for emergency orders is 5 - 15 seconds, and the average scheduling time is 10 seconds. In terms of distribution timeliness, the distribution time interval for ordinary orders is 30 - 60 minutes, and the average distribution time is 45 minutes; the distribution time interval for emergency orders is 10 - 20 minutes, and the average distribution time is 15 minutes. Among them, the interval difference in distribution time is mainly affected by factors such as distribution distance, road conditions, and time periods.

[0127] Finally, the trained scheduling policy model is deployed on the edge computing unit of the distribution vehicle, and the federated learning framework is used to continuously optimize the model performance. The actual operation data for 6 months shows that the method of the present invention achieves high distribution efficiency and drug safety guarantee. As shown in Table 1 below, the method of the present invention is compared and analyzed with the prior art:

[0128] Table 1 Performance comparison between the present invention and the prior art

[0129] Evaluation index Traditional method Existing intelligent scheduling method Method of the present invention Temperature deviation rate ±1.2℃ ±0.8℃ ±0.3℃ Drug integrity rate 92.5% 95.8% 99.5% Distribution accuracy rate 85.6% 91.3% 97.2% Vehicle utilization rate 65.2% 75.8% 85.3% Reduction rate of operating cost - 15.2% 28.6%

[0130] It can be seen from the comparison in Table 1 that the method proposed by the present invention has achieved remarkable results in aspects such as drug safety guarantee, distribution efficiency improvement, and resource utilization through the collaborative optimization of multi-modal feature extraction, adaptive grid division, graph neural network path planning, and reinforcement learning scheduling. Compared with the existing methods, the present invention has achieved a significant improvement in various key indicators, fully verifying the application value of this method in the actual pharmaceutical logistics scenario.

[0131] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. An efficient pharmaceutical logistics order scheduling and distribution method, characterized in that: including Construct a multimodal drug feature extraction module. The multimodal drug feature extraction module generates a drug feature vector by fusing multimodal drug information, identifies drug storage requirements and distribution restriction conditions based on the drug feature vector, and outputs an order constraint matrix; Establish an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input and environmental features as the auxiliary input, dynamically divides the distribution area into grids, and generates a multi-level distribution area model; Construct a double-layer heterogeneous graph neural network. Based on the multi-level distribution area model, map the order nodes and distribution area nodes to the graph data structure, perform graph convolution operations to extract the correlation features between nodes, and set the drug timeliness weight and storage condition weight in the graph convolution layer to generate a set of candidate solutions for the distribution path planning considering drug characteristics; Set up a reinforcement learning scheduling system. Use the temperature field state feature and road network state feature as the state space, and the set of candidate solutions for the distribution path planning as the action space. Construct a double-objective reward function, and iteratively train through the policy gradient algorithm to obtain a distribution scheduling policy model; Install an edge computing unit on the distribution vehicle to collect distribution environment data, use the federated learning framework to aggregate multi-region distribution experience data, and periodically update the scheduling policy model based on the differential privacy mechanism.

2. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 1, wherein: The multimodal drug information includes drug image information, instruction text information, and storage condition information; The construction of the multimodal drug feature extraction module includes the following steps: Collect drug image data, and extract the drug appearance features through the hierarchical attention visual network of the drug packaging recognizer to generate a drug image feature matrix; Perform word segmentation on the drug instruction text, and use the domain-adaptive BERT language model integrating the drug knowledge graph to extract the key attribute information of the drug to generate a drug text feature matrix; Establish a drug storage condition quantization module, perform linear interval mapping on the drug storage parameters to generate a drug storage feature matrix; Construct a multimodal feature fusion network, input the drug image feature matrix, the drug text feature matrix, and the drug storage feature matrix into the feature fusion layer, integrate different modal features through the feature weighted fusion mechanism, and output a drug feature vector; Design a bidirectional attention network, use the drug feature vector as the query input, use the preset distribution condition library as the key-value input, adopt a multi-head attention structure to extract the drug distribution requirements, and set up an independent timeliness evaluation module to generate drug distribution restriction features; Establish a constraint matrix generation module, combine the drug distribution restriction features with the order basic information, and form an order constraint matrix through the gradient constraint propagation algorithm.

3. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 1, characterized in that: The environmental features include road condition sensing data, temperature monitoring data, and population density data of the distribution area; Establishing an adaptive spatio-temporal allocation network according to the order constraint matrix includes the following steps: Construct a spatio-temporal constraint analysis module to analyze the order constraint matrix and form a spatio-temporal constraint intensity distribution map of the distribution area; Design a multi-source road condition perception network, collect road condition sensing data through a distributed sensor array, and establish a road network state feature in combination with the spatio-temporal distribution of medical institutions; Construct a dynamic modeling network for the temperature field. Use a temperature interpolation network based on key monitoring points to model the regional temperature distribution, incorporate the temperature threshold constraints for drug storage and transportation, and generate the temperature field state features. Design a population activity predictor. Based on the population density data of the distribution area, establish a regional activity prediction model and output the population activity state features. Construct a constraint-driven adaptive feature extractor. Perform non-linear feature transformation on the road network state features, temperature field state features, and population activity state features to generate a time-varying environmental feature field. Overlay the constraint intensity distribution map with the time-varying environmental feature field. Determine the grid deformation direction by solving the variational equation for minimizing the distribution cost, and obtain the adaptive grid division result. Perform hierarchical processing on the adaptive grid division result, and perform hierarchical merging according to the regional adjacency relationship and environmental feature similarity, and output a multi-level distribution area model.

4. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 1, characterized in that: The construction of the double-layer heterogeneous graph neural network includes the following steps: Decompose the order constraint matrix, extract the key constraint parameters, and construct a multi-dimensional feature vector of the order nodes. Analyze the multi-level distribution area model, obtain the spatio-temporal features and environmental attributes of the area nodes, and construct a multi-dimensional feature vector of the area nodes. Design a graph structure generator that perceives drug constraints. Calculate the node connection strength according to the constraint matching degree between the order nodes and the area nodes, and establish a double-layer heterogeneous graph with drug distribution constraints. Construct a time-sensitive adaptive edge weight calculation module, and convert the drug distribution time requirement into an edge weight coefficient. Construct a storage condition constraint factor calculation module, generate a constraint factor matrix according to the drug storage condition requirements, and regulate the feature transfer in the graph convolution operation. Adopt a graph attention mechanism, set different weights for the feature aggregation channels based on the constraint factor matrix, and update the node representations. Design a path evaluation module based on node representations. Considering the timeliness matching degree, storage condition satisfaction degree, and path feasibility, generate a candidate set of distribution path planning schemes.

5. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 1, characterized in that: The setting of the reinforcement learning scheduling system includes the following steps: Integrate the temperature field state features and the road network state features to generate a distribution environment state matrix. Map the candidate set of distribution path planning schemes to an action feature vector, and construct an action description matrix that includes path attributes and drug distribution requirements. Design a dual-objective evaluation network architecture based on the distribution environment state matrix and the action description matrix, including a quality value network for evaluating the drug temperature control quality and an efficiency value network for evaluating the distribution timeliness. The evaluation results of the two networks jointly guide the policy network to generate distribution path planning decisions. Design a dual-objective reward function based on drug characteristics, including a quality maintenance reward function based on temperature deviation and light exposure, and a timeliness performance reward function based on the distribution on-time rate. Construct a policy gradient algorithm framework. Based on the dual-objective reward function, iteratively optimize the parameters of the quality value network and the efficiency value network, and guide the update of the policy network parameters to train a distribution scheduling policy model.

6. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 5, characterized in that: The evaluation process of the quality value network is as follows: Extract the environmental state features from the distribution environment state matrix through a convolutional neural network, and extract the path planning features from the action description matrix through a graph convolutional network. By designing an environment-path interaction module, the environmental state features and path planning features are fused with attention weighting to generate drug quality impact features; A quality impact assessment module is constructed to process the drug quality impact features. The quality impact assessment module includes a differential calculation unit, a spatio-temporal mapping unit, and an accumulation effect unit. The differential calculation unit detects the mutation degree of environmental parameters. The spatio-temporal mapping unit evaluates the environmental exposure duration of drugs in different path segments. The accumulation effect unit calculates the long-term impact of environmental fluctuations based on the environmental exposure duration; The output result of the quality impact assessment module is weighted based on the quality sensitivity of the drug type to obtain a quality assessment score.

7. The efficient pharmaceutical logistics order scheduling and distribution method according to claim 5, characterized in that: The specific formula of the dual-objective reward function is as follows: R = λ q ·R q + λ t ·R t Among them, λ q , λ t are the weight coefficients of the quality retention reward function and the timeliness performance reward function, R q is the quality retention reward function, R t is the timeliness performance reward function. R q = Q v - η1·∫|D(T t )|dt - η2·L Among them, Q v is the quality evaluation score output by the quality value network, T t is the actual temperature at time t during transportation, L is the cumulative value of light exposure, η1 and η2 are the penalty coefficients for temperature deviation integration and light exposure, D(T t ) is the temperature deviation function, E v is the efficiency evaluation score output by the efficiency value network, t actual is the actual delivery time, t required is the required delivery time, Δt is the delivery delay time, γ1 and γ2 are the penalty coefficients for time ratio and delay time.

8. An efficient pharmaceutical logistics order scheduling and distribution system, based on the efficient pharmaceutical logistics order scheduling and distribution method according to any one of claims 1 to 7, characterized in that: It also includes, A feature extraction module for constructing a multi-modal drug feature extraction module. The multi-modal drug feature extraction module generates a drug feature vector by fusing multi-modal drug information, identifies drug storage requirements and distribution restriction conditions based on the drug feature vector, and outputs an order constraint matrix; A region allocation module for establishing an adaptive spatio-temporal allocation network according to the order constraint matrix. The adaptive spatio-temporal allocation network takes the order constraint matrix as the main input and environmental features as the auxiliary input, dynamically divides the distribution area into grids, and generates a multi-level distribution area model; A path planning module for constructing a two-layer heterogeneous graph neural network. Based on the multi-level distribution area model, the order nodes and distribution area nodes are mapped to the graph data structure, graph convolution operations are performed to extract the correlation features between nodes, and the drug timeliness weight and storage condition weight are set in the graph convolution layer to generate a set of candidate distribution path planning schemes considering drug characteristics; A scheduling strategy module for setting up a reinforcement learning scheduling system. Using the temperature field state feature and road network state feature as the state space, and the set of candidate distribution path planning schemes as the action space, a reward function is constructed based on the drug quality maintenance index and distribution timeliness rate, and a distribution scheduling strategy model is obtained through iterative training by the policy gradient algorithm; A model update module for collecting distribution environment data when the distribution vehicle is equipped with an edge computing unit, aggregating multi-region distribution experience data using the federated learning framework, and periodically updating the scheduling strategy model based on the differential privacy mechanism.

Citation Information

Cited By

  • Risk identification method of medicine order and related device

    CN120893845A

  • Medicine supply chain scheduling method and system based on reinforcement learning

    CN121617585A

  • A medicine supply chain scheduling method and system based on reinforcement learning

    CN121617585B

  • Multi-modal distribution resource optimal configuration and scheduling method for drug online platform

    CN122114519A

  • Method for optimizing configuration and scheduling of multi-modal distribution resources of online drug platform

    CN122114519B