Order Allocation Method, Model Construction Method and Device under the Shared Manufacturing Model
The DQN algorithm optimizes order allocation in shared manufacturing systems by integrating historical data and simulation, addressing inefficiencies and adaptability issues, resulting in significantly reduced order completion times.
Patent Information
- Application Number
- CN202411079969.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-08-07
AI Technical Summary
The existing hybrid integer programming and heuristic algorithms are difficult to fully restore complex dynamic shared manufacturing systems, which makes order allocation strategies difficult to adapt to actual scenario changes and lacks rapid adaptability.
Build a shared manufacturing simulation model, combine the DQN algorithm to train the order allocation strategy model, and optimize order allocation decisions by obtaining historical data and manufacturer characteristics.
It improves the accuracy and adaptability of order allocation strategies, and can make order allocation decisions efficiently in a complex and dynamic shared manufacturing environment.
Smart Images

Figure CN119398358B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of shared manufacturing, and particularly relates to an order allocation method, a model construction method and a device under a shared manufacturing mode. Background Art
[0002] Shared manufacturing, as an emerging manufacturing mode, provides hierarchical shared manufacturing services in a peer-to-peer manner, supporting sharing consumption between individuals or enterprises. Specifically, the tangible or intangible assets of suppliers are encapsulated into shared manufacturing services and integrated into platform services to support consumers to access in a peer-to-peer manner. In this way, the utility of goods and services can be maximized, and the breadth and depth of sharing can be expanded. According to factors such as platform main body characteristics, business models, and shared content, the shared manufacturing mode can be divided into four types: intermediary type, mass innovation type, service type, and collaborative type. The intermediary type platform is built by a third-party Internet enterprise, which docks the supply and demand sides by integrating the production capacity of multiple parties; the mass innovation type platform is led by large manufacturing enterprises and provides comprehensive services for small and micro start-up enterprises and internal entrepreneurship teams; the service type platform is built by industrial technology-based enterprises, leases intelligent equipment to manufacturing enterprises and provides technical services; the collaborative type platform is jointly built by a third party or small and micro enterprises, and the participating entities are mainly small and medium-sized enterprises.
[0003] The intermediary type shared manufacturing platform is built, operated and managed by a third-party enterprise. This type of platform does not own manufacturing capacity itself and is mainly responsible for matching and dispatching transactions between the supply and demand sides. Therefore, under this mode, the research on the matching and scheduling of the supply and demand sides - the research on the order allocation strategy, has greater difficulty and prominent practical significance.
[0004] At present, a modeling method of mixed integer programming is usually adopted and combined with a heuristic algorithm to solve the supply and demand matching and production scheduling problems. However, the modeling method of mixed integer programming is difficult to fully restore the complex dynamic shared manufacturing system, and thus it is difficult to formulate an allocation strategy that can effectively respond to the actual order allocation scenario. The shared manufacturing system involves multiple main bodies such as the demand side and production enterprises. The modeling of its order allocation problem is inseparable from the description of the demand generation process and the enterprise production process. And various demands have heterogeneity in characteristics such as geographical location and demand quantity, and there are also heterogeneities in elements such as the number of devices and production rate among production enterprises. In addition, there is randomness in the internal production process of a single enterprise and the external demand of the shared manufacturing system. The mathematical programming model is difficult to fully describe the order allocation problem of the shared manufacturing system. In addition, although the heuristic algorithm can solve the problem modeled by the simulation method, it can only solve the order allocation strategy for one shared manufacturing scenario, that is, one parameter combination at a time. Once the scenario changes, it is necessary to re-find the optimal solution, that is, it lacks the ability to quickly adapt to the dynamic manufacturing environment.
[0005] Therefore, there is an urgent need for a new order allocation method to solve the order allocation problem in the shared manufacturing model. Summary of the Invention
[0006] To solve the above problems, the present invention proposes an order allocation method, a model construction method and a device in the shared manufacturing model, which can more comprehensively and accurately reflect various factors and interactions in the shared manufacturing system, and depict the order allocation problem in the shared manufacturing system in a way closer to the actual situation, with high operating efficiency and good performance, and can be applied to complex and dynamic shared manufacturing environments.
[0007] In the first aspect, an order allocation strategy model construction method provided by the present invention includes:
[0008] Obtain historical data, where the historical data at least includes historical order data and production enterprise characteristic data;
[0009] Construct a shared manufacturing simulation model according to the historical data and the production enterprise operation process;
[0010] Obtain an initial order allocation strategy model, where the initial order allocation strategy model is constructed according to the shared manufacturing simulation model and a DQN algorithm model with the shared manufacturing simulation model as the interaction object;
[0011] Train the initial order allocation strategy model based on the DQN (Deep Q-Network) algorithm to obtain a trained order allocation strategy model.
[0012] In an optional embodiment, the shared manufacturing simulation model at least includes a state acquisition module, a revenue acquisition module, an order arrival module, a queue module, an order allocation module and multiple production enterprise modules. Among them, the order arrival module, the queue module and the order allocation module are connected in sequence, and each production enterprise module includes a production sub-module, a production equipment sub-module and a truck sub-module respectively connected to the production sub-module; the order allocation module is respectively connected to the production sub-modules in each production enterprise module;
[0013] The state acquisition module is used to determine state variable information and send the state variable information to the DQN algorithm model;
[0014] The revenue acquisition module is used to determine revenue information and send the revenue information to the DQN algorithm model; where the revenue information is expressed as:
[0015] FT j =WT j +PT j +DT j ;
[0016] In the above formula, FT j is the completion time of the j-th order; WT j is the waiting time of the j-th order; PT j is the production time of the j-th order; DT j is the transportation time of the j-th order;
[0017] The order arrival module is used to randomly generate multiple orders according to demand points and determine the order information corresponding to each order;
[0018] The queue module is used to queue each order, send the orders sequentially to the order allocation module according to the queuing order; and determine the waiting time of each order;
[0019] The order allocation module is used to determine the production enterprise module corresponding to each order according to the order allocation strategy information output by the DQN algorithm model, and sequentially send multiple order information to the corresponding production enterprise modules;
[0020] The production sub-module is used to call the production equipment sub-module to produce products according to the order information;
[0021] The production equipment sub-module is used to determine the production time required for the target production equipment to produce a single product;
[0022] The truck sub-module is used to transport the products to the demand points according to the pre-established GIS map model and determine the transportation time.
[0023] In an alternative embodiment, the historical order data at least includes the location data of the demand point to which each historical order belongs and the product quantity data in each historical order.
[0024] In an alternative embodiment, the order arrival module includes a demand point sub-module, a product quantity sub-module, and an order arrival interval sub-module;
[0025] The demand point sub-module is used to determine the location information of the demand point to which each order belongs according to the location data of the demand point to which each historical order belongs and the location distribution function;
[0026] The product quantity module is used to determine the product quantity information included in each order according to the product quantity data in each historical order and the data distribution function;
[0027] The order arrival interval module is used to determine the time interval at which each order arrives, so as to generate multiple orders according to the time interval at which each order arrives, wherein each generated order includes order information, and the order information includes the location information of the demand point to which each order belongs and the product quantity information included in each order.
[0028] In an alternative embodiment, the production enterprise characteristic data at least includes the location data of each production enterprise, the number of production devices of each production enterprise, and the time data for each device in each production enterprise to produce a single product, the failure time data, and the maintenance time data.
[0029] In an alternative embodiment, the production device sub-module includes a processing capacity module, a mean time between failures (MTBF) module, and a maintenance time module;
[0030] The failure time module is used to determine the failure occurrence time of the target production device according to the failure time data;
[0031] The maintenance time module is used to determine the maintenance time of the target production device according to the maintenance time data;
[0032] The processing capacity module is used to determine the production time used by the target production device to produce a single product according to the time data for producing a single product;
[0033] Among them, the production time of the target production device for producing the single product is determined according to the production time used by the target production device to produce a single product, the maintenance time of the target production device, and the failure occurrence time of the target production device.
[0034] In an alternative embodiment, the state variable information includes:
[0035] s t =(d t ,l t ,c it *k it ,w it );
[0036] In the above formula, s t is the state variable in the t-th round; d t is the number of products to be produced for the order to be allocated in the t-th round; l t is the location of the demand point to which the order to be allocated in the t-th round belongs; c i *k it is the idle production capacity of the i-th production enterprise; where c i is the reciprocal of the average time used to produce a single product; k it is the t-th round, the number of idle devices of the i-th production enterprise; w it is the number of products to be produced by the i-th production enterprise.
[0037] In an alternative embodiment, the initial model of the order allocation strategy is trained based on the DQN algorithm to obtain a trained order allocation strategy model, including:
[0038] Initialize the network parameters of the experience pool and the DQN algorithm model. Among them, the DQN algorithm model includes an action neural network and a target neural network, and the network parameters of the DQN algorithm model include the network parameters of the action neural network and the network parameters of the target neural network;
[0039] Repeat the following steps until the preset number of training times is reached:
[0040] Reset the shared manufacturing simulation model;
[0041] Repeat the following steps until the production of all orders is completed:
[0042] Obtain the current state variable information sent by the shared manufacturing simulation model;
[0043] Determine the order allocation strategy information according to the current state variable information and the ε-greedy policy algorithm, and send the order allocation strategy information to the shared manufacturing simulation model;
[0044] Obtain the revenue information and the next state variable information sent by the shared manufacturing simulation model;
[0045] Store the experience tuple in the experience pool, where the experience tuple includes the current state variable information, the order allocation strategy information, the revenue information, and the next state variable information;
[0046] Randomly select a preset number of experience tuples from the experience pool for training;
[0047] Determine the target value corresponding to each experience tuple according to the action neural network and the target neural network, and determine the loss function value;
[0048] Use the backpropagation algorithm and the gradient descent algorithm to update the network parameters of the action neural network. After updating the preset number of times, copy the value of the network parameters of the action neural network to the target neural network to make the training process stable;
[0049] Output the trained order allocation strategy model.
[0050] In a second aspect, an order allocation method under a shared manufacturing mode provided by the present invention includes:
[0051] Obtain the order to be allocated;
[0052] Input the order to be allocated into the trained order allocation strategy model, and output the allocation strategy of the order to be allocated, where the order allocation strategy model is constructed according to the order allocation strategy model construction method described in the first aspect.
[0053] In a third aspect, the present invention provides an apparatus for constructing an order allocation strategy model, including:
[0054] A data acquisition module, configured to acquire historical data, where the historical data at least includes historical order data and production enterprise characteristic data;
[0055] A simulation model construction module, configured to construct a shared manufacturing simulation model according to the historical data and the production enterprise operation process;
[0056] An initial model acquisition module, configured to acquire an initial order allocation strategy model, where the initial order allocation strategy model is constructed according to the shared manufacturing simulation model and a DQN algorithm model with the shared manufacturing simulation model as an interaction object;
[0057] A training module, configured to train the initial order allocation strategy model based on the DQN algorithm to obtain a trained order allocation strategy model.
[0058] In a fourth aspect, the present invention provides an order allocation apparatus under a shared manufacturing mode, characterized by including:
[0059] An order acquisition module, configured to acquire orders to be allocated;
[0060] An order allocation module, configured to input the orders to be allocated into the trained order allocation strategy model and output an allocation strategy for the orders to be allocated, where the order allocation strategy model is constructed by the order allocation strategy model construction method according to any one of the first aspect.
[0061] In a fifth aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the processor executes the computer program, the steps of the method according to any one of the foregoing embodiments are implemented.
[0062] In a sixth aspect, the present invention provides a computer-readable medium having non-volatile program code executable by a processor, where the program code causes the processor to execute the method according to any one of the foregoing embodiments.
[0063] The beneficial effects brought by the technical solution provided by the embodiments of the present invention are as follows: In the order allocation method, model construction method and device under the shared manufacturing mode of the present invention, in the model construction method, a shared manufacturing simulation model is constructed based on historical data and the operation process of production enterprises, so that various factors and interactions in the shared manufacturing system can be more comprehensively and accurately reflected, and the order allocation problem in the shared manufacturing system can be described in a way closer to the actual situation, solving the limitation that it is difficult for the mathematical programming model to fully restore the complex and dynamic shared manufacturing system, and providing support for solving the order allocation problem; in the model construction method, the initial order allocation strategy model is trained based on the DQN algorithm with the shared manufacturing simulation model as the interaction object, that is, the initial order allocation strategy model based on the DQN algorithm is trained with the shared manufacturing simulation model as the training environment, so that the trained order allocation strategy model can directly make order allocation decisions for the orders in the shared manufacturing system without repeatedly running the simulation model to search for order allocation decisions, thereby improving the operation efficiency; during the training process, the shared manufacturing system in different scenarios can be simulated, thereby improving the adaptability; thus, it can be applied to complex and dynamic shared manufacturing environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 It is a schematic flowchart of the order allocation strategy model construction method provided by the embodiments of the present invention;
[0065] Figure 2 It is a schematic diagram of the principle of the application scenario of the shared manufacturing system provided by the embodiments of the present invention;
[0066] Figure 3 It is a schematic diagram of the interaction principle between the shared manufacturing simulation model and the DQN algorithm model provided by the embodiments of the present invention;
[0067] Figure 4 It is a schematic diagram of the shared manufacturing simulation model provided by the embodiments of the present invention;
[0068] Figure 5 It is a schematic flowchart of the order allocation method under the shared manufacturing mode provided by the embodiments of the present invention;
[0069] Figure 6 It is a schematic diagram for comparing the effects of the order allocation strategy generated by using the order allocation method under the shared manufacturing mode and the nearest-neighbor allocation strategy provided by the embodiments of the present invention;
[0070] Figure 7 It is a sensitivity analysis diagram of different order arrival rates under the order allocation strategy generated by using the order allocation method under the shared manufacturing mode provided by the embodiments of the present invention;
[0071] Figure 8Sensitivity analysis diagram of order arrival following uniform distribution under the order allocation strategy generated by the order allocation method using the shared manufacturing mode provided by the embodiments of the present invention;
[0072] Figure 9 System schematic diagram of the order allocation strategy model construction device provided by the embodiments of the present invention;
[0073] Figure 10 System schematic diagram of the order allocation device under the shared manufacturing mode provided by the embodiments of the present invention;
[0074] Figure 11 System schematic diagram of the electronic device provided by the embodiments of the present invention.
[0075] In the figure: 11 - data acquisition module; 12 - simulation model construction module; 13 - initial model acquisition module; 14 - training module; 21 - order acquisition module; 22 - order allocation module; 400 - electronic device; 401 - communication interface; 402 - processor; 403 - memory; 404 - bus. Detailed implementation manners
[0076] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.
[0077] Refer to Figure 1 An order allocation strategy model construction method provided by the embodiments of the present invention includes steps S110 to S140.
[0078] Step S110, obtain historical data, where the historical data includes historical order data and production enterprise characteristic data; the historical order data at least includes location data and product quantity data. Here, the location data is used to determine the location of the demand point to which each order belongs, and the product quantity data is used to determine the quantity of products included in each order. The production enterprise characteristic data at least includes the location data of each production enterprise, the number of production equipment of each production enterprise, and the time data for each device in each production enterprise to produce a single product, failure time data, and repair time data.
[0079] Step S120, construct a shared manufacturing simulation model according to the historical data and the production enterprise operation process. Specifically, in this embodiment, simulation software is used to construct the shared manufacturing simulation model.
[0080] When constructing the shared manufacturing simulation model, it is necessary to construct it in combination with the application scenario of the shared manufacturing system - the production enterprise operation process under the shared manufacturing system. The application scenario of the shared manufacturing system in this embodiment is as Figure 2As shown in the figure, it includes a shared manufacturing platform, an enterprise cluster composed of multiple enterprises producing the same product, and several demand points. Demand orders are randomly generated at the demand points and sent to the shared manufacturing platform. After receiving the order information, the platform needs to decide which enterprise to allocate the order to for production. If the order is allocated to a production enterprise with idle production capacity, production will start immediately; otherwise, the order will enter the waiting queue. An order contains multiple products, a production enterprise contains several production devices, and a production device can only produce one product at a time and has a certain probability of malfunctioning. When a malfunction occurs during the production process, the task being processed will be suspended until the device is successfully repaired and restored to its normal working state. After the production enterprise produces all the products required by the order, it will transport them to the demand point of the order. Among them, the demand point is used to represent the geographical location where the product demander is located, and the demander will generate one or more orders. And the order information includes the quantity information of the products required by the order and the delivery address after the order is completed (i.e., the location of the demand point to which the order belongs).
[0081] Step S130, obtain the initial model of the order allocation strategy, where the initial model of the order allocation strategy is constructed according to the shared manufacturing simulation model and the DQN algorithm model with the shared manufacturing simulation model as the interaction object.
[0082] Among them, the DQN algorithm model is a model constructed based on the DQN algorithm. The main process of the DQN algorithm includes the following (1)-(5).
[0083] (1) The DQN algorithm uses a neural network to predict the reward values of various actions in a given state. The input of this neural network is the state s, and the output is the estimation of the long-term reward value for all possible actions a.
[0084] (2) The DQN algorithm uses an experience replay mechanism to store and replay past experiences (i.e., states, actions, rewards, and new states). This can disrupt the correlation between experiences during the training process, making the update of the neural network more effective.
[0085] (3) The DQN algorithm also uses a fixed Q-target mechanism, that is, it uses two neural networks with the same structure but different parameters (an action network and a target network) to predict Q values at the same time. The action network is used to select actions, and the target network is used to calculate the target values of Q values. This mechanism can make the target values not affected by the latest parameters, improving the stability of the neural network.
[0086] (4) During the training process, the DQN algorithm updates the parameters of the neural network by minimizing the gap between the predicted Q values and the target Q values. This gap is used as the loss function and is optimized by methods such as backpropagation and gradient descent.
[0087] (5) To balance exploration and exploitation, the DQN algorithm adopts the ε-greedy strategy algorithm, which selects random actions (exploration) with a certain probability and selects the actions currently estimated to be optimal with a higher probability (exploitation).
[0088] Step S140: Train the initial order allocation policy model based on the DQN algorithm to obtain a trained order allocation policy model.
[0089] Refer to Figure 3 , which shows the interaction principle between the shared manufacturing simulation model and the DQN algorithm model. The shared manufacturing simulation model includes m demand points, n manufacturing enterprises, and a shared manufacturing platform, and is used to simulate the operation process of the shared manufacturing system. The DQN algorithm model continuously optimizes the initial order allocation policy model through interactions with the shared manufacturing simulation model. Among them, each demand point in the shared manufacturing simulation model randomly generates orders. After generating an order, the DQN algorithm model makes an order allocation decision a t , that is, decides which enterprise will produce this order. The shared manufacturing simulation model returns a reward value r t , which is used to evaluate the quality of the decision made based on the order allocation policy and the policy itself. Based on multiple groups of (s t , a t , r t , s t , s t+1 ) to train and update the DQN algorithm model. After training, a trained order allocation policy model is obtained. This order allocation policy model can output an order allocation policy, and the order allocation policy is how the shared manufacturing system should allocate orders in different states.
[0090] The shared manufacturing simulation model of this embodiment includes a state acquisition module, a revenue acquisition module, an order arrival module, a queue module, an output order module, an order allocation module, and multiple manufacturing enterprise modules. Among them, the order arrival module, the queue module, the order allocation module, and the output order module are connected in sequence. Each manufacturing enterprise module includes a receiving order sub-module, a production sub-module, an output product information module connected in sequence, and a production equipment sub-module and a truck sub-module respectively connected to the production sub-module; the order allocation module is respectively connected to the production sub-modules in each manufacturing enterprise module.
[0091] Specifically, the constructed shared manufacturing simulation model is as Figure 4 shown. Figure 4In (a), demand_points[...] are agents representing a cluster of demand point objects; orders[..] are agents representing a cluster of order objects, i.e., the order objects generated by Order_arrival are stored in orders[..] for order data statistics; factories[..] are agents representing a cluster of factory objects for factory data statistics; order_id is a variable representing an order identifier; amoutDistribution represents a quantity distribution function; l ocationDistribution represents a location distribution function.
[0092] max_order_size is a parameter of the Order_arrival order arrival module, representing the maximum order size. For example, when the value of max_order_size is 1000, it means that when the order quantity reaches 1000, the shared manufacturing simulation model stops running; arrival_interval is a parameter of the Order_arrival order arrival module, representing the order arrival interval time. For example, when the value of arrival_interval is 10 minutes, it means that an order object is generated every 10 minutes; dqn represents the DQN algorithm parameter. When dqn is the first preset threshold, the DQN algorithm is not used, but the proximity allocation strategy is used (i.e., the Proximity_allocation proximity allocation function is called); when dqn is the second preset threshold, the DQN algorithm is used; episode_stop is a parameter used to judge whether the training is over. For example, when the value of episode_stop is 1000, it means that the training is carried out for 1000 rounds.
[0093] Proximity_allocation represents a proximity allocation function, which is used to find the production enterprise closest to the order demand point and allocate the order to this enterprise; getstate represents a state acquisition function, i.e., the state acquisition module mentioned above; getreward represents a reward acquisition function, i.e., the reward acquisition module mentioned above; takeaction is a function for the shared manufacturing simulation model to execute the order allocation decision generated by the DQN algorithm.
[0094] Order_arrival represents the order arrival module; queue represents the queue module, decision represents the order allocation module; exit represents the output order module; in the figure, the Order_arrival order arrival module, queue queue module, decision order allocation module, and exit output order module are connected in sequence.
[0095] Figure 4In (b), trucks[..] represents the truck sub-module; Take_order represents the order receiving sub-module; produce represents the production sub-module; equipments represents the production equipment sub-module; and exit represents the product output module. In the figure, the Take_order order receiving sub-module and the produce production sub-module are connected in sequence, and the exit product output module is connected to the trucks[..] truck sub-module and the equipments production equipment sub-module respectively.
[0096] Figure 4 In (b), name is a parameter of the produce production sub-module, representing the name of the production enterprise; process_capacity is a parameter of the equipments production equipment sub-module, representing the production capacity of a single device; num_eguipment is a parameter of the equipments production equipment sub-module, representing the number of devices; and num_wait_process is a variable representing the number of products to be produced.
[0097] In Figure 4 , the status acquisition module getstate is used to determine the status variable information of the current order and send the status variable information to the DQN algorithm model. Among them, the status variable information is shown in Equation (1).
[0098] s t =(d t ,l t ,c i *k it ,w it ), (1).
[0099] In the above formula, s t is the status variable in the t-th round; d t is the number of products to be produced for the order to be allocated in the t-th round; l t is the location of the demand point to which the order to be allocated in the t-th round belongs; c i *k it is the idle production capacity of the i-th production enterprise; among them, c i is the reciprocal of the average time taken to produce a single product; k it is the number of idle devices of the i-th production enterprise in the t-th round; w it is the number of products to be produced by the i-th production enterprise in the t-th round. Among them, "round" is defined as the time interval between two consecutive orders.
[0100] The revenue acquisition module getreward is used to determine the revenue information and send the revenue information to the DQN algorithm model; where the revenue information is shown in Equation (2). And the optimization goal of this embodiment is to minimize the order completion time in the shared manufacturing system.
[0101] FT j =WT j +PT j +DT j , (2).
[0102] In the above formula (2), FT j is the completion time of the j-th order; WT j is the waiting time of the j-th order; PT j is the production time of the j-th order; DT j is the transportation time of the j-th order.
[0103] The order arrival module Order_arrival is used to randomly generate multiple orders according to the demand points and determine the order information corresponding to each order. Among them, the order arrival module includes a demand point sub-module (i.e., the previous locationDistribution location distribution function), a product quantity sub-module (amoutDistribution quantity distribution function), and an order arrival interval sub-module (that is, when setting the parameters of the order arrival module, associate demand_points[...], amoutDistribution, and locationDistribution with Order_arrival). The demand point sub-module represents the location information of the demand point to which the order belongs, and it determines the location information of the demand point by calling the location data of the demand point in the historical order data and the location distribution function it contains. The product quantity module represents the quantity of products included in the order, and it determines the quantity of products included in the order by calling the product quantity data in the historical order data and the data distribution function it contains. The order arrival interval module is triggered by an event based on the time interval, and the order arrival interval sub-module arrival_interval is used to determine the time interval for each order to arrive, so as to generate multiple orders according to the time interval for each order to arrive. These orders are attached with order information, where the order information includes not only the time interval for each order to arrive, but also the location information of the demand point to which each order belongs and the product quantity information included in each order. In this way, the location information of the demand point to which each order belongs and the product quantity information included in each order will be transmitted to the production equipment sub-module of each manufacturing enterprise along with the order. The production equipment sub-module will produce the corresponding quantity of products according to the product quantity information, and the products produced by the production equipment sub-module will also carry the location information of the demand point. When the products are transported to the truck sub-module, the truck sub-module transports the products according to the location information of the demand point.
[0104] The queue module queue is used to queue each order and sequentially send the orders to the order allocation module according to the queuing order; and determine the waiting time WT of each order. The order allocation module is used to determine the production enterprise module corresponding to each order according to the order allocation policy information output by the DQN algorithm model. The output order module exit is used to output the order to the corresponding production enterprise module and sequentially send multiple order information to the corresponding production enterprise module.
[0105] The production capacity of each production enterprise is different, and the work processes of each enterprise are implemented using the process modeling library module. Figure 4 (b) lists the model of one of the established production enterprise modules. The order receiving sub-module Take_order in the production enterprise module is used to receive the order and order information and output the order and order information to the production sub-module. The production sub-module produce is used to call the production equipment sub-module to produce products according to the order information. The production equipment sub-module equipments is used to determine the production time of a single product by the target production equipment; the output order module is used to output the produced product to the truck sub-module after the order is completed. The truck sub-module is used to transport the product to the demand point according to the pre-established GIS map model and determine the transportation time. In addition, the truck sub-module can also customize its transportation speed according to the actual situation.
[0106] Preferably, the production equipment sub-module includes a processing capacity module, a mean time between failures module, and a repair time module. The failure time module is used to determine the failure occurrence time of the target production equipment according to the failure time data of each equipment in each production enterprise in the production enterprise characteristic data; the repair time module is used to determine the repair time of the target production equipment according to the repair time data of each equipment in each production enterprise in the production enterprise characteristic data; the processing capacity module is used to determine the production time of a single product by the target production equipment according to the production time data of a single product of each equipment in each production enterprise in the production enterprise characteristic data; wherein, the production time of a single product by the target production equipment is determined according to the production time of a single product by the target production equipment, the number of products in a single order, the repair time of the target production equipment, and the failure occurrence time of the target production equipment.
[0107] In some embodiments, the initial model of the order allocation strategy is trained based on the DQN algorithm to obtain a trained order allocation strategy model, including the following steps (1)-(3).
[0108] (1) Initialize the network parameters of the experience pool and the DQN algorithm model. Among them, the DQN algorithm model includes an action neural network and a target neural network, and the network parameters of the DQN algorithm model include the network parameters of the action neural network and the network parameters of the target neural network.
[0109] (2) Repeat the following steps (21)-(22) until the maximum number of training times is reached.
[0110] (21) Reset the shared manufacturing simulation model.
[0111] (22) Repeat the following steps (221)-(227) until all order productions are completed.
[0112] (221) Obtain the current state variable information sent by the shared manufacturing simulation model.
[0113] (222) Determine the order allocation strategy information according to the current state variable information and the ε-greedy policy algorithm, and send the order allocation strategy information to the shared manufacturing simulation model.
[0114] (223) Obtain the revenue information and the next state variable information sent by the shared manufacturing simulation model.
[0115] (224) Store the experience tuple in the experience pool, where the experience tuple includes the current state variable information, the order allocation strategy information, the revenue information, and the next state variable information.
[0116] (225) Randomly select a preset number of experience tuples from the experience pool for training.
[0117] (226) Determine the target value corresponding to each experience tuple according to the action neural network and the target neural network, and determine the loss function value.
[0118] (227) Use the backpropagation algorithm and the gradient descent algorithm to update the network parameters of the action neural network. After updating a preset number of times, copy the value of the updated network parameters of the action neural network to the target neural network to make the training process stable.
[0119] (3) Output the trained order allocation strategy model.
[0120] In the above steps, the corresponding algorithm pseudocode is shown in Table 1.
[0121] Table 1
[0122]
[0123]
[0124] In each round, the DQN algorithm model needs to decide which manufacturing enterprise to allocate the order to. The decision variable is a = (a1, a2,.., a i ), if the order is allocated to the i-th factory, a i = 1, otherwise 0. In this embodiment, the optimization objective of the order allocation strategy model is to minimize the order completion time. Therefore, the revenue value is the negative value of the order completion time in this round. Here, the smaller the order completion time, the better. Therefore, in the algorithm, the order completion time is taken as a negative value, and the larger the revenue value, the better.
[0125] In Table 1, at the beginning of each round, the aforementioned algorithm model needs to call the "getstate" function in the shared manufacturing simulation model to obtain the state variable s t , and then make a decision a t according to s t and the ε-greedy strategy, that is, decide which manufacturing enterprise to allocate the order to. Subsequently, the algorithm model calls the "takeaction" function and passes the order allocation decision a t into the shared manufacturing simulation model and executes it. At the end of the round, the algorithm model calls the "getstate" function again to obtain the current production system state variable s t+1 . Since the order may not be completed at the end of the round, the corresponding order completion time cannot be obtained. So at the end of each round, that is, when the order completion times of each round are known, the start state s t of each round, the action a t taken, the corresponding revenue value r t and the end state s t+1 of the round are stored in the experience pool. In the algorithm training stage, J samples are drawn from the experience pool each time, and the target value of sample j is calculated, that is, the sum of the action revenue value r j and the discounted long-term revenue value of the optimal action predicted by the target neural network for the next state; while the predicted value is the long-term total revenue value of the action predicted by the action neural network for the current state. Based on the difference between the target value and the predicted value of the drawn samples, the parameters of the action neural network are continuously adjusted, and after C adjustments, the parameters of the action neural network are copied to the parameters of the target neural network. Each round of training (episode) contains several rounds, and the shared manufacturing simulation model will return to the initial stage at the end of each round. By repeatedly training and continuously optimizing the neural network parameters, this algorithm can more accurately predict the long-term total revenue of taking various actions in different states and generate an order allocation strategy that optimizes the target.
[0126] It should be noted that the focus of this embodiment is on the order allocation strategy of the platform to optimize the supply-demand matching and scheduling process, and does not consider the costs, revenues, etc. of the platform. That is, this embodiment makes decisions by comprehensively considering the demand volume and the production capacity of each manufacturing enterprise.
[0127] The method of this embodiment can effectively solve the order allocation problem under the shared manufacturing mode. As Figure 5 shown, it is the training curve of the order allocation strategy model of this embodiment, that is, the change of the average completion time of all orders in this round with the increase of the training rounds. Every 5 rounds of training are completed, the shared manufacturing simulation model will run 20 times according to the DQN strategy at this time, and different random seeds will be set for each run. Subsequently, the mean value of the average completion time of the orders in these 20 simulation runs and the 95% confidence interval are plotted. The horizontal dotted line represents the average completion time of the orders when the nearest-neighbor allocation strategy is executed. The results show that the order allocation strategy model of this embodiment has a relatively fast convergence speed, and after convergence, compared with the nearest-neighbor allocation strategy used by real enterprises, the order completion time is shortened by 25.5%.
[0128] Referring to Figure 6 , an order allocation method under the shared manufacturing mode provided by this embodiment includes steps S210 - S220.
[0129] Step S210, obtain the orders to be allocated.
[0130] Step S220, input the orders to be allocated into the trained order allocation strategy model, and output the allocation strategy of the orders to be allocated, where the order allocation strategy model is constructed by the aforementioned order allocation strategy model construction method.
[0131] Using the order allocation method under the shared manufacturing mode of this embodiment does not require additional training and can be applied to different order arrival rates and order arrival distributions. ① The allocation strategy trained with the order arrival following a Poisson distribution with λ of 22 is used for the scenario where the order arrival follows a Poisson distribution with λ from 1 to 50. As Figure 7 shown, the results show that in different order arrival rate scenarios, the DQN optimization strategy is better than the nearest-neighbor allocation strategy, the order completion time is shortened by an average of 55.47%, and at least 51.81% can be shortened, and it is less affected by the change of the order arrival rate. ② The allocation strategy trained with the order arrival following a Poisson distribution with λ of 22 is used for the scenario where the order arrival follows a uniform distribution. As Figure 8 shown, compared with the nearest-neighbor allocation strategy, the DQN strategy shortens the average order completion time by 22.44% under U[17, 27]; by 18.66% under U[12, 32]; by 24.32% under U[7, 37]; and by 22.05% under U[2, 42].
[0132] Referring to Figure 9, an order allocation strategy model construction device provided in this embodiment includes: a data acquisition module 11, a simulation model construction module 12, an initial model acquisition module 13, and a training module 14. The data acquisition module 11 is used to acquire historical data, which at least includes historical order data and production enterprise characteristic data. The simulation model construction module 12 is used to construct a shared manufacturing simulation model according to the historical data and the production enterprise operation process. The initial model acquisition module 13 is used to acquire an initial order allocation strategy model, where the initial order allocation strategy model is constructed according to the shared manufacturing simulation model and a DQN algorithm model with the shared manufacturing simulation model as the interaction object. The training module 14 is used to train the initial order allocation strategy model based on the DQN algorithm to obtain a trained order allocation strategy model.
[0133] In an alternative embodiment, the shared manufacturing simulation model at least includes a state acquisition module, a revenue acquisition module, an order arrival module, a queue module, an order allocation module, and multiple production enterprise modules. Among them, the order arrival module, the queue module, and the order allocation module are connected in sequence. Each production enterprise module includes a production sub-module, a production equipment sub-module and a truck sub-module respectively connected to the production sub-module; the order allocation module is respectively connected to the production sub-modules in each production enterprise module;
[0134] The state acquisition module is used to determine state variable information and send the state variable information to the DQN algorithm model. The revenue acquisition module is used to determine revenue information and send the revenue information to the DQN algorithm model; among them, the revenue information is expressed as:
[0135] FT j =WT j +PT j +DT j ;
[0136] In the above formula, FT j is the completion time of the jth order; WT j is the waiting time of the jth order; PT j is the production time of the jth order; DT j is the transportation time of the jth order.
[0137] The order arrival module is used to randomly generate multiple orders according to demand points and determine the order information corresponding to each order. The queue module is used to queue each order and sequentially send the orders to the order allocation module according to the queuing order; and determine the waiting time of each order. The order allocation module is used to determine the production enterprise module corresponding to each order according to the order allocation policy information output by the DQN algorithm model, and sequentially send the multiple order information to the corresponding production enterprise modules. The production sub-module is used to call the production equipment sub-module to produce products according to the order information. The production equipment sub-module is used to determine the production time for the target production equipment to produce a single product. The truck sub-module is used to transport the products to the demand points according to the pre-established GIS map model and determine the transportation time.
[0138] In an alternative embodiment, the historical order data at least includes the location data of the demand point to which each historical order belongs and the product quantity data in each historical order.
[0139] In an alternative embodiment, the order arrival module includes a demand point sub-module, a product quantity sub-module, and an order arrival interval sub-module. The demand point sub-module is used to determine the location information of the demand point to which each order belongs according to the location data of the demand point to which each historical order belongs and the location distribution function. The product quantity module is used to determine the product quantity information included in each order according to the product quantity data in each historical order and the data distribution function. The order arrival interval module is used to determine the time interval for each order to arrive, and generate multiple orders according to the time interval for each order to arrive, where each generated order includes order information, and the order information includes the location information of the demand point to which each order belongs and the product quantity information included in each order.
[0140] In an alternative embodiment, the production enterprise characteristic data at least includes the location data of each production enterprise, the number of production equipment of each production enterprise, and the production time data, failure time data, and repair time data for each equipment in each production enterprise to produce a single product.
[0141] In an alternative embodiment, the production equipment sub-module includes a processing capacity module, a failure interval time module, and a repair time module. The failure time module is used to determine the failure occurrence time of the target production equipment according to the failure time data. The repair time module is used to determine the repair time of the target production equipment according to the repair time data; the processing capacity module is used to determine the production time for the target production equipment to produce a single product according to the production time data for producing a single product; wherein, the production time for the target production equipment to produce a single product is determined according to the production time for the target production equipment to produce a single product, the repair time of the target production equipment, and the failure occurrence time of the target production equipment.
[0142] In an alternative embodiment, the state variable information includes:
[0143] s t = (d t , l t , c it *k it , w it );
[0144] In the above formula, s t is the state variable in the t-th round; d t is the quantity of products to be produced for the order to be allocated in the t-th round; l t is the location of the demand point to which the order to be allocated in the t-th round belongs; c i *k it is the idle production capacity of the i-th production enterprise; where c i is the reciprocal of the average time used to produce a single product; k it is the number of idle devices of the i-th production enterprise in the t-th round; w it is the quantity of products to be produced by the i-th production enterprise in the t-th round.
[0145] In an alternative embodiment, the training module 14 includes: an initialization module, a repetition module, a selection module, a loss function module, an update parameter module, and an output module.
[0146] The initialization module is used to initialize the experience pool and the network parameters of the DQN algorithm model. Among them, the DQN algorithm model includes an action neural network and a target neural network, and the network parameters of the DQN algorithm model include the network parameters of the action neural network and the network parameters of the target neural network.
[0147] The repetition module is used to repeatedly execute the following modules: a first acquisition sub-module, a determination sub-module, a second acquisition sub-module, and a storage sub-module until the production of all orders is completed. The first acquisition sub-module is used to acquire the current state variable information sent by the shared manufacturing simulation model. The determination sub-module is used to determine the order allocation policy information according to the current state variable information and the ε-greedy policy algorithm, and send the order allocation policy information to the shared manufacturing simulation model. The second acquisition sub-module is used to acquire the revenue information and the next state variable information sent by the shared manufacturing simulation model. The storage sub-module is used to store the experience tuple in the experience pool, where the experience tuple includes the current state variable information, the order allocation policy information, the revenue information, and the next state variable information.
[0148] The selection module is used to randomly select a preset number of experience tuples from the experience pool for training. The loss function module is used to determine the target value corresponding to each experience tuple according to the action neural network and the target neural network, and determine the loss function value. The parameter update module is used to update the network parameters of the action neural network using the backpropagation algorithm and the gradient descent algorithm. After updating a preset number of times, the values of the updated network parameters of the action neural network are copied to the target neural network to make the training process stable. The output module is used to output the order allocation policy model that reaches the maximum number of training times.
[0149] Referring to Figure 10 , an order allocation device under a shared manufacturing mode provided in this embodiment includes an order acquisition module 21 and an order allocation module 22. The order acquisition module 21 is used to acquire orders to be allocated. The order allocation module 22 is used to input the orders to be allocated into the trained order allocation policy model and output the allocation policy of the orders to be allocated, where the order allocation policy model is constructed according to the aforementioned order allocation policy model construction method.
[0150] By using the device provided in the embodiment of the present application, since the device has the same inventive concept as the above method provided in the embodiment of the present application, and on the premise that the method can solve the technical problem, the device can also solve the technical problem, which will not be elaborated here.
[0151] Referring to Figure 11 , an embodiment of the present invention further provides an electronic device 400, including a communication interface 401, a processor 402, a memory 403, and a bus 404. The processor 402, the communication interface 401, and the memory 403 are connected through the bus 404; the above memory 403 is used to store a computer program that supports the processor 402 to execute the above order allocation policy model construction method and the order allocation method under the shared manufacturing mode, and the above processor 402 is configured to execute the program stored in the memory 403.
[0152] Optionally, an embodiment of the present invention further provides a computer-readable medium having non-volatile program code executable by a processor 402, and the program code causes the processor 402 to execute the order allocation policy model construction method and the order allocation method under the shared manufacturing mode as described in the above embodiments.
[0153] It is known by common technical knowledge that the present invention can be implemented by other embodiments that do not depart from its spirit or essential characteristics. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or equivalent to the present invention are encompassed by the present invention.
Claims
1. A method for constructing an order allocation strategy model, characterized in that Including: Obtain historical data, where the historical data at least includes historical order data and production enterprise characteristic data; The production enterprise characteristic data at least includes the location data of each production enterprise, the number of production equipment of each production enterprise, and the time data for a single product production, failure time data, and repair time data of each device in each production enterprise; Construct a shared manufacturing simulation model based on the historical data and the production enterprise operation process; Obtain an initial order allocation strategy model, where the initial order allocation strategy model is constructed based on the shared manufacturing simulation model and a DQN algorithm model with the shared manufacturing simulation model as the interaction object; Train the initial order allocation strategy model based on the DQN algorithm to obtain a trained order allocation strategy model; Among them, the shared manufacturing simulation model at least includes a state acquisition module, a revenue acquisition module, an order arrival module, a queue module, an order allocation module, and multiple production enterprise modules. Among them, the order arrival module, the queue module, and the order allocation module are connected in sequence. Each production enterprise module includes a production sub-module, a production equipment sub-module and a truck sub-module respectively connected to the production sub-module; the order allocation module is respectively connected to the production sub-modules in each production enterprise module; The state acquisition module is used to determine state variable information and send the state variable information to the DQN algorithm model; The revenue acquisition module is used to determine revenue information and send the revenue information to the DQN algorithm model; among them, the revenue information is expressed as: ; In the above formula, is the completion time of the th order; is the waiting time of the th order; is the production time of the th order; is the shipping time of the th order; The order arrival module is used to randomly generate multiple orders and determine the order information corresponding to each order; The queue module is used to queue each order, send the orders to the order allocation module in sequence according to the queuing order; and determine the waiting time of each order; The order allocation module is used to determine the production enterprise module corresponding to each order according to the order allocation strategy information output by the DQN algorithm model, and send multiple order information to the corresponding production enterprise modules in sequence; The production sub-module is used to call the production equipment sub-module to produce products according to the order information; The production equipment sub-module is used to determine the production time for a single product produced by the target production equipment; The truck sub-module is used to transport the product to the demand point according to the pre-established GIS map model and determine the transportation time.
2. The method for constructing an order allocation strategy model according to claim 1, characterized in that The historical order data at least includes the location data of the demand point to which each historical order belongs and the product quantity data in each historical order.
3. The method for constructing an order allocation strategy model according to claim 2, wherein The order arrival module includes a demand point sub-module, a product quantity sub-module, and an order arrival interval sub-module; The demand point sub-module is used to determine the location information of the demand point to which each order belongs according to the location data of the demand point to which each historical order belongs and the location distribution function; The product quantity module is used to determine the product quantity information included in each order according to the product quantity data in each historical order and the data distribution function; The order arrival interval module is used to determine the time interval between the arrivals of each order, and generate a plurality of orders according to the time interval between the arrivals of each order. Each generated order includes order information, and the order information includes the location information of the demand point to which each order belongs and the quantity information of the products included in each order.
4. The method for constructing an order allocation strategy model according to claim 1, wherein The production enterprise characteristic data at least includes the location data of each production enterprise, the number of production equipment of each production enterprise, and the time data for each device in each production enterprise to produce a single product, the failure time data, and the repair time data.
5. The method for constructing an order allocation strategy model according to claim 4, wherein The production equipment sub-module includes a processing capacity module, a failure interval time module, and a repair time module; The failure time module is used to determine the failure occurrence time of the target production equipment according to the failure time data; The repair time module is used to determine the repair time of the target production equipment according to the repair time data; The processing capacity module is used to determine the production time used by the target production equipment to produce a single product according to the time data for producing a single product; Among them, the production time of the target production equipment for producing the single product is determined according to the production time used by the target production equipment to produce a single product, the repair time of the target production equipment, and the failure occurrence time of the target production equipment.
6. The method for constructing an order allocation strategy model according to claim 1, wherein The state variable information includes: ; In the above formula, is the state variable in the t-th round; is t the quantity of products to be produced for the order to be allocated in the round; t is the location of the demand point to which the order to be allocated in the round belongs; is the idle production capacity of the -th production enterprise; where is the reciprocal of the average time taken to produce a single product; is the number of idle devices of the -th production enterprise in the round; is the quantity of products to be produced by the -th production enterprise in the round.
7. The method for constructing the order allocation strategy model according to claim 1, wherein Training the initial order allocation strategy model based on the DQN algorithm to obtain a trained order allocation strategy model, including the following steps (1)-(3): (1) Initialize the experience pool and the network parameters of the DQN algorithm model. Among them, the DQN algorithm model includes an action neural network and a target neural network, and the network parameters of the DQN algorithm model include the network parameters of the action neural network and the network parameters of the target neural network; (2) Repeat the following steps (21)-(22) until the preset number of training times is reached; (21) Reset the shared manufacturing simulation model; (22) Repeat the following steps (221)-(227) until the production of all orders is completed; (221) Obtain the current state variable information sent by the shared manufacturing simulation model; (222) Based on the current state variable information and - determine the order allocation policy information according to the greedy policy algorithm, and send the order allocation policy information to the shared manufacturing simulation model; (223) Obtain the revenue information and the next state variable information sent by the shared manufacturing simulation model; (224) Store the experience tuple into the experience pool, where the experience tuple includes the current state variable information, the order allocation strategy information, the revenue information, and the next state variable information; (225) Randomly select a preset number of experience tuples from the experience pool for training; (226) Determine the target value corresponding to each experience tuple according to the action neural network and the target neural network, and determine the loss function value; (227) Use the backpropagation algorithm and the gradient descent algorithm to update the network parameters of the action neural network. After updating a preset number of times, copy the value of the network parameters of the action neural network to the target neural network to make the training process stable; (3) Output the trained order allocation strategy model.
8. An order allocation method under a shared manufacturing model, characterized in that, Include: Obtain the order to be allocated; Input the order to be allocated into the trained order allocation strategy model, and output the allocation strategy of the order to be allocated, where the order allocation strategy model is constructed according to the order allocation strategy model construction method described in any one of claims 1-7.
9. An order allocation strategy model construction device for executing the order allocation strategy model construction method according to any one of claims 1-7, characterized in that It includes: A data acquisition module for acquiring historical data, where the historical data at least includes historical order data and production enterprise feature data; The production enterprise feature data at least includes the location data of each production enterprise, the number of production equipment of each production enterprise, and the time data for each equipment in each production enterprise to produce a single product, the failure time data, and the maintenance time data; A simulation model construction module for constructing a shared manufacturing simulation model according to the historical data and the production enterprise operation process; An initial model acquisition module for acquiring an initial order allocation strategy model, where the initial order allocation strategy model is constructed according to the shared manufacturing simulation model and a DQN algorithm model with the shared manufacturing simulation model as the interaction object; A training module for training the initial order allocation strategy model based on the DQN algorithm to obtain a trained order allocation strategy model.
Citation Information
Patent Citations
Prefabricated part production scheduling optimization method and system based on reinforcement learning
CN115204497A
Enterprise order collaborative manufacturing scheduling optimization method and system based on improved SPSA
CN119205048A