Self-adaptive production scheduling method and system based on multi-dimensional environment perception
Through multi-dimensional environmental perception and adaptive production scheduling methods, workshop data is collected in real time, equipment efficiency attenuation model is constructed, and production scheduling is optimized using NSGA-II and DRL algorithms, which solves the problem of environmental parameters not being considered in the existing technology, and achieves efficient and flexible production scheduling.
Patent Information
- Application Number
- CN202510579984.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-15
AI Technical Summary
The existing production scheduling system fails to effectively consider the workshop environmental parameters, resulting in large deviations from the scheduling plan and the actual working conditions, insufficient dynamic response capabilities, limited multi-objective optimization capabilities, lack of real-time feedback and adaptive learning capabilities, and insufficient emergency response capabilities, resulting in high risk of production interruption.
A multi-dimensional environment perception method is used to collect workshop environmental parameters and equipment status data in real time, build a equipment efficiency attenuation model, combine NSGA-II and DRL algorithms to generate production scheduling solutions, and use MCTS algorithm to deal with abnormal situations to achieve adaptive production scheduling.
It achieves rapid response without manual intervention, improves production efficiency and quality, reduces the risk of production interruption, improves equipment utilization and energy consumption efficiency, and enhances the flexibility and adaptability of the system.
Smart Images

Figure CN120491568A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of production scheduling, and in particular to an adaptive production scheduling method and system based on multi-dimensional environmental perception. Background Art
[0002] Production scheduling systems (APSs) are indispensable tools in modern manufacturing. Their core function is to improve production efficiency, reduce costs, and ensure on-time order delivery by optimizing the sequence and timing of production tasks. The fundamental principle of APS systems is to generate optimal or near-optimal production plans based on mathematical models and optimization algorithms, taking into account constraints such as a company's resource capabilities, order requirements, and process flows. APS systems typically require input from multiple data sources, including order data, resource data, process data, and various constraints. This data is typically integrated and shared through enterprise resource planning (ERP), manufacturing execution systems (MES), or other production management systems. In terms of optimization algorithms, production scheduling systems employ a variety of techniques, from simple heuristics such as "shortest processing time first" or "earliest delivery date first" in the early days, to mathematical programming methods such as linear programming and mixed integer programming, and now to artificial intelligence technologies such as deep learning and reinforcement learning. These technological advancements have enabled production scheduling systems to handle increasingly complex constraints and are gradually moving towards intelligent control.
[0003] Although production scheduling systems have been widely used in the manufacturing industry, existing production scheduling systems still have the following limitations: (1) The lack of production environment factors leads to a large deviation between the scheduling plan and the actual working conditions: Existing production scheduling systems (such as traditional APS systems) mainly rely on order data and equipment static parameters, and do not consider the impact of workshop environmental parameters (such as temperature and humidity, vibration, and power fluctuations) on production efficiency and quality. For example, high temperature may cause equipment performance to decline, high humidity may increase the rejection rate of electronic components, and excessive vibration may cause equipment failure. (2) Insufficient dynamic response capabilities lead to extended production interruption time: Existing production scheduling systems face sudden environmental changes (such as equipment failure, order insertion, raw material shortage) or dynamic changes in the production process. Manual intervention and adjustment are required, and the response time is long (usually more than 2 hours). It is impossible to adjust the scheduling plan in time and cannot meet the requirements of modern manufacturing for real-time and flexibility. (3) Limited multi-objective optimization capabilities lead to low resource utilization efficiency: Existing production scheduling algorithms (such as genetic algorithms and linear programming) are difficult to optimize multiple objectives (such as equipment utilization, energy consumption costs, and on-time delivery rates of orders, etc.) at the same time, resulting in equipment utilization rates generally below 75%, high energy consumption and costs, and the inability to achieve optimal configuration in terms of resource utilization and cost control. (4) Lack of real-time feedback and adaptive learning capabilities lead to poor system flexibility: Existing production scheduling systems are unable to perceive changes in the production environment in real time and dynamically adjust scheduling strategies. They lack adaptive learning capabilities, making it difficult for the system to cope with complex and changing production scenarios, and unable to dynamically optimize scheduling plans based on real-time data, resulting in poor flexibility. (5) Insufficient emergency response capabilities lead to a high risk of production interruption: Existing production scheduling systems lack effective emergency response mechanisms when faced with abnormal emergencies such as sudden equipment failures and production accidents, and are unable to quickly generate feasible emergency scheduling plans, resulting in extended production interruption time and a high risk of production interruption. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide an adaptive production scheduling method and system based on multi-dimensional environmental perception that does not require human intervention, has a short response time, high flexibility, and has adaptive learning ability, and can dynamically optimize production scheduling plans according to real-time production environment data and quickly generate emergency scheduling plans.
[0005] The technical solution adopted by the present invention to solve the technical problem is: an adaptive production scheduling method based on multi-dimensional environmental perception, comprising the following steps:
[0006] S1. Collection of production environment data: real-time collection of workshop environment parameters, equipment status data, order data, and material status data;
[0007] S2. Construction of an equipment efficiency decay model: Based on the LSTM time series prediction algorithm, an equipment efficiency decay model is constructed to analyze the relationship between production environment data and equipment efficiency.
[0008] S3. Construction of a multi-objective optimization model: The multi-objective optimization model includes optimization objectives and constraints. The optimization objectives include maximizing equipment utilization, minimizing order delay rate, and minimizing energy consumption cost. The constraints include equipment capacity constraints, process sequence constraints, resource constraints, and delivery time constraints.
[0009] S4. Production scheduling plan generation: Based on the multi-objective optimization model and the equipment efficiency decay model, the NSGA-II algorithm is used to generate the production scheduling plan;
[0010] S5. Dynamic response to production environment changes: Based on the multi-objective optimization model and equipment efficiency attenuation model, when the production environment data change is lower than the preset threshold, the DRL algorithm is used to quickly adjust the production scheduling plan; when the production environment data change exceeds the preset threshold, the MCTS algorithm is used to generate an emergency scheduling plan.
[0011] Furthermore, the equipment efficiency decay model in step S2 includes an input layer, an LSTM layer and a fully connected layer. The input layer receives time series data of workshop environmental parameters and equipment status data. The LSTM layer contains multiple LSTM units for capturing dynamic changes in time series data. The fully connected layer maps the output of the LSTM layer to the predicted value of the efficiency decay coefficient.
[0012] Furthermore, the LSTM unit includes a forget gate, an input gate, and an output gate. The forget gate determines how much long-term information to retain and which long-term information to forget. The input gate determines how much new information to absorb into the long-term memory and updates the unit state used to control the long-term memory. The output gate filters out the short-term information that is most suitable for the current time step from the new long-term information.
[0013] Furthermore, the forget gate uses the following formula to determine how much long-term information to retain:
[0014] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0015] Where, f t is the ratio of the forget gate output, f t ∈[0,1]; σ is the sigmoid function; W f is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x tis the information of the current time step; b f is the offset matrix; t is the time step;
[0016] The input gate uses the following formula to determine how much new information to incorporate into long-term memory:
[0017]
[0018] Where, To integrate new information into long-term memory; is the total new information absorbed in the current time step; i t is the ratio of screening new information; σ is the sigmoid function; W t 、W c is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b t 、b c is the offset matrix;
[0019] The input gate updates the cell state used to control long-term memory using the following formula:
[0020]
[0021] Where C t is the long-term information of the current time step; f t is the ratio of the forget gate output; C t-1 is the long-term information of the previous time step; To integrate new information into long-term memory;
[0022] The output gate uses the following formula to filter out the short-term information that is most suitable for the current time step from the new long-term information:
[0023]
[0024] Where h t To filter out the short-term information that is most suitable for the current time step from the new long-term information; t is the proportion of screening new information; C t is the long-term information of the current time step; σ is the sigmoid function; W o is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b o is the offset matrix.
[0025] Furthermore, the specific steps of generating a production scheduling plan using the NSGA-II algorithm in step S4 are as follows:
[0026] S41. Initialize the population: Randomly generate an initial population, where each individual represents a production schedule, and calculate the objective function value for each individual.
[0027] S42. Non-dominated sorting and crowding calculation: Perform non-dominated sorting on the initial population, dividing the population into different non-dominated hierarchies; calculate the crowding degree of each individual in the non-dominated hierarchy;
[0028] S43. Selection, crossover, and mutation operations: Select individuals from the current population to generate the next generation population; exchange partial gene fragments of two parent individuals to generate new offspring individuals; and randomly change certain genes of individuals.
[0029] S44. Determine whether the termination condition is met: If the termination condition is met, output the Pareto optimal solution set in the current population; otherwise, execute step S45;
[0030] S45. Population merging: All individuals of the parent and offspring generations are mixed to form a larger population, and then return to step S42 for iteration.
[0031] Furthermore, the specific steps of using the DRL algorithm to quickly adjust the production scheduling plan in step S5 are as follows:
[0032] S51. Define the state space, action space, and reward function: The state space includes workshop environment parameters, equipment status data, order data, and material status data. The action space includes equipment start and stop, production batch splitting, line change sequence adjustment, equipment parameter adjustment, and order priority adjustment. The reward function formula is as follows:
[0033] Reward=α×OEE+β×OnTimeDelivery-γ×(η×EnergyCost)
[0034] Where Reward is the reward value obtained at each time step; OEE is an indicator to measure equipment utilization; OnTimeDelivery is a measure of the proportion of orders completed on time; EnergyCost is a measure of energy consumption in the production process; η is the equipment efficiency attenuation factor calculated based on the workshop environmental parameters; α, β, and γ are weight coefficients used to balance the importance of different objectives.
[0035] S52. Use the improved DDPG+TD3 algorithm as the basic framework: The basic framework includes the actor network, critic network, target network, and buffer. Initialize the actor network, critic network, target network, and buffer, and set hyperparameters.
[0036] S53. Introducing hierarchical reinforcement learning: High-level controllers with longer action cycles are responsible for production line scheduling, while low-level controllers with shorter action cycles are responsible for equipment parameter adjustment. The high-level and low-level controllers work together.
[0037] S54. Introducing multi-agent collaboration: Each device has an independent agent responsible for local decision-making. The central critic network coordinates the decisions of each agent to obtain the globally optimal production scheduling strategy.
[0038] S55. Use historical production data to train the iterative DRL algorithm: Use the experience replay mechanism to randomly sample historical production environment data from the buffer for training until the set termination condition is reached, and continuously update.
[0039] Furthermore, the specific steps of generating the emergency production scheduling plan using the MCTS algorithm in step S5 are as follows:
[0040] S51' initializes the root node of the tree, the root node is the starting point of the current production state, contains all the data of the current production environment;
[0041] S52′. Select branch nodes according to the UCB1 formula until reaching an incompletely expanded branch node. Select an unexplored child node on the incompletely expanded branch node for expansion. Generate a new state based on the current state and possible decision actions, and add it to the tree as a new branch node.
[0042] S53′. Starting from the expanded node, simulate the production situation in the future until the termination state is reached;
[0043] S54'. Back propagate the simulation results back to the root node, update the information of each node on the path, and optimize the decision path;
[0044] S55 '. Loop steps S52 '-S54 ', evaluating various possible decision paths;
[0045] S56′. Select the optimal decision path and generate an emergency production scheduling plan.
[0046] An adaptive production scheduling system based on multi-dimensional environmental perception adopts the above-mentioned adaptive production scheduling method based on multi-dimensional environmental perception, including a sensor layer, an edge computing layer, a cloud platform layer and an execution terminal; the sensor layer collects workshop environmental parameters, equipment status data, order data and material status data in real time, transmits the collected data to the edge computing layer in real time, and uploads it to the cloud platform layer at the same time; the edge computing layer generates a production scheduling plan based on the received data, generates scheduling instructions based on the production scheduling plan and sends them to the execution terminal in a timely manner, and uploads a summary of key data to the cloud platform layer at the same time; the cloud platform layer is used for data storage, global optimization, model iteration and result display.
[0047] Furthermore, the sensor layer includes temperature and humidity sensors, vibration sensors, smart meters and RFID tags.
[0048] Furthermore, the cloud platform layer includes a database, a model training module, a visualization dashboard and a federated learning center.
[0049] The beneficial effects of the present invention are:
[0050] (1) The present invention is based on a multi-objective optimization model and an equipment efficiency attenuation model, and adopts the NSGA-II algorithm to generate a production scheduling plan, so that the production scheduling plan fully integrates the production environment factors and can achieve multi-objective optimization, thereby ensuring production efficiency and quality, and achieving the optimal configuration of multiple objectives; at the same time, when the production environment data changes below the preset threshold (i.e., normal changes), the DRL algorithm is used to quickly adjust the production scheduling plan without manual intervention, with a short response time and adaptive learning ability, and can dynamically optimize the production scheduling plan according to real-time production environment data, with high flexibility; when the production environment data changes exceed the preset threshold (i.e., abnormal changes), the MCTS algorithm is used to generate an emergency scheduling plan, which realizes the rapid generation of the emergency scheduling plan, avoids long-term production interruption, and reduces the risk of production interruption.
[0051] (2) The DRL algorithm in the present invention adopts the improved DDPG+TD3 algorithm as the basic framework, and introduces hierarchical reinforcement learning and multi-agent collaboration. The target network and double delay mechanism in the improved DDPG+TD3 algorithm improve the stability and convergence speed of the algorithm; hierarchical reinforcement learning decomposes the decision-making process into high-level controllers and low-level controllers, which improves the decision-making efficiency and flexibility of the algorithm; multi-agent collaboration allows each device to make independent decisions, and at the same time coordinates through the central critic network to obtain the globally optimal production scheduling strategy; thereby being able to more effectively cope with dynamic changes and complex constraints in the production environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] The present invention will be further described below with reference to the accompanying drawings and examples.
[0053] Figure 1 It is a flow chart of the adaptive production scheduling method based on multi-dimensional environmental perception in the present invention;
[0054] Figure 2 It is a framework diagram of the adaptive production scheduling system based on multi-dimensional environmental perception in the present invention. DETAILED DESCRIPTION
[0055] The present invention will now be further described with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.
[0056] Example 1
[0057] like Figure 1 As shown, an adaptive production scheduling method based on multi-dimensional environmental perception includes the following steps:
[0058] S1. Collection of production environment data: Real-time collection of workshop environment parameters, equipment status data, order data and material status data.
[0059] S2. Construction of equipment efficiency decay model: Based on the LSTM (long short-term memory) time series prediction algorithm, an equipment efficiency decay model is constructed to show the relationship between production environment data and equipment efficiency.
[0060] S3. Construction of a multi-objective optimization model: The multi-objective optimization model includes optimization objectives and constraints. The optimization objectives include maximizing equipment utilization, minimizing order delay rate, and minimizing energy consumption cost. The constraints include equipment capacity constraints, process sequence constraints, resource constraints, and delivery time constraints.
[0061] S4. Generation of production scheduling plan: Based on the multi-objective optimization model and the equipment efficiency decay model, the NSGA-II (non-dominated sorting genetic) algorithm is used to generate the production scheduling plan.
[0062] S5. Dynamic response to changes in the production environment: Based on a multi-objective optimization model and an equipment efficiency attenuation model, when the change in production environment data is lower than a preset threshold, the DRL (deep reinforcement learning) algorithm is used to quickly adjust the production scheduling plan; when the change in production environment data exceeds the preset threshold, the MCTS (Monte Carlo Tree Search) algorithm is used to generate an emergency scheduling plan.
[0063] Based on the multi-objective optimization model and the equipment efficiency attenuation model, the NSGA-II algorithm is used to generate a production scheduling plan, so that the production scheduling plan fully integrates the production environment factors and can achieve multi-objective optimization, thereby ensuring production efficiency and quality, and achieving the optimal configuration of multiple objectives; at the same time, when the production environment data change is lower than the preset threshold (i.e. normal change), the DRL algorithm is used to quickly adjust the production scheduling plan without manual intervention, with a short response time and adaptive learning ability. It can dynamically optimize the production scheduling plan according to real-time production environment data and has high flexibility; when the production environment data change exceeds the preset threshold (i.e. abnormal change), the MCTS algorithm is used to generate an emergency scheduling plan, which can realize the rapid generation of the emergency scheduling plan, avoid long-term production interruption, and reduce the risk of production interruption.
[0064] The equipment efficiency decay model in step S2 includes an input layer, an LSTM layer, and a fully connected layer. The input layer receives time series data of workshop environmental parameters and equipment status data. The LSTM layer contains multiple LSTM units to capture dynamic changes in time series data. The fully connected layer maps the output of the LSTM layer to the predicted value of the efficiency decay coefficient.
[0065] The equipment efficiency attenuation model calculates the equipment's operating efficiency attenuation coefficient under different production environments, providing an important basis for subsequent production scheduling optimization.
[0066] The LSTM unit consists of a forget gate, an input gate, and an output gate. The forget gate determines how much long-term information to retain and which long-term information to forget. The input gate determines how much new information to absorb into the long-term memory and updates the unit state used to control the long-term memory. The output gate filters out the short-term information that is most suitable for the current time step from the new long-term information.
[0067] The forget gate uses the following formula to determine how much long-term information to retain:
[0068] f t =σ(W f ·[h t-1 ,x t ]+b f )
[0069] Where, f t is the ratio of the forget gate output, f t ∈[0,1]; σ is the sigmoid function; W f is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b f is the offset matrix; t is the time step.
[0070] The forget gate can combine contextual information, gradient information from the loss function, and historical information to calculate a new, retained long-term memory.
[0071] The input gate uses the following formula to determine how much new information to incorporate into long-term memory:
[0072]
[0073] Where, To integrate new information into long-term memory; is the total new information absorbed in the current time step; i t is the ratio of screening new information; σ is the sigmoid function; W t 、W c is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b t 、b c is the offset matrix.
[0074] The input gate updates the cell state, which controls the long-term memory, using the following formula:
[0075]
[0076] Where C t is the long-term information of the current time step; f t is the ratio of the forget gate output; C t-1 is the long-term information of the previous time step; To incorporate new information into long-term memory.
[0077] The output gate uses the following formula to filter out the short-term information that is most suitable for the current time step from the new long-term information:
[0078]
[0079] Where h t To filter out the short-term information that is most suitable for the current time step from the new long-term information; t is the proportion of screening new information; C t is the long-term information of the current time step; σ is the sigmoid function; W o is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b o is the offset matrix.
[0080] The specific steps of generating a production scheduling plan using the NSGA-II algorithm in step S4 are as follows:
[0081] S41. Initialize the population: Randomly generate an initial population. Each individual in the initial population represents a production scheduling plan, and calculate the objective function value of each individual.
[0082] Specifically, the individual encoding method can be machine allocation, process sequence, etc.; the objective function value includes equipment utilization, order delay rate and energy consumption cost.
[0083] S42. Non-dominated sorting and crowding calculation: Perform non-dominated sorting on the initial population and divide the population into different non-dominated levels; calculate the crowding of individuals in each non-dominated level.
[0084] Specifically, the first layer contains all non-dominated individuals. After removing the first layer individuals from the population, the remaining individuals are sorted again by non-domination to obtain the second layer, and so on. The method for calculating the crowding degree is: for each objective function, the neighborhood distance of the individual on its target value is calculated, and the neighborhood distances of all target values are added together to obtain the total crowding degree of the individual.
[0085] S43. Selection, crossover, and mutation operations: Select individuals from the current population to generate the next generation population; exchange some gene fragments of two parent individuals to generate new offspring individuals; and randomly change certain genes of an individual.
[0086] Specifically, during the selection operation, the individual with the lowest non-dominated hierarchy (i.e., the strongest dominance relationship) is selected first. Within the same non-dominated hierarchy, individuals with high crowding are given priority to ensure the diversity of the population. During the mutation operation, certain genes of the individual are randomly changed to introduce new genetic mutations to prevent the algorithm from falling into a local optimum.
[0087] S44. Determine whether the termination condition is met: If the termination condition is met, output the Pareto optimal solution set in the current population; otherwise, execute step S45.
[0088] Specifically, the termination condition may be a preset number of iterations, or the population diversity being reduced to a certain level.
[0089] S45. Population merging: All individuals of the parent and offspring generations are mixed to form a larger population, and then return to step S42 for iteration.
[0090] The specific steps for using the DRL algorithm to quickly adjust the production scheduling plan in step S5 are as follows:
[0091] S51. Define the state space, action space, and reward function: The state space includes workshop environment parameters, equipment status data, order data, and material status data. The action space includes equipment start and stop, production batch splitting, line change sequence adjustment, equipment parameter adjustment, and order priority adjustment. The reward function formula is as follows:
[0092] Reward=α×OEE+β×OnTimeDelivery-γ×(η×EnergyCost)
[0093] Where Reward is the reward value obtained at each time step; OEE is an indicator to measure equipment utilization; OnTimeDelivery is a measure of the proportion of orders completed on time; EnergyCost is a measure of energy consumption in the production process; η is the equipment efficiency attenuation factor calculated based on the workshop environmental parameters; α, β, and γ are weight coefficients used to balance the importance of different objectives.
[0094] Specifically, workshop environmental parameters include workshop temperature and humidity, vibration intensity, power load fluctuations, etc.; equipment status data includes equipment load rate, fault code, comprehensive efficiency, etc.; order data includes order urgency, delivery time, quantity, etc.; material status data includes raw material inventory level, supplier delivery punctuality rate, etc.; equipment start and stop determines which equipment should be started or stopped, production batch splitting determines whether to split a production batch into multiple small batches, line change sequence adjustment determines the switching sequence of the production line, equipment parameter adjustment adjusts the equipment's operating parameters (such as speed, temperature, etc.), and order priority adjustment rearranges the production order of orders.
[0095] The state space provides the DRL algorithm with a comprehensive "environmental snapshot," enabling it to make informed decisions based on the current state of the production environment. The action space defines all possible actions the DRL algorithm can perform at each time step, which directly impacts the efficiency and flexibility of the production process. The reward function, which defines the reward value the DRL algorithm receives at each time step, is a key factor in measuring the quality of decisions within the DRL algorithm.
[0096] S52. Use the improved DDPG+TD3 algorithm as the basic framework: The basic framework includes the actor network, critic network, target network and buffer. Initialize the actor network, critic network, target network and buffer, and set hyperparameters.
[0097] Specifically, the improved DDPG+TD3 algorithm is used to learn the mapping relationship between states and actions and is applicable to continuous action spaces; hyperparameters include learning rate, discount factor, experience replay, exploration strategy and batch size. The initial value of the learning rate is 3e-4, and it is adjusted using the cosine annealing strategy. The discount factor is 0.99, which is used to measure the importance of long-term rewards. The experience replay capacity is 1e6 samples, and the priority experience replay strategy is adopted. The batch size is set to 256; the update frequency of the target network is synchronized once every 1000 steps.
[0098] The improved target network and dual-delay mechanism in the DDPG+TD3 algorithm improve the algorithm's stability and convergence speed. The dual-delay mechanism refers to the TD3 algorithm's introduction of two critic networks on top of the DDPG actor-critic network structure.
[0099] S53. Introduce hierarchical reinforcement learning (HRL): A high-level controller with a longer action cycle (such as an action cycle of 10 minutes) is responsible for production line-level scheduling, and a low-level controller with a shorter action cycle (such as an action cycle of 1 minute) is responsible for equipment parameter adjustment. The high-level controller and the low-level controller work together.
[0100] Hierarchical reinforcement learning decomposes the decision-making process into high-level controllers and low-level controllers, improving the decision-making efficiency and flexibility of the algorithm.
[0101] S54. Introduce multi-agent collaboration: Each device sets up an independent agent responsible for local decision-making, and coordinates the decisions of each agent through the central critic network to ensure global optimization.
[0102] Multi-agent collaboration allows each device to make independent decisions while coordinating through a central critic network to obtain the globally optimal scheduling strategy.
[0103] S55. Use historical production data to train the iterative DRL algorithm: Use the experience replay mechanism to randomly sample historical production data from the buffer for training until the set conditions are met, and continuously update.
[0104] The DRL algorithm uses the improved DDPG+TD3 algorithm as the basic framework, and introduces hierarchical reinforcement learning and multi-agent collaboration, which can more effectively cope with dynamic changes and complex constraints in the production environment.
[0105] When production environment data fluctuations fall below a preset threshold (i.e., normal fluctuations), the DRL algorithm can quickly adjust the production schedule, achieving real-time dynamic optimization. Normal fluctuations include environmental changes, sudden orders, and so on.
[0106] The specific steps for generating an emergency production scheduling plan using the MCTS algorithm in step S5 are as follows:
[0107] S51′. Initialize the root node of the tree. The root node is the starting point of the current production status and contains all the data of the current production environment.
[0108] S52′. Select branch nodes according to the UCB1 formula until you reach an incompletely expanded branch node (i.e., there are unexplored child nodes), select an unexplored child node on the incompletely expanded branch node for expansion, generate a new state based on the current state and possible decision actions, and add it to the tree as a new branch node.
[0109] Specifically, in the process of selecting branch nodes, those with higher potential value are given priority; the UCB1 formula is the existing technology; branch nodes are different decision paths that the system can take based on the current production status, and each branch node corresponds to a possible decision action.
[0110] S53′. Starting from the expanded node, simulate the production situation in the future period until the termination state is reached.
[0111] Specifically, during the simulation process, factors such as production speed, burst priority, and other equipment status are considered until a terminal state is reached (such as completing a production task or reaching a certain time point); the terminal state is a leaf node, that is, the final result of a decision path; the simulation adopts a random strategy or a heuristic strategy.
[0112] S54′. Backpropagate the simulation results (such as reward values) back to the root node, update the information of each node on the path, and optimize the decision path.
[0113] Through backpropagation, the system can learn which paths are more likely to lead to good outcomes.
[0114] S55′. Loop steps S52′-S54′ to evaluate various possible decision paths.
[0115] S56′. Select the optimal decision path and generate an emergency production scheduling plan.
[0116] When changes in production environment data exceed a preset threshold (i.e., abnormal changes), the MCTS algorithm can quickly evaluate various possible decision paths and select the optimal emergency scheduling plan to ensure production continuity and on-time order delivery. Abnormal changes include equipment failures, production accidents, and other such changes.
[0117] According to experimental data (as shown in the table below), this application can significantly improve equipment utilization, reduce energy consumption costs, and shorten production interruption time.
[0118] index Traditional methods This application Overall Equipment Efficiency 68%-72% 82%-87% Order delay rate 15%-20% <5% Abnormal response time >2 hours <5 minutes Energy consumption cost per unit output value Baseline value 100% Reduce by 20%-25%
[0119] Example 2
[0120] like Figure 2 As shown, an adaptive production scheduling system based on multi-dimensional environmental perception adopts the adaptive production scheduling method based on multi-dimensional environmental perception in Example 1, including a sensor layer, an edge computing layer, a cloud platform layer and an execution terminal; the sensor layer collects workshop environmental parameters, equipment status data, order data and material status data in real time, transmits the collected data to the edge computing layer in real time, and uploads it to the cloud platform layer at the same time; the edge computing layer generates a production scheduling plan based on the received data, generates scheduling instructions based on the production scheduling plan and sends them to the execution terminal in a timely manner, and uploads key data summaries to the cloud platform layer; the cloud platform layer is used for data storage, global optimization, model iteration and result display. Specifically, the sensor layer includes temperature and humidity sensors, vibration sensors, smart meters and RFID tags; the cloud platform layer includes a database, a model training module, a visual dashboard and a federated learning center. The database stores all historical data, the visual dashboard displays the comprehensive efficiency of equipment, energy consumption heat map and production scheduling Gantt chart, and the federated learning center is used for cross-enterprise model collaborative optimization; the execution terminal includes an MES system and a PLC controller.
[0121] During operation, the sensor layer collects workshop environmental parameters, equipment status data, order data, and material status data in real time. It transmits this collected production environment data to the edge computing layer via industrial protocols for preprocessing and computational optimization. The sensor layer then uploads the collected data to the cloud platform layer, while the edge computing layer also uploads the optimization plan and results to the cloud platform layer. Preprocessing at the edge computing layer includes operations such as data filtering and normalization to ensure data quality and availability. Furthermore, the cloud platform layer regularly updates and optimizes the global model based on data uploaded by the edge computing layer and distributes the new model parameters to the edge computing nodes, enabling continuous learning and improvement of the system.
[0122] The above embodiments are only for illustrating the technical concept and features of the present invention. Their purpose is to enable people familiar with this technology to understand the content of the present invention and implement it. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made according to the spirit of the present invention should be included in the scope of protection of the present invention.
Claims
1. An adaptive production scheduling method based on multi-dimensional environmental perception, characterized in that: The steps include: S1. Collection of production environment data: real-time collection of workshop environment parameters, equipment status data, order data, and material status data; S2. Construction of an equipment efficiency decay model: Based on the LSTM time series prediction algorithm, an equipment efficiency decay model is constructed to analyze the relationship between production environment data and equipment efficiency. S3. Construction of a multi-objective optimization model: The multi-objective optimization model includes optimization objectives and constraints. The optimization objectives include maximizing equipment utilization, minimizing order delay rate, and minimizing energy consumption cost. The constraints include equipment capacity constraints, process sequence constraints, resource constraints, and delivery time constraints. S4. Production scheduling plan generation: Based on the multi-objective optimization model and the equipment efficiency decay model, the NSGA-II algorithm is used to generate the production scheduling plan; S5. Dynamic response to production environment changes: Based on the multi-objective optimization model and equipment efficiency attenuation model, when the production environment data change is lower than the preset threshold, the DRL algorithm is used to quickly adjust the production scheduling plan; when the production environment data change exceeds the preset threshold, the MCTS algorithm is used to generate an emergency scheduling plan.
2. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 1 is characterized in that: The equipment efficiency decay model in step S2 includes an input layer, an LSTM layer and a fully connected layer. The input layer receives time series data of workshop environmental parameters and equipment status data. The LSTM layer contains multiple LSTM units for capturing dynamic changes in time series data. The fully connected layer maps the output of the LSTM layer to the predicted value of the efficiency decay coefficient.
3. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 2 is characterized in that: The LSTM unit includes a forget gate, an input gate, and an output gate. The forget gate determines how much long-term information to retain and which long-term information to forget. The input gate determines how much new information to absorb into the long-term memory and updates the unit state used to control the long-term memory. The output gate filters out the short-term information that is most suitable for the current time step from the new long-term information.
4. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 3 is characterized in that: The forget gate uses the following formula to determine how much long-term information to retain: f t =σ(W f ·[h t-1 ,x t ]+b f ) Where, f t is the ratio of the forget gate output, f t ∈[0,1]; σ is the sigmoid function; W f is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b f is the offset matrix; t is the time step; The input gate uses the following formula to determine how much new information to incorporate into long-term memory: Where, To integrate new information into long-term memory; is the total new information absorbed in the current time step; i t is the ratio of screening new information; σ is the sigmoid function; W t 、W c is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b t 、b c is the offset matrix; The input gate updates the cell state used to control long-term memory using the following formula: Where C t is the long-term information of the current time step; f t is the ratio of the forget gate output; C t-1 is the long-term information of the previous time step; To integrate new information into long-term memory; The output gate uses the following formula to filter out the short-term information that is most suitable for the current time step from the new long-term information: Where h t To filter out the short-term information that is most suitable for the current time step from the new long-term information; t is the proportion of screening new information; C t is the long-term information of the current time step; σ is the sigmoid function; W o is a parameter that dynamically affects the final weight; h t-1 is the short-term information of the previous time step; x t is the information of the current time step; b o is the offset matrix.
5. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 1 is characterized in that: The specific steps of using the NSGA-II algorithm to generate a production scheduling plan in step S4 are as follows: S41. Initialize the population: Randomly generate an initial population, where each individual represents a production schedule, and calculate the objective function value for each individual. S42. Non-dominated sorting and crowding calculation: Perform non-dominated sorting on the initial population and divide the population into different non-dominated levels; Calculate the crowding degree for each individual in the non-dominated hierarchy; S43. Selection, crossover, and mutation operations: Select individuals from the current population to generate the next generation population; exchange partial gene fragments of two parent individuals to generate new offspring individuals; and randomly change certain genes of individuals. S44. Determine whether the termination condition is met: If the termination condition is met, output the Pareto optimal solution set in the current population; Otherwise, execute step S45; S45. Population merging: All individuals of the parent and offspring generations are mixed to form a larger population, and then return to step S42 for iteration.
6. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 1 is characterized in that: The specific steps of using the DRL algorithm to quickly adjust the production scheduling plan in step S5 are as follows: S51. Define the state space, action space, and reward function: The state space includes workshop environment parameters, equipment status data, order data, and material status data. The action space includes equipment start and stop, production batch splitting, line change sequence adjustment, equipment parameter adjustment, and order priority adjustment. The reward function formula is as follows: Reward=α×OEE+β×OnTimeDelivery-γ×(η×EnergyCost) Where Reward is the reward value obtained at each time step; OEE is an indicator to measure equipment utilization; OnTimeDelivery is a measure of the proportion of orders completed on time; EnergyCost is a measure of energy consumption in the production process; η is the equipment efficiency attenuation factor calculated based on the workshop environmental parameters; α, β, and γ are weight coefficients used to balance the importance of different objectives; S52. Use the improved DDPG+TD3 algorithm as the basic framework: The basic framework includes the actor network, critic network, target network, and buffer. Initialize the actor network, critic network, target network, and buffer, and set hyperparameters. S53. Introducing hierarchical reinforcement learning: High-level controllers with longer action cycles are responsible for production line-level scheduling, while low-level controllers with shorter action cycles are responsible for equipment parameter adjustment. The high-level and low-level controllers work together. S54. Introducing multi-agent collaboration: Each device has an independent agent responsible for local decision-making. The central critic network coordinates the decisions of each agent to obtain the globally optimal production scheduling strategy. S55. Use historical production data to train the iterative DRL algorithm: Use the experience replay mechanism to randomly sample historical production environment data from the buffer for training until the set termination condition is reached, and continuously update.
7. The adaptive production scheduling method based on multi-dimensional environmental perception according to claim 1 is characterized in that: The specific steps of using the MCTS algorithm to generate an emergency production scheduling plan in step S5 are as follows: S51' initializes the root node of the tree, the root node is the starting point of the current production state, contains all the data of the current production environment; S52′. Select branch nodes according to the UCB1 formula until reaching an incompletely expanded branch node. Select an unexplored child node on the incompletely expanded branch node for expansion. Generate a new state based on the current state and possible decision actions, and add it to the tree as a new branch node. S53′. Starting from the expanded node, simulate the production situation in the future until the termination state is reached; S54'. Back propagate the simulation results back to the root node, update the information of each node on the path, and optimize the decision path; S55 '. Loop steps S52 '-S54 ', evaluating various possible decision paths; S56′. Select the optimal decision path and generate an emergency production scheduling plan.
8. An adaptive production scheduling system based on multi-dimensional environmental perception, characterized in that: An adaptive production scheduling method based on multi-dimensional environmental perception as described in any one of claims 1-7 is adopted, including a sensor layer, an edge computing layer, a cloud platform layer and an execution terminal; the sensor layer collects workshop environmental parameters, equipment status data, order data and material status data in real time, transmits the collected data to the edge computing layer in real time, and uploads it to the cloud platform layer at the same time; the edge computing layer generates a production scheduling plan based on the received data, generates scheduling instructions based on the production scheduling plan and sends them to the execution terminal in a timely manner, and uploads a summary of key data to the cloud platform layer at the same time; the cloud platform layer is used for data storage, global optimization, model iteration and result display.
9. The adaptive production scheduling system based on multi-dimensional environmental perception according to claim 8, characterized in that: The sensor layer includes temperature and humidity sensors, vibration sensors, smart meters and RFID tags.
10. The adaptive production scheduling system based on multi-dimensional environmental perception according to claim 8, characterized in that: The cloud platform layer includes a database, a model training module, a visualization dashboard, and a federated learning center.