Workshop multi-agent deep reinforcement learning scheduling method based on real-time production condition

Through distributed sensor networks and multi-agent deep reinforcement learning, combined with virtual twin simulation environment, the multi-objective collaboration and data perception problems of the workshop production scheduling system under dynamic production conditions are solved, real-time scheduling optimization and solution robustness are achieved.

CN120370867AActive Publication Date: 2025-07-25ZHEJIANG YUEXIN PRINTING & DYEING CO LTD

Patent Information

Application Number
CN202510507802.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-25
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

When facing dynamic production conditions such as equipment failures, order priority adjustments and material supply fluctuations, the existing workshop production scheduling system is insufficient in dynamic adaptability, the multi-target coordination efficiency is low, and it is unable to effectively balance energy consumption economy, time efficiency and abnormal fault tolerance. Intelligent algorithms are prone to fall into local optimality in multi-agent collaboration scenarios, lack full-dimensional data perception and closed-loop feedback mechanisms, the fusion efficiency of multi-source heterogeneous data is low, and abnormal events are difficult to quantify and model, resulting in lack of robustness of the scheduling scheme.

Method used

A distributed sensor network is used to collect multi-dimensional production data in real time, build a three-dimensional decision space, and coordinate scheduling with multi-agent deep deterministic strategy gradient algorithm, combine it with a virtual twin simulation environment to perform multi-objective verification, establish real-time urgency and collaborative efficiency factors, optimize the scheduling scheme through improved course learning strategies, and conduct online parameter updates.

Benefits of technology

Real-time perception and dynamic response of the workshop production scheduling system are realized, and time efficiency, energy consumption economy and abnormal fault tolerance are adaptively balanced, resource coordination efficiency and scheduling solutions are improved, and the lag and offline verification problems of traditional scheduling systems are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120370867A_ABST
    Figure CN120370867A_ABST
Patent Text Reader

Abstract

The invention provides a workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions, and relates to the technical field of production scheduling and industrial automation, and the method comprises the steps: collecting production data in real time through a distributed sensor network, and constructing a multi-dimensional time sequence matrix; establishing a three-dimensional decision space containing time, resource and task dimensions, and adopting an improved multi-agent depth deterministic strategy gradient algorithm; dynamically calculating the real-time production urgency degree of the production batch, and analyzing cross-unit cooperation to generate a cooperation efficiency factor; the two are combined to optimize a scheduling scheme through a course learning strategy; and verifying the scheme by using a virtual twin environment and feeding back and updating model parameters. The method breaks through the limitations of insufficient dynamic adaptability, low efficiency of multi-target coordination and the like of traditional scheduling, realizes real-time response to dynamic production conditions such as equipment faults and order adjustment, balances time efficiency, energy consumption economy and abnormal fault tolerance through a multi-agent coordination and closed-loop verification mechanism, and improves resource coordination efficiency under complex production constraints.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of production scheduling and industrial automation, and particularly to a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions. Background Art

[0002] With the acceleration of the intelligent transformation of the manufacturing industry, workshop production scheduling, as the core link of the manufacturing system, faces increasingly complex dynamic environments and multi-objective optimization requirements. Traditional scheduling methods are mostly based on static rules or offline optimization models, and it is difficult to adapt to real-time changing production conditions such as sudden equipment failures, order priority adjustments, and material supply fluctuations. Especially in the scenario of multi-variety and small-batch customized production, where the process connection is highly intensive and resource competition is fierce, existing scheduling systems often lead to problems such as decreased production efficiency, increased energy consumption, and the spread of abnormal events due to lagging response or rigid decision-making, and there is an urgent need to break through the inherent limitations of the static scheduling mode.

[0003] Current mainstream scheduling technologies have significant deficiencies in aspects such as dynamic adaptability, multi-objective coordination, practicality of intelligent algorithms, and real-time data fusion. Strategies based on fixed rules (such as first-come, first-served) or periodic rescheduling cannot real-time perceive dynamic factors such as equipment status and environmental parameters, resulting in the disconnection between the scheduling plan and the actual situation. For example, it is difficult to adjust the processing sequence in a timely manner to avoid quality defects when the equipment temperature is abnormal, and it is also impossible to quickly reconstruct the path optimization plan when the material flow is delayed; most adopt single-objective optimization (such as the shortest completion time) or simple weighted summation methods, and it is difficult to balance the conflicts between energy consumption economy, time efficiency, and abnormal fault tolerance, and may cause chain problems or production capacity waste due to excessive pursuit of a certain goal; intelligent algorithms such as reinforcement learning are prone to falling into local optima in multi-agent cooperation scenarios, and rely on simulation environments with a lack of high-fidelity modeling capabilities for training, and the strategy effect significantly decreases when migrated to the actual workshop; most systems lack a full-dimensional data perception system and an online feedback mechanism, and cannot achieve the closed-loop iteration of "decision-making - execution - verification - optimization", resulting in the disconnection between decision-making and the actual production process.

[0004] Although the industry has tried to improve the dynamic response ability of the scheduling system through technologies such as Internet of Things perception and digital twin, it still faces multiple challenges: the real-time fusion and feature extraction efficiency of multi-source heterogeneous data (such as vibration spectra, temperature curves, and logistics trajectories) is low, the learning stability and interpretability of multi-agent cooperation strategies under complex production constraints are insufficient, the randomness and propagation effects of abnormal events are difficult to quantify and model, and the multi-objective trade-off in a dynamic environment relies on manual experience and lacks an adaptive optimization mechanism. In this context, there is an urgent need for a scheduling method that deeply integrates real-time perception, intelligent decision-making, and closed-loop verification to break through the limitations of traditional technologies and achieve the efficient coordination of workshop resources and the dynamic optimization of the production process. Summary of the Invention

[0005] To solve the technical problems in the prior art, such as insufficient dynamic adaptability in workshop production scheduling, difficulty in dealing with sudden equipment failures, order priority adjustments, material supply fluctuations and other dynamic production conditions in real time, low multi-objective collaborative efficiency, inability to effectively balance the conflicts among energy consumption economy, time efficiency and exception tolerance, the intelligent algorithm is prone to fall into local optimum in the multi-agent collaboration scenario and relies on a simulation environment lacking high-fidelity modeling ability, resulting in poor policy migration effect, lack of real-time data fusion and decision-making closed-loop, lack of a full-dimensional data perception system covering equipment, materials and environment and a "decision-execution-validation-optimization" closed-loop mechanism, low efficiency of real-time fusion and feature extraction of multi-source heterogeneous data, insufficient stability and interpretability of multi-agent collaborative strategy learning under complex production constraints, difficulty in quantifying and modeling the randomness and propagation effect of abnormal events, resulting in lack of robustness of the scheduling scheme, and the multi-objective trade-off in the dynamic environment relying on manual experience and lacking an adaptive optimization mechanism, the present invention provides a workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions.

[0006] The technical solution provided by the present invention is as follows:

[0007] A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions provided by the present invention includes:

[0008] S1. Real-time production data collection: Real-time collect the processing status data of each production unit through a distributed sensor network, including equipment operation parameters, material flow progress and environmental monitoring indicators, and construct a multi-dimensional time series data matrix;

[0009] S2. Dynamic scheduling decision modeling: Establish a three-dimensional decision space including time dimension, resource dimension and task dimension, and adopt a multi-agent deep deterministic policy gradient algorithm to cooperate with the scheduling model, and each agent corresponds to an independent decision-making unit;

[0010] S3. Real-time production urgency calculation: Dynamically calculate the real-time production urgency of each production batch based on three dimensions of process connection tightness, remaining processing time margin and equipment load balance;

[0011] S4. Generation of collaborative efficiency factor: Generate a multi-dimensional collaborative efficiency factor by analyzing the cross-unit material flow path, equipment cooperation working mode and abnormal event propagation path;

[0012] S5. Online scheduling decision optimization: Take the real-time production urgency and collaborative efficiency factor as state input features, dynamically adjust the agent exploration rate through an improved curriculum learning strategy, and output a joint scheduling plan including process sorting, equipment allocation and material path;

[0013] S6. Scheduling Scheme Verification and Feedback: Establish a virtual twin simulation environment to conduct multi-objective verification on the generated scheduling scheme. The verification metrics include energy consumption economy, time efficiency, and exception fault tolerance. Feed the verification results back to the scheduling model for online parameter update.

[0014] The beneficial effects brought by the technical solution provided by the present invention at least include:

[0015] (1) In the present invention, multi-dimensional production data is collected in real time through a distributed sensor network, a three-dimensional decision-making space including time, resources, and tasks is constructed, and combined with the dynamic calculation of real-time production urgency and collaborative efficiency factor, the scheduling system can perceive the changes in equipment status, material flow, and environmental parameters in real time, effectively respond to dynamic production conditions such as sudden equipment failures and order priority adjustments, and break through the lag limitation of traditional static scheduling;

[0016] (2) In the present invention, an improved multi-agent deep reinforcement learning algorithm is adopted, an attention mechanism and a variance constraint term are introduced to optimize the agent cooperation strategy, and the exploration rate is dynamically adjusted in combination with the curriculum learning strategy, which can achieve an adaptive balance among time efficiency, energy consumption economy, and exception fault tolerance, avoid the production capacity waste or the risk of chain failures caused by single-objective optimization, and improve the resource collaboration efficiency under complex production constraints;

[0017] (3) In the present invention, a virtual twin simulation environment is established to conduct multi-objective verification on the scheduling scheme. Through the comprehensive evaluation of indicators such as energy consumption economy, time efficiency, and exception fault tolerance and Pareto front analysis, a closed-loop feedback mechanism of "decision-making - execution - verification - optimization" is formed, solving the problem of the disconnection between the offline verification of the traditional scheduling system and actual production, realizing the online iterative optimization of the scheduling model, and significantly improving the robustness and long-term adaptability of the scheme. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions provided by an embodiment of the present invention;

[0020] Figure 2 It is a schematic flowchart of a construction method for a three-dimensional decision-making space of a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions provided by an embodiment of the present invention;

[0021] Figure 3 Schematic flow chart of the calculation method of the real-time production urgency P of a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions provided by an embodiment of the present invention;

[0022] Figure 4 Schematic flow chart of the calculation method of the collaborative efficiency factor CE of a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions provided by an embodiment of the present invention. Detailed implementation manners

[0023] The technical solutions in the present invention will be described below with reference to the accompanying drawings.

[0024] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as an "example" in the present invention should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of the word "example" is intended to present concepts in a specific manner. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one of the two.

[0025] In the embodiments of the present invention, "image" and "picture" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same. "(of)", "corresponding", and "corresponding" can sometimes be used interchangeably. It should be noted that when the difference is not emphasized, the meanings they express are the same.

[0026] In the embodiments of the present invention, sometimes subscripts such as W1 may be miswritten as non-subscript forms such as W1. When the difference is not emphasized, the meanings they express are the same.

[0027] To make the technical problems, technical solutions and advantages to be solved by the present invention clearer, the following will be described in detail with reference to the accompanying drawings and specific embodiments.

[0028] Refer to the attached specification Figure 1 , which shows a schematic flow chart of a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions provided by an embodiment of the present invention.

[0029] The embodiments of the present invention provide a multi-agent deep reinforcement learning scheduling method for a workshop based on real-time production conditions. The processing flow may include the following steps:

[0030] S1. Real-time production data collection: Real-time collect the processing status data of each production unit through a distributed sensor network, including equipment operation parameters, material flow progress and environmental monitoring indicators, and construct a multi-dimensional time series data matrix.

[0031] It should be noted that in this step, a multi-dimensional time series data matrix is constructed through a distributed sensor network, and its core function is to achieve a comprehensive situation awareness of the production site. The adoption of a distributed architecture avoids the risk of single-point failure. The introduction of environmental monitoring indicators can capture parameters such as temperature and humidity that have implicit impacts on processing quality, providing multi-physical field coupling data support for subsequent decision-making. Compared with traditional single-dimensional data acquisition, this method can effectively retain the dynamic correlation characteristics in the production process through a time-series data organization method, laying a data foundation for building an accurate scheduling decision model.

[0032] S2. Dynamic scheduling decision-making modeling: Establish a three-dimensional decision space including the time dimension, resource dimension, and task dimension, and adopt a multi-agent deep deterministic policy gradient algorithm (MADDPG) collaborative scheduling model, where each agent corresponds to an independent decision-making unit.

[0033] It should be noted that in this step, a three-dimensional decision space is constructed, and the three dimensions of time, resources, and tasks are orthogonally decomposed. The sliding window mechanism in the time dimension realizes the organic combination of historical experience and future prediction. The equipment capacity vector in the resource dimension incorporates the maintenance demand index and temperature influence coefficient into the quantitative evaluation system for the first time. The feature matrix in the task dimension breaks through the limitation of traditional scheduling methods that only consider the delivery date. Using the MADDPG algorithm to construct a collaborative scheduling model effectively solves the contradiction problem of equipment resource competition and cooperation through a multi-agent architecture, providing an algorithm framework for distributed decision-making in a complex production environment.

[0034] In a possible implementation manner, as Figure 2 shown, the construction method of the three-dimensional decision space in S2 includes:

[0035] S201. The time dimension adopts a sliding window mechanism, intercepting historical data for a period of Δt before the current moment as a reference, and predicting the scheduling requirements for the period of t + Δt backward;

[0036] S202. The resource dimension establishes an equipment capacity vector E = (e1, e2,..., e m ), where e i = α·U i + β·M i + γ·T i , U i is the real-time utilization rate of equipment i, M i is the maintenance demand index, T i is the temperature influence coefficient, and α, β, and γ are dynamic weight parameters;

[0037] S203. Construct the feature matrix F = [f ij n×k , where f ij represents the quantization value of task j in the i-th feature dimension. The feature dimensions include process complexity, quality requirements, and delivery date sensitivity.

[0038] It can be understood that the model interpretability is improved by structuring and decomposing decision-making elements. The time dimension sliding window mechanism (S201) combines historical data analysis and future demand prediction to achieve dynamic time modeling. The adaptive adjustment mechanism of the window length Δt can match different production rhythm requirements. The resource dimension equipment capacity vector (S202) breaks through the traditional single evaluation mode of utilization rate, and introduces a maintenance requirement index and a temperature influence coefficient. The maintenance requirement index integrates the equipment vibration spectrum characteristics and maintenance cycle data, and the temperature influence coefficient correlates the material thermal expansion coefficient with the machining accuracy requirements. The task dimension feature matrix (S203) quantifies the processing difficulty level through process complexity, maps the detection standard strictness in the quality requirement dimension, and introduces the customer level weight in the delivery date sensitivity, forming a task portrait system facing multiple constraint conditions, providing a structured input for subsequent agent decision-making.

[0039] It should be noted that the dynamic weight parameters in S202 specifically include:

[0040] The determination methods of the dynamic weight parameters α, β, and γ include:

[0041] S501. Construct the judgment matrix J = [j xy 3×3 , where j xy represents the importance degree of the x-th index relative to the y-th index;

[0042] S502. Calculate the information entropy k is the normalization coefficient;

[0043] S503. Calculate the difference coefficient G x = 1 - E x ;

[0044] S504. The final weight

[0045] ​​It is understandable that the traditional method is improved to achieve data-driven dynamic weight allocation. In the construction of the judgment matrix (S501), an improved three-scale method (1-3-5 grading) is adopted to automatically generate the relative importance scores between indicators through the device operation log, eliminating subjective judgment biases. In the information entropy calculation (S502), Laplace smoothing is introduced to avoid the zero-probability problem, and the normalization coefficient k = 1 / ln(3) ensures a reasonable entropy value range. The difference coefficient transformation (S503) quantifies the information uncertainty as the basis for weight allocation, and finally the normalization process (S504) guarantees the mathematical completeness of the weight parameters. This method enables the weight ratios of equipment utilization rate, maintenance requirements, and temperature impact to reflect the characteristics of the current production environment in real time, automatically increasing the weight of the temperature coefficient in the high-temperature season and strengthening the weight of maintenance requirements in the equipment aging stage.

[0046] It should be noted that the improvement of the multi-agent deep deterministic policy gradient algorithm MADDPG algorithm in S2 includes:

[0047] S801. Add an attention mechanism to the critic network to calculate the correlation weights between agents:

[0048]

[0049] Among them, Q i is the query vector of the current agent, K j is the key vector of other agents, and d is the vector dimension;

[0050] S802. Adopt a double-delay update strategy, with the main network parameter update period Δτ = 1000 steps and the target network update period Δτ' = 2000 steps;

[0051] S803. Introduce a variance constraint term in the policy gradient: t is the current training step, and T max is the maximum training step.

[0052] It is understandable that the multi-agent cooperation efficiency is improved through algorithm-level innovation. The attention mechanism (S801) adopts a multi-head attention structure (4 heads), and the d = 64-dimensional vector space ensures feature decoupling. The Q and K matrices are generated by splicing the device type encoding and the real-time state. The double-delay update strategy (S802) sets the main network update step size η = 0.001 and the target network soft update coefficient τ = 0.005 to balance the learning speed and stability. The variance constraint term (S803) introduces the calculation of the moving average variance, with the window size set to 100 experience samples, and the β coefficient decays linearly with the training progress, encouraging exploration in the initial stage (β = 0.1) and focusing on exploitation in the later stage (β = 0.01).

[0053] S3. Real-time production urgency calculation: Based on three dimensions, namely the tightness of process connection, the margin of remaining processing time, and the equipment load balance, dynamically calculate the real-time production urgency of each production batch.

[0054] It should be noted that this step creatively proposes a three-dimensional urgency evaluation system to solve the problem of rigid priority setting in traditional scheduling methods. The tightness of process connection quantifies the coupling relationship in the production process. The margin of remaining processing time introduces an exponential decay function to enhance time sensitivity. The equipment load balance uses the standard deviation measurement to avoid local overload. Through the adaptive adjustment of dynamic weight coefficients, the dynamic balance between emergency order response and system stable operation is achieved. The calculation of this parameter provides a quantitative decision basis for subsequent scheduling optimization, enabling the system to have the elastic ability to handle emergencies.

[0055] In a possible implementation, as Figure 3 shown, the calculation method of the real-time production urgency P includes:

[0056] S301. Calculate the time urgency component: where T r is the remaining processing time, T a is the maximum allowable delay time, T h is the process half-life;

[0057] S302. Calculate the coupling degree component: C d is the coupling degree between the current process and the downstream process, C t is the threshold coupling degree;

[0058] S303. Calculate the load balance component: P1 = ω3·σ(S1), where S1 is the standard deviation of the load of similar equipment, x0 is the reference load difference value;

[0059] S304. Synthesize the total urgency: P = P t + P c + P1, where ω1, ω2, ω3 are dynamic weight coefficients, satisfying ω1 + ω2 + ω3 = 1 and ω i > 0.

[0060] It can be understood that an itemized weighted model is established to solve the problem of multi-source parameter fusion. The time urgency component (S301) uses an exponential decay function to simulate the non-linear change of process urgency. The process half-life T hThe setting enables different process types (such as heat treatment and machining) to exhibit differentiated attenuation curves. The coupling degree component (S302) eliminates the inherent coupling differences between processes through threshold normalization processing, and the logarithmic function compresses the influence of extreme values to avoid decision-making oscillations. The load balancing component (S303) innovatively uses the relative standard deviation for measurement, and the reference load difference value x0 is dynamically calculated according to the production capacity of the equipment group to eliminate the evaluation bias caused by equipment heterogeneity. The dynamic weight coefficient (S304) introduces a hidden Markov model to track the production status in real time, automatically increases the time weight during the peak order period, and strengthens the load balancing weight during the equipment failure period to achieve intelligent weight migration.

[0061] S4. Generation of collaborative efficiency factor: By analyzing the cross-unit material flow path, equipment collaborative working mode, and abnormal event propagation path, a multi-dimensional collaborative efficiency factor is generated.

[0062] It should be noted that in this step, aiming at the problem of cross-unit collaborative scheduling, a multi-dimensional evaluation model integrating path optimization, resource collaboration, and abnormal propagation is proposed. The path optimization potential value quantifies the improvement space of material flow efficiency, the resource collaboration benefit reveals the improvement potential of equipment utilization rate, and the abnormal influence factor evaluates the risk conduction effect through the probability product model. The min-max normalization strategy is used to handle the fusion problem of parameters with different dimensions, and the dynamic adjustment mechanism of the balance coefficient η enables the system to automatically focus on the efficiency or stability goal at different production stages, significantly improving the overall collaborative efficiency of complex production networks.

[0063] In a possible implementation manner, as Figure 4 shown, the calculation method of the collaborative efficiency factor CE in S4 includes:

[0064] S401. Calculate the path optimization potential value: where is the original path length, is the optimized path length;

[0065] S402. Calculate the resource collaboration benefit: is the theoretical maximum utilization rate of the equipment, is the current utilization rate;

[0066] S403. Calculate the abnormal influence factor: ρ k is the occurrence probability of the k-th type of abnormal event, and A k is the influence coefficient corresponding to the abnormal;

[0067] S404. Generate the final collaborative efficiency factor: CE = η·A + (1 - η)·min{μ1·L p , μ2·R e}, where T cis the current cumulative processing time, T t is the time threshold, μ1 and μ2 are benefit weights and μ1 + μ2 = 1.

[0068] It can be understood that a multi - path collaborative evaluation model is constructed to break through the efficiency - robustness trade - off dilemma. The path optimization potential value (S401) is measured by the relative improvement rate to avoid the dimensional difference of different transportation vehicles. Among them, the optimized path length is generated by the hybrid strategy of A* algorithm and tabu search. The resource collaborative benefit (S402) introduces the theoretical maximum utilization rate whose value is dynamically corrected according to the historical peak of the equipment OEE (Overall Equipment Effectiveness). The abnormal influence factor (S403) is constructed through fault tree analysis ρ k and A k quantitative relationship, and the product formula accurately depicts the risk of multi - abnormal chain reaction. The final synthesis formula (S404) uses the min function to constrain the collaborative short - board effect, and the sigmoid function design of the balance coefficient η enables the system to automatically focus on abnormal prevention when the cumulative processing time is close to the threshold, forming an environment - adaptive collaborative optimization mechanism.

[0069] S5. Online scheduling decision optimization: Taking the real - time production urgency and collaborative efficiency factor as state input features, dynamically adjust the agent exploration rate through an improved curriculum learning strategy, and output a joint scheduling plan including process sequencing, equipment allocation, and material path.

[0070] It should be noted that this step realizes progressive agent training through an improved curriculum learning strategy, solving the problem of low exploration efficiency in traditional reinforcement learning. The high exploration rate in the initial stage ensures the learning of basic scheduling strategies. The sine function fluctuation introduced in the dynamic adjustment formula prevents local optimality, and the exponential decay mechanism enhances the strategy stability in the later training. Taking the urgency and collaborative factor as joint state features enables the agent to simultaneously focus on task urgency and system coordination. The three - dimensional scheduling plan (process, equipment, path) output forms a complete decision - making closed - loop, ensuring the executability and optimality of the plan.

[0071] In a possible implementation, the realization of the curriculum learning strategy in S5 includes:

[0072] S601. Set the exploration rate ε0 = 0.9 in the initial stage and preferentially learn single - equipment scheduling tasks;

[0073] S602. When the scheduling accuracy rate of the model on the validation set reaches 85% for 3 consecutive epochs, introduce equipment collaborative constraint conditions and adjust the exploration rate to t is the current training step, and T is the cycle parameter;

[0074] S603. When the device utilization rate reaches the 75% threshold, activate the exception handling module and at the same time reduce the exploration rate to ε2 = 0.5·(1 - e -0.01t ).

[0075] It can be understood that a progressive training framework is designed to solve the complex scheduling learning problem. In the initial stage (S601), the learning difficulty is reduced by simplifying the state space (only including the basic parameters of the device), and the high exploration rate of 0.9 ensures full traversal of the single-device scheduling strategy. The accuracy trigger mechanism (S602) introduces early stopping to prevent overfitting, and the device cooperation constraint conditions gradually increase the material buffer limit and energy consumption threshold. The sine wave term in the dynamic formula of the exploration rate breaks the strategy convergence deadlock, and the setting of the cycle parameter T = 5000 steps matches the rhythm of typical production batches. The activation condition of the exception handling module (S603) is associated with the device's MTBF (Mean Time Between Failures) data, and the exponential decay function makes the exploration rate decline rate negatively correlated with the device reliability, realizing a risk-adaptive exploration strategy.

[0076] S6. Scheduling plan verification and feedback: Establish a virtual twin simulation environment to conduct multi-objective verification on the generated scheduling plan. The verification metrics include energy consumption economy, time efficiency, and exception tolerance. Feed the verification results back to the scheduling model for online parameter update.

[0077] It should be noted that the virtual twin simulation environment constructed in this step breaks through the limitations of traditional offline verification and realizes the comprehensive evaluation of the plan through a multi-objective verification system. The energy consumption economy metric is associated with the device power and running time, the time efficiency metric uses a normalized tardiness penalty function, and the exception tolerance metric quantifies the system robustness through the weighted frequency reciprocal. The Pareto front analysis in the three-dimensional evaluation space ensures the non-dominance of the plan, and the online parameter update mechanism forms a reinforcement learning closed-loop of "decision-making - verification - optimization", significantly improving the continuous optimization ability of the scheduling system.

[0078] In a possible implementation, the specific method for multi-objective verification in S6 includes:

[0079] S701. Calculate the energy consumption economy metric: P i is the power of device i, t i is the running time, P max is the maximum allowable power consumption of the system;

[0080] S702. Calculate the time efficiency metric: d j is the delivery date, a j is the actual completion time;

[0081] S703. Calculate the abnormal fault tolerance index: w k is the weight of the k-th type of anomaly, and f k is the occurrence frequency;

[0082] S704. Establish a three-dimensional evaluation space and screen the non-dominated solution set through Pareto front analysis.

[0083] It can be understood that the construction of a quantitative evaluation system solves the multi-attribute optimization problem. The energy consumption economy index (S701) introduces a power factor correction coefficient, P max which is dynamically adjusted according to the workshop distribution capacity to avoid overload risks. The time efficiency index (S702) adopts a piecewise penalty function, sets a 3-fold penalty weight for key customer orders, and strengthens the VIP guarantee ability. The abnormal fault tolerance index (S703) determines w k weights through Failure Mode and Effects Analysis (FMEA), and the frequency statistics window is set to a rolling 24-hour cycle. The three-dimensional evaluation space (S704) uses the NSGA-II algorithm for Pareto front search, sets the crowding distance threshold to 0.1 to maintain the diversity of the solution set, and provides a strategy pool containing 5-8 non-dominated solutions for decision-makers.

[0084] It should be noted that the construction method of the virtual twin simulation environment in S6 includes:

[0085] S901. Establish a digital model of the equipment, import historical processing data to train the LSTM prediction network, and predict the equipment state transition probability;

[0086] S902. Construct a directed graph of material flow G=(V, E), where the node V represents the processing site, the edge E represents the transmission path, and the edge weight includes the transmission time and energy consumption cost;

[0087] S903. Design an abnormal event injection module, generate equipment failure events according to the Weibull distribution, and generate material shortage events according to the Poisson process;

[0088] S904. Set the simulation acceleration factor α sim ∈[1, 10], and dynamically adjust the simulation speed according to the GPU computing power.

[0089] It is understandable that a high-fidelity simulation environment is created for closed-loop verification. The device digital model (S901) integrates a physics engine to simulate mechanical transmission errors, and the input of the LSTM network includes 12-dimensional features such as current waveforms and noise spectra. The edge weights of the material flow diagram (S902) are calculated by integrating the time cost (0.5 yuan / minute) and the energy consumption cost (1.2 yuan / kWh), and the path optimization is constrained by a maximum turning angle of 45 degrees. In the abnormal event injection (S903), the shape parameter β of the Weibull distribution is 2.1 to simulate wear failures, and the intensity λ of the Poisson process is 0.02 times / minute to fit material shortages. The simulation acceleration (S904) uses time scaling technology and enables double-precision floating-point operations when the GPU memory is sufficient to ensure numerical stability.

[0090] In a possible implementation manner, it further includes a dynamic parameter adjustment method:

[0091] S1001. Real-time monitor the device temperature volatility T avg is the historical average temperature;

[0092] S1002. When ΔT > 0.15, trigger an emergency cooling strategy and insert a forced maintenance time window into the scheduling plan;

[0093] S1003. Dynamically adjust the process priority weights: ω base is the basic weight;

[0094] S1004. Update the time threshold in the calculation of the collaborative efficiency factor t is the cumulative running time.

[0095] It is understandable that an environment-responsive adaptive mechanism is constructed. The temperature monitoring (S1001) uses a sliding Z-score algorithm to detect abnormal fluctuations, with a window size N = 30 minutes and a threshold Z = 2.5. The forced maintenance strategy (S1002) inserts a 15-minute cooling time window, during which a standby device is started to take over the tasks. The priority adjustment (S1003) sets the basic weight ω base = 0.7, and the temperature compensation coefficient is limited in the interval [0.7, 1.3] to prevent the weights from getting out of control. The time threshold decay (S1004) introduces a learning rate decay mechanism, reducing the initial threshold by 5% every 1000 steps, but setting a minimum threshold to avoid excessive compression of the time margin.

[0096] The beneficial effects brought by the technical solution provided by the embodiment of the present invention at least include:

[0097] (1) In the present invention, multi-dimensional production data is collected in real time through a distributed sensor network, a three-dimensional decision space including time, resources, and tasks is constructed, and through the dynamic calculation of real-time production urgency and collaborative efficiency factors, the scheduling system can perceive in real time changes in equipment status, material flow, and environmental parameters, effectively respond to dynamic production conditions such as sudden equipment failures and order priority adjustments, and break through the lag limitation of traditional static scheduling;

[0098] (2) In the present invention, an improved multi-agent deep reinforcement learning algorithm is adopted, an attention mechanism and a variance constraint term are introduced to optimize the agent collaboration strategy, and the exploration rate is dynamically adjusted in combination with the curriculum learning strategy, which can achieve an adaptive balance among time efficiency, energy consumption economy, and anomaly fault tolerance, avoid production capacity waste or chain failure risks caused by single-objective optimization, and improve the resource collaboration efficiency under complex production constraints;

[0099] (3) In the present invention, a virtual twin simulation environment is established to conduct multi-objective verification on the scheduling scheme. Through comprehensive evaluation of indicators such as energy consumption economy, time efficiency, and anomaly fault tolerance and Pareto frontier analysis, a closed-loop feedback mechanism of "decision-making - execution - verification - optimization" is formed, solving the problem of the disconnection between offline verification of traditional scheduling systems and actual production, realizing online iterative optimization of the scheduling model, and significantly improving the robustness and long-term adaptability of the scheme.

[0100] The above content is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

[0101] The following points need to be explained:

[0102] (1) The attached drawings of the embodiments of the present invention only relate to the structures involved in the embodiments of the present invention, and other structures can refer to the general design.

[0103] (2) For clarity, in the attached drawings used to describe the embodiments of the present invention, the thickness of layers or regions is enlarged or reduced, that is, these drawings are not drawn according to the actual ratio. It can be understood that when an element such as a layer, film, region, or substrate is referred to as being "on" or "under" another element, the element can be "directly" on or under another element or there can be intermediate elements.

[0104] (3) Without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other to obtain new embodiments.

[0105] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions, characterized in that, Including: S1. Real-time production data collection: The processing status data of each production unit is collected in real time through a distributed sensor network, including equipment operation parameters, material flow progress, and environmental monitoring indicators, to construct a multi-dimensional time series data matrix; S2. Dynamic scheduling decision-making modeling: A three-dimensional decision-making space including time dimension, resource dimension, and task dimension is established, and a multi-agent deep deterministic policy gradient algorithm is used to coordinate the scheduling model, with each agent corresponding to an independent decision-making unit; S3. Real-time production urgency calculation: Based on three dimensions of process connection tightness, remaining processing time margin, and equipment load balance, the real-time production urgency of each production batch is dynamically calculated; S4. Collaborative efficiency factor generation: By analyzing cross-unit material flow paths, equipment collaborative working modes, and abnormal event propagation paths, multi-dimensional collaborative efficiency factors are generated; S5. Online scheduling decision optimization: The real-time production urgency and collaborative efficiency factors are used as state input features, and the agent exploration rate is dynamically adjusted through an improved curriculum learning strategy, and a joint scheduling plan including process sequencing, equipment allocation, and material path is output; S6. Scheduling plan verification and feedback: A virtual twin simulation environment is established to conduct multi-objective verification on the generated scheduling plan. The verification indicators include energy consumption economy, time efficiency, and anomaly fault tolerance, and the verification results are fed back to the scheduling model for online parameter update.

2. The workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that The construction method of the three-dimensional decision-making space described in S2 includes: S201. The time dimension adopts a sliding window mechanism, intercepting historical data in a Δt time period forward based on the current moment, and predicting the scheduling requirements in a t+Δt time period backward; S202. Establish the device capability vector E = (e1, e2,..., e m ), where e i = α·U i + β·M i + γ·T i , U i is the real-time utilization rate of device i, M i is the maintenance requirement index, T i is the temperature influence coefficient, and α, β, and γ are dynamic weight parameters; S203. Construct a feature matrix F = [f ij n×k , where f ij represents the quantization value of task j in the i-th feature dimension, and the feature dimensions include process complexity, quality requirements, and delivery date sensitivity.​ 3. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that, The calculation method of the real-time production urgency P includes: S301. Calculate the time urgency component: where T r is the remaining processing time, T a is the maximum allowable delay time, T h is the half-life of the process; S302. Calculate the coupling degree component: C d is the coupling degree between the current process and the downstream process, and C t is the threshold coupling degree; S303. Calculate the load balancing component: P1 = ω3·σ(S1), where S1 is the standard deviation of the loads of the same type of devices. x0 is the reference load difference value. S304. Synthetic total urgency: P = P t + P c + P1, where ω1, ω2, ω3 are dynamic weight coefficients, satisfying ω1 + ω2 + ω3 = 1 and ω i > 0.

4. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that, The calculation method of the collaborative efficiency factor CE described in S4 includes: S401. Calculate the path optimization potential value: Where is the original path length, is the path length after optimization; S402. Calculate the collaborative benefit of computing resources: is the theoretical maximum utilization rate of the device, is the current utilization rate; S403. Calculate the abnormal influence factor: ρ k is the occurrence probability of the k-th type of abnormal event, and A k is the influence coefficient corresponding to the abnormality; S404. Generate the final collaborative efficiency factor: CE = η·A + (1 - η)·min{μ1·L p , μ2·R e}, where T c is the current cumulative processing time, T t is the time threshold, μ1 and μ2 are benefit weights and μ1 + μ2 = 1.

5. The workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 2, wherein, The dynamic weight parameters described in S202 specifically include: The determination method of the dynamic weight parameters α, β, γ includes: S501. Construct a judgment matrix \(J = [j_{xy}]\), xy 3×3 where \(j_{xy}\) xy represents the importance degree of the \(x\)th index relative to the \(y\)th index;​ S502. Calculate the information entropy k is a normalization coefficient; S503. Calculate the coefficient of variation G x = 1 - E x ; S504, Final Weight 6. The workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that The implementation of the curriculum learning strategy described in S5 includes: S601. In the initial stage, the exploration rate ε0 = 0.9 is set, and single-equipment scheduling tasks are preferentially learned; S602. After the scheduling accuracy rate of the model on the validation set reaches 85% for three consecutive epochs, introduce the device cooperation constraint condition and adjust the exploration rate to where t is the current training step and T is the cycle parameter; S603. When the device utilization rate reaches the 75% threshold, activate the exception handling module and at the same time reduce the exploration rate to ε2 = 0.5·(1 - e -0.01t ).

7. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that The specific method of multi-objective verification in S6 includes: S701. Calculate the energy consumption economy index: P i is the power of device i, and t i is the running time. P max is the maximum allowable power consumption of the system; S702. Calculate the time efficiency index: d j is the delivery date, and a j is the actual completion time; S703. Calculate the abnormal fault tolerance index: w k is the weight of the k-th type of anomaly, and f k is the occurrence frequency; S704. A three-dimensional evaluation space is established, and non-dominated solution sets are screened through Pareto front analysis.

8. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that The improvement of the MADDPG algorithm described in S2 includes: S801. An attention mechanism is added to the critic network to calculate the correlation weights between agents: Among them, Q i is the query vector of the current agent, and K j is the key vector of other agents, where d is the vector dimension; S802. A double-delay update strategy is adopted, with the main network parameter update period Δτ = 1000 steps and the target network update period Δτ' = 2000 steps; S803. Introduce a variance constraint term in the policy gradient: t is the current training step, and T max is the maximum training step.

9. The workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that The construction method of the virtual twin simulation environment in S6 includes: S901. Establish an equipment digital model, import historical processing data to train the LSTM prediction network, and predict the equipment state transition probability; S902. Construct a directed material flow graph G=(V,E), where the node V represents the processing site, the edge E represents the transmission path, and the edge weight includes transmission time and energy consumption cost; S903. Design an abnormal event injection module, generate equipment failure events according to the Weibull distribution, and generate material shortage events according to the Poisson process; S904. Set the simulation acceleration factor α sim ∈[1, 10], and dynamically adjust the simulation speed according to the GPU computing power.

10. A workshop multi-agent deep reinforcement learning scheduling method based on real-time production conditions according to claim 1, characterized in that It also includes a dynamic parameter adjustment method: S1001. Real-time monitoring of the temperature volatility of the device T avg is the historical average temperature; S1002. When ΔT > 0.15, trigger the emergency cooling strategy and insert a forced maintenance time window into the scheduling plan; S1003. Dynamically adjust the priority weight of the process: ω base is the basic weight; S1004. Update the time threshold in the calculation of the collaborative efficiency factor t is the cumulative running time.

Citation Information

Patent Citations

  • Multi-agent depth deterministic strategy gradient method based on course learning

    CN113449458A

  • Flexible job shop energy-saving scheduling method based on multi-agent architecture

    CN115509188A

  • Multi-agent path planning method based on distributed collaborative deep reinforcement learning model

    CN116225016A

  • Reconfigurable workshop dynamic scheduling method and system based on multi-agent near-end strategy optimization algorithm and efficient action decoding

    CN118780416A

  • Production and logistics cooperative scheduling method for dynamic flexible job shop

    CN119338156A

Cited By

  • Cooperative control method and system for die-casting pot production line based on data driving

    CN120821256A

  • A data-driven collaborative control method and system for die-casting pot production line

    CN120821256B

  • Big data-based production comprehensive supervision system

    CN120848430A

  • Multi-laser SLM layer vector data multi-source collaborative dynamic distribution method

    CN120875457A

  • A multi-laser SLM layer vector data multi-source collaborative dynamic allocation method

    CN120875457B