Dynamic production scheduling optimization method based on reinforcement learning and ERP (Enterprise Resource Planning) integrated system

Through a dynamic production scheduling optimization method based on reinforcement learning, combined with edge computing and ERP integrated system, the problem of insufficient flexibility and adaptability of traditional manufacturing scheduling methods under multiple processes and multiple constraints is solved, and efficient and flexible production scheduling optimization is achieved.

CN120258252AInactive Publication Date: 2025-07-04SHENZHEN NUOFEI TECH DEV CO LTD +1
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510739552.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the dynamic scheduling problem under multi-process and multi-constraint conditions, the traditional methods are insufficient in flexibility and adaptability. Especially when facing real-time changing production environments, existing reinforcement learning technologies still face challenges when dealing with high-dimensional state space and real-time requirements.

Method used

The dynamic production scheduling optimization method based on reinforcement learning is adopted, and the production process state is quantized, convolutional neural network and spatial attention mechanism are used to extract feature, combined with edge computing and lightweight framework deployment, the edge device operation of the computing task is realized, and the relationship between different strategies is analyzed through game theory models, and the ERP integrated system is added to optimize the scheduling rules.

Benefits of technology

It improves the intelligence level of production processes, meets the PLC-level soft real-time requirements, balances delay, energy consumption and safety goals, realizes the flexibility and adaptability of production scheduling, and reduces production cycle fluctuations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258252A_ABST
    Figure CN120258252A_ABST
Patent Text Reader

Abstract

The invention discloses a reinforcement learning-based dynamic production scheduling optimization method and an ERP (Enterprise Resource Planning) integrated system, relates to the technical field of intelligent manufacturing and information technologies, and is used for solving the problem that the rapid development of the modern manufacturing industry puts forward higher requirements on production scheduling optimization. The flexibility and the adaptability are insufficient in the face of a real-time changing production environment; the intelligent level of the production process is improved through the multi-dimensional data processing capacity, the multi-dimensional quantization coding technology is adopted, dimension raising or dimension reduction processing is carried out based on the income-cost ratio of cross-domain technical parameters, and it is ensured that a calculation task is transferred to edge equipment and the PLC-level soft real-time requirement is met through edge calculation and lightweight framework deployment. On-time process completion and delay penalty are integrated, a global optimization item is combined, a meta-learning algorithm is used for dynamically adjusting the weight, and balance of delay, energy consumption and a safety target is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent manufacturing and information technology. More specifically, the present invention relates to a dynamic production scheduling optimization method based on reinforcement learning and an ERP integration system. Background Art

[0002] The rapid development of modern manufacturing has put forward higher requirements for production scheduling optimization. Especially in the dynamic scheduling problems under multiple processes and multiple constraint conditions, traditional methods gradually show certain limitations. The scheduling rules in enterprise resource planning systems are usually based on predefined static logics. Although they can meet the basic production requirements to a certain extent, their flexibility and adaptability are insufficient when facing the real-time changing production environment. In recent years, the application of reinforcement learning technology in the field of dynamic scheduling has gradually received attention, but existing methods still face challenges in dealing with high-dimensional state spaces and real-time requirements.

[0003] In view of the above problems, the present invention proposes a solution. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a dynamic production scheduling optimization method based on reinforcement learning and an ERP integration system to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions: A dynamic production scheduling optimization method based on reinforcement learning, comprising the steps of: Step S1: Quantify and encode the production process state, splice the state quantification encoding and the overall progress index into a two-dimensional matrix as the state input, and use a convolutional neural network combined with a spatial attention mechanism to extract features therefrom, and output multiple priority rules as the action space; Step S2: Considering the inference speed and resource consumption of the model, use edge computing or lightweight framework deployment to transfer the computing tasks to edge devices; Step S3: Combine the on-time completion of processes and delay penalties to obtain a basic reward and set a reward function in combination with a global optimization term to reduce the production cycle; Step S4: In the case of coexistence of multiple priority rules, use a game theory model to analyze the mutual relationships and their benefits among different strategies to find the optimal solution, and add an ERP integration system to extract predefined scheduling rules from the ERP system.

[0006] In a preferred embodiment, the following steps are included: Quantitatively encode the status data of the production process, including multi-dimensional features such as the overall progress index, carbon emission factor, material completeness confidence, and fatigue index, and splice them with the overall progress index into a two-dimensional matrix as the initial state input; Introduce cross-domain technical parameters φ = {carbon emission factor c, material completeness confidence q, fatigue index f}, and calculate the benefit-cost ratio of each domain in real time: ; where Measure the partial derivative of the objective function with respect to the cross-domain parameter, that is, the information benefit, Represents the computational amount corresponding to this parameter, that is, the usage cost, Is the exponentially weighted term for drift rate estimation; According to the benefit-cost ratio Perform dimensionality increase or decrease processing: If Is higher than the preset threshold, then stack the corresponding dimension as an independent channel onto the time dimension to generate a 3D or 4D tensor; otherwise, keep the two-dimensional matrix form and scale the weights through a broadcast operation; When the benefit-cost ratio Is higher than the threshold, trigger the dimensionality increase module to generate a high-dimensional tensor. For the dimensions that do not meet the dimensionality increase conditions, keep the two-dimensional matrix form and scale the weights through a broadcast operation; Use the spatial attention mechanism to calculate the attention weights at each position, generate the attention weight map A, and multiply it with the original feature map F to obtain the weighted feature map F'; Flatten the feature map F' and input it into the fully connected layer, map it to the dimensional space related to the number of priority rules, and output the probability distribution of the priority rules as the action space.

[0007] In a preferred embodiment, it includes the following steps: Step B1: Implement hierarchical monitoring through eBPF monitoring extension, monitor resource consumption for different computing layers, including memory bandwidth occupancy rate and network I / O latency; Step B2: Adopt model reversible sharding and reinforcement learning decision-making, and dynamically optimize the migration decision according to the current load, historical load trend, network bandwidth, and cloud resource availability; Step B3: When the increase in computational amount caused by the dimensionality increase of cross-domain parameters exceeds 30% or the load exceeds the threshold, trigger the model reversible sharding module, and hot migrate the attention layer and the subsequent fully connected layer to the cloud WebAssembly sandbox, and keep the front-end convolutional layer running locally; Step B4: Collaboratively optimize through the teacher-student distillation module, channel pruning module, and INT8 quantization module to ensure that the delay during the migration process does not exceed 35ms and meets the PLC-level soft real-time requirements; Remove redundant structures through the pruning module, transfer the knowledge of the teacher model to the student model through the distillation module, and reduce the model's computational load through the INT8 quantization module.

[0008] In a preferred embodiment, the following steps are included: Design on-time completion rewards and delay penalties : ; where is an indicator function, is the process weight, is the actual completion time of the i-th process, is the expected completion time of the i-th process; Introduce a global optimization term to reduce production cycle fluctuations: ; where is the standard deviation of the production cycle, reflecting scheduling volatility; is the average completion time; is the stability adjustment coefficient; Construct a multi-domain coupling reward function: ; where , , are weight coefficients, adaptively adjusted through the meta-learning algorithm; Instantly re-weight the gradients of the new dimensions in the 3D / 4D form to ensure the balance of delay, energy consumption, and safety goals.

[0009] In a preferred embodiment, the following steps are included: Step D1: Treat each ERP rule as an agent, construct a game theory model, design a revenue function that includes local and global goals, and dynamically adjust the resource competition relationship between agents through the conflict weight; Step D2: Represent the ERP rules through a decision tree and convert them into a neural network model for optimization through backpropagation; Step D3: Input real-time data using the OPC standard protocol, dynamically adjust the rule trigger conditions, and use moving window averaging or a prediction model to fill in missing data; Step D4: Use the multi-agent deep deterministic policy gradient algorithm for potential game optimization, and limit the extreme choices of agent strategies through the potential function; Step D4-1: Design a potential function as a regularization term to adjust the penalty intensity of cross-domain costs; Step D4-2: Update agent behavior through policy gradients to avoid local optima; Step D4-3: Jointly calculate rewards and update policy gradients during training to ensure the authenticity of the simulation environment; Step D5: Conduct multi-round game training in the simulation environment, and use the OPC-UA-driven dynamic closed-loop mechanism to update parameters in real time and feedback the execution effect.

[0010] Step D5-1: Dynamically adjust the agent parameters using real-time data, including equipment status and production progress; Step D5-2: Collect execution data to evaluate the policy effect and feedback it to the policy network to optimize the parameters.

[0011] The present invention discloses a dynamic production scheduling optimization ERP integration system based on reinforcement learning, including an intelligent decision-making module, an edge lightweight module, a dynamic reward and punishment optimization module, and a game collaboration module, with signal connections between the modules.

[0012] Intelligent decision-making module: Quantify and encode the production process status, splice the status quantification encoding and the overall progress index into a two-dimensional matrix as the status input, and use a convolutional neural network combined with a spatial attention mechanism to extract features from it, and output multiple priority rules as the action space; Edge lightweight module: Considering the inference speed and resource consumption of the model, use edge computing or lightweight frameworks for deployment, and transfer the computing tasks to edge devices; Dynamic reward and punishment optimization module: Combine the on-time completion of the process and the delay penalty to obtain the basic reward and set the reward function in combination with the global optimization item to reduce the production cycle; Game collaboration module: In the case of multiple coexisting priority rules, use game theory models to analyze the mutual relationships and their benefits between different strategies to find the optimal solution, add an ERP integration system, and extract predefined scheduling rules from the ERP system.

[0013] Technical effects and advantages of the dynamic production scheduling optimization method and ERP integration system based on reinforcement learning of the present invention: Through the multi-dimensional data processing ability, the intelligent level of the production process is improved. The multi-dimensional quantization encoding technology is adopted to perform dimensionality increase or decrease processing based on the benefit-cost ratio of cross-domain technical parameters. Through edge computing and lightweight framework deployment, it is ensured that the computing tasks are transferred to edge devices and meet the PLC-level soft real-time requirements. The on-time completion and delay penalty of the process are integrated, combined with the global optimization item and the meta-learning algorithm is used to dynamically adjust the weights to ensure the balance of delay, energy consumption, and safety goals. Through seamless docking with the ERP system, it is convenient for enterprises to optimize production scheduling using existing data and rules. Description of the Drawings

[0014] Figure 1 It is a schematic structural diagram of the dynamic production scheduling optimization method and ERP integration system based on reinforcement learning of the present invention.

[0015] Figure 2 This is the flowchart of the two-dimensional matrix convolution output of the present invention.

[0016] Figure 3 This is the implementation flowchart of the dynamic production scheduling optimization method combining the ERP integration system and reinforcement learning of the present invention.

[0017] Figure 4 This is the flowchart of the ERP integration system for dynamic production scheduling optimization based on reinforcement learning of the present invention. Detailed implementation manners

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0019] Embodiment 1 The present invention discloses a dynamic production scheduling optimization method and an ERP integration system based on reinforcement learning, including the steps of: Step S1: Quantitatively encode the production process state, splice the state quantitative encoding and the overall progress index into a two-dimensional matrix as the state input, and use a convolutional neural network combined with a spatial attention mechanism to perform feature extraction on it, and output multiple priority rules as the action space; Step S2: Considering the inference speed and resource consumption of the model, use edge computing or lightweight framework deployment to transfer the computing task to the edge device; Step S3: Combine the on-time completion of the process and the delay penalty to obtain the basic reward, and set the reward function in combination with the global optimization item to reduce the production cycle; Step S4: In the case of multiple coexisting priority rules, use a game theory model to analyze the mutual relationship and benefits between different strategies to find the optimal solution, and add an ERP integration system to extract predefined scheduling rules from the ERP system.

[0020] In step S1, the production process state is quantitatively encoded, the state quantitative encoding and the overall progress index are spliced into a two-dimensional matrix as the state input, and a convolutional neural network combined with a spatial attention mechanism is used to perform feature extraction on it, and multiple priority rules are output as the action space. The specific content includes: Quantify and encode the status data of the production process, including multi-dimensional features such as the overall progress index, carbon emission factor, material availability confidence, and fatigue index, and splice them with the overall progress index into a two-dimensional matrix as the initial state input. The two-dimensional matrix is used to represent the basic status information of the current production environment, facilitating subsequent feature extraction operations. The data format of the two-dimensional matrix input forms a unified two-dimensional matrix form after status quantization encoding and serves as the initial input of the convolutional neural network. The core function of the convolutional neural network is to extract spatial correlation and time-dependence features; On the premise of maintaining the "two-dimensional matrix" as the benchmark input form, introduce cross-domain technical parameters φ = {carbon emission factor c, material availability q, fatigue index f...}. The system calculates the benefit-cost ratio of each domain in real time according to the following formula: ; where Measures the partial derivative of the objective function with respect to the cross-domain parameter, that is, the information benefit, Represents the computational amount corresponding to this parameter, that is, the usage cost, Is the exponentially weighted term for drift rate estimation; Dimensionality increase and decrease feature data processing mechanism: Dimensionality increase: When the benefit-cost ratio of certain dimensions Is higher than the preset threshold, the dimensionality increase processing module is triggered, and the corresponding dimension is stacked as an independent channel to the time dimension to generate a 3D or 4D tensor, so as to quickly capture trend changes in emergency scenarios. For example, if the benefit-cost ratio of the tool life parameter is higher than the threshold, the information of the tool life dimension is used as a new channel and stacked with the original two-dimensional matrix in the time dimension to form a 3D tensor, which can better capture the change trend of this dimension in the time series and ensure that trend changes can be quickly captured in emergency scenarios; Dimensionality decrease: For dimensions that do not meet the dimensionality increase conditions, they are still represented in the form of a two-dimensional matrix, and the weights are scaled through a "broadcast" operation. This design not only retains the advantages of the lightweight 2D-CNN but also enhances the representation ability of multi-dimensional features, can dynamically adjust the input dimensions according to real-time data, and adapt to complex and changeable production environments. Overall, this solution aims to achieve effective monitoring and analysis of the production process status through reasonable data processing and model design to cope with various changes in the production process.

[0021] On the processed feature data, use the spatial attention mechanism to calculate the attention weight of each position, and generate an attention weight map by weighted summation of different channels and spatial positions of the feature map through two fully connected layers. For example, for a feature map F, first map the channel dimension to an intermediate dimension through a fully connected layer, and then map it back to the same channel dimension as F through another fully connected layer to obtain the attention weight map A; Feature weighting: Multiply the attention weight map A by the original feature map F to obtain the weighted feature map F′, enabling the model to pay more attention to the features in the spatial positions and channels that are important for the current task, suppressing unimportant features, and thus improving the effectiveness of feature extraction; Fully connected layer mapping: Flatten the feature map F′ processed by the spatial attention mechanism. If it is a 3D or 4D tensor, perform appropriate dimensionality reduction operations first and then flatten it, and then input it into one or more fully connected layers. The fully connected layer maps the features to a suitable dimensional space, and the size of this dimension corresponds to the number of priority rules, enabling the model to focus more specifically on different spatial positions and channels and identify the features that are most important for the current task; The fully connected layer maps the features to a dimensional space related to the number of priority rules, and finally outputs a probability distribution, indicating the likelihood of each priority rule being selected: Apply a suitable activation on the output of the fully connected layer to convert the output values into a probability distribution form. Each probability value corresponds to the weight of a priority rule, and these probability values constitute the action space, representing the likelihood of different priority rules being selected, so that corresponding production decisions, such as equipment scheduling, parameter adjustment, etc., can be made according to these priority rules.

[0022] In step S2, considering the inference speed and resource consumption of the model, use edge computing or lightweight frameworks for deployment, and transfer the computing tasks to edge devices. The specific content includes: Establish a resource-dimension coordination mechanism. Considering the inference speed and resource consumption of the model, use edge computing or lightweight frameworks for deployment, and transfer the computing tasks to edge devices: Step B1: Extended implementation of eBPF monitoring. It is reasonable to increase the memory bandwidth occupancy rate and network I / O latency. Especially when the system involves cloud communication, network performance and latency directly affect the overall system performance. Adopting a hierarchical monitoring method to monitor resource consumption for different computing layers can provide more accurate information for subsequent sharding decisions. For deep neural networks, consider combining the computational volume and memory occupancy of each layer for multi-dimensional analysis; Step B2: Model reversible sharding and reinforcement learning decision-making. For the refinement of sharding granularity, focusing on the migration of individual sub-modules can significantly reduce the migration overhead, which helps to reduce the burden on edge devices and improve the overall efficiency. Factors such as the current load, historical load trend, network bandwidth, and cloud resource availability are considered. These inputs help the reinforcement learning model better evaluate how to optimize the migration decision. Actions such as the migration layer, migrating to cloud nodes, and compressing the model need to consider the real-time requirements of the system and resource utilization efficiency. The reinforcement learning model can make dynamic adjustments based on real-time feedback; Step B3: Automatic migration mechanism. In addition to load shedding, the strategy of using time window judgment is very meaningful. It can reduce frequent migrations caused by short-term fluctuations. It is necessary to ensure that the migration trigger conditions are highly robust to prevent excessive migrations due to short-term load fluctuations. Predicting future load trends based on historical load patterns and migrating in advance is a very effective optimization method. Combined with time series analysis, it can help the system to schedule loads more intelligently and reduce unnecessary migration operations. When the cross-domain parameter dimension increase causes the computational volume to increase by more than 30% or the load exceeds the threshold, the model reversible sharding module is triggered, and the online reinforcement learning decision will hot migrate the attention layer and the subsequent fully connected layer to the cloud WebAssembly sandbox, and the front-end convolutional layer will remain running locally.

[0023] Step B4: Model optimization technology works together. The teacher-student distillation module, channel pruning module and INT8 quantization module work together to ensure that the delay of the migration process does not exceed 35ms and meets the PLC-level soft real-time requirements. The loss of accuracy after quantization is an issue that needs to be balanced, especially in critical tasks, where the loss of accuracy may affect the stability of the system. The loss of accuracy can be mitigated by using a high-quality calibration data set. In addition, dynamic quantization technology can also reduce quantization errors. Pruning first and then distilling is a more effective strategy because pruning can reduce redundant structures and make the information in the distillation process more concentrated. The pruned model is more concise and more conducive to the optimization of the distillation process.

[0024] In step S3, the basic reward is obtained by combining the on-time completion of the process with the delay penalty, and the reward function is set in combination with the global optimization item to reduce the production cycle. The specific contents include: Combining timing constraints, energy consumption, and carbon emissions, we design a multi-domain coupled reward function to evaluate the performance of the scheduling strategy, comprehensively consider different objectives, and balance them in an adaptive way to optimize system performance: Balanced design of basic rewards and penalties: (Reward for on-time completion): It is essentially an efficiency incentive, directly related to the achievement rate of the production plan, and is achieved through the following formula: ;in, is the indicator function, is the process weight, is the actual completion time of the ith process, is the expected completion time of the i-th process; (Delay Penalty): Adoption Exponential Growth ;in, is the penalty intensity coefficient, which avoids treating small delays the same as large delays; Global Optimization : ; among them, is the standard deviation of the production cycle, reflecting the scheduling volatility; is the average completion time; is the stability adjustment coefficient. If the of multiple processes fluctuates greatly, will increase, thereby reducing the reward value, prompting the scheduling strategy to pursue a more stable production cycle, suppressing the oscillation behavior of the scheduling strategy through variance minimization, and enhancing long-term stability; Multi-domain coupling reward function: ; among them, is the basic reward for completing the process on time, reflecting the efficiency of task completion on time; is the delay penalty, reflecting the negative impact brought by the task not being completed on time; is the cost penalty for carbon emissions. The higher the carbon emissions, the greater the penalty, prompting the system to minimize the carbon footprint as much as possible; is the cost penalty for energy consumption. Saving energy costs will be considered when optimizing the scheduling strategy; is the global optimization term for reducing the production cycle fluctuation, ensuring that the production process is more stable and reducing large-scale cycle changes; Adaptive adjustment: Through the meta-learning algorithm, automatically adjust the parameter β in the reward function according to the current system state and requirements, so as to balance each goal in different situations, making the reward mechanism flexible and adaptable in a dynamic environment, and being able to adjust the weight of each goal according to the actual situation. The update rule is: ; among them, is the learning rate, is the gradient calculated by the policy gradient method; Reward function adjustment in 3D / 4D form: When the system expands to a higher dimension, the reward function will instantaneously re-weight the gradients of the new dimensions. This ensures that in the new dimension, the three major goals of delay, energy consumption, and safety can always be comparable, making the performance evaluation of the system in multiple dimensions still reasonable and effective.

[0025] In step S4, in the case of multiple priority rules coexisting, use the game theory model to analyze the mutual relationship and benefits between different strategies to find the optimal solution, add the ERP integration system, and extract the predefined scheduling rules from the ERP system. The specific content includes: Step D1: Game theory model construction Regard each ERP rule as an agent. The behavior strategy of the agent depends on the state, and the state includes environmental factors and historical decisions. The revenue function is designed as: Local objective: Optimization of the rules themselves, such as reducing delivery delays. Global objective: Optimization of system collaboration efficiency, such as improving resource utilization. Conflict weight design: Considering resource competition among multiple agents (ERP rules), design conflict weights to measure the competition relationship between agents. The conflict weights will be dynamically adjusted according to the production cycle or environmental changes to ensure reasonable resource allocation. Use the Lagrange multiplier method to transform the constraints into penalty terms, thus incorporating global constraints into the objective function. The global objective is defined by the constraint function; Step D2: Fusion of ERP rules and real-time data Represent the ERP rules through a decision tree. The decision tree nodes represent conditional judgments, and the leaves represent actions, intuitively transforming the rules into a decision model. Then transform the decision tree into a neural network model and optimize it through backpropagation; Adopt the OPC standard protocol to input real-time data into the system and dynamically adjust the rule trigger conditions. If data delays or missing occur, a sliding window average or prediction model can be used to fill in the missing data to ensure the robustness of the system; Step D3: MADDPG - Potential game optimization Select the multi-agent deep deterministic policy gradient algorithm for its adaptability to continuous action spaces and partially observable environments. In the environment of multi-agent collaborative games, update the agents' behaviors through policy gradients to avoid local optimum problems.

[0026] Function of the potential function: The potential function, as a regularization term, can limit the extreme choices of agents' strategies, avoiding only pursuing local optima while ignoring the global objective. By adjusting the penalty intensity of the cross-domain cost, the dynamic balance between rules can be achieved, and thus global optimization can be reached.

[0027] Training process: Conduct multiple rounds of game training in the simulation environment to accumulate agents' experience and improve the decision-making quality of the system. The mechanism design of joint reward calculation and policy gradient update during training is reasonable, and pay attention to the authenticity of the simulation environment to accurately simulate the complexity in the real production scenario.

[0028] Step D4: OPC-UA-driven dynamic closed-loop Dynamic parameter update mechanism: Utilize real-time data, including equipment status and production progress, to dynamically adjust the parameters of agents, ensuring the flexible response of the system.

[0029] Execution feedback closed-loop: Collect execution data to evaluate the execution effect of the strategy, and adjust the parameters of the policy network according to the feedback data to further optimize the system performance.

[0030] The present invention discloses an ERP integration system for optimizing dynamic production scheduling based on reinforcement learning, which includes an intelligent decision-making module, an edge lightweight module, a dynamic reward and punishment optimization module, and a game collaboration module, and the modules are connected by signals.

[0031] Intelligent decision-making module: Quantitatively encode the status of production processes, splice the status quantitative encoding and the overall progress index into a two-dimensional matrix as the status input, and use a convolutional neural network combined with a spatial attention mechanism to extract features from it, and output multiple priority rules as the action space; Edge lightweight module: Considering the inference speed and resource consumption of the model, use edge computing or lightweight frameworks for deployment, and transfer the computing tasks to edge devices; Dynamic reward and punishment optimization module: Combine the on-time completion of processes and delay penalties to obtain a basic reward and set a reward function in combination with global optimization items to reduce the production cycle; Game collaboration module: In the case of multiple coexisting priority rules, use game theory models to analyze the mutual relationships and their benefits between different strategies to find the optimal solution, add an ERP integration system, and extract predefined scheduling rules from the ERP system.

[0032] The above formulas are all dimensionless and take their numerical calculations. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.

[0033] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product.

[0034] Those of ordinary skill in the art can realize that the modules and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application of the technical solution and the invention constraints. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0035] In addition, in each embodiment of the present application, the various functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.

[0036] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims described above.

[0037] Finally: The above description is only the preferred embodiment of the present invention and is not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A dynamic production scheduling optimization method based on reinforcement learning, characterized in that Including the steps: Step S1: Quantitatively encode the status of the production process, splice the status quantitative encoding and the overall progress index into a two-dimensional matrix as the status input, and use a convolutional neural network combined with a spatial attention mechanism to extract features from it, and output multiple priority rules as the action space; Step S2: Considering the inference speed and resource consumption of the model, use edge computing or lightweight frameworks for deployment, and transfer the computing tasks to edge devices; Step S3: Combine the on-time completion of the process and the delay penalty to obtain the basic reward, and set the reward function in combination with the global optimization item to reduce the production cycle; Step S4: In the case of multiple coexisting priority rules, use a game theory model to analyze the mutual relationship and benefits between different strategies to find the optimal solution, add an ERP integration system, and extract predefined scheduling rules from the ERP system.

2. The dynamic production scheduling optimization method based on reinforcement learning according to claim 1, characterized in that: Quantitatively encode the status data of the production process, including multi-dimensional features such as the overall progress index, carbon emission factor, material completeness confidence, and fatigue index, and splice it with the overall progress index into a two-dimensional matrix as the initial state input; Introduce cross-domain technical parameters φ = {carbon emission factor c, material matching confidence q, fatigue index f}, and calculate the benefit-cost ratio of each domain in real time: ; among them, The partial derivative of the objective function with respect to the cross-domain parameter is the information benefit, It represents the computational amount corresponding to this parameter, that is, the usage cost, It is the exponentially weighted term for drift rate estimation; According to the benefit-cost ratio Perform dimensionality increase or decrease processing: If is higher than the preset threshold, stack the corresponding dimension as an independent channel onto the time dimension to generate a 3D or 4D tensor; otherwise, maintain the two-dimensional matrix form and scale the weights through a broadcast operation; Use the spatial attention mechanism to calculate the attention weight of each position, generate the attention weight map A, and multiply it with the original feature map F to obtain the weighted feature map F'; Flatten the feature map F' and input it into the fully connected layer, map it to the dimensional space related to the number of priority rules, and output the probability distribution of the priority rules as the action space.

3. The dynamic production scheduling optimization method based on reinforcement learning according to claim 2, wherein The specific implementation of the dimension increase and decrease processing includes: Step A3-1: When the benefit-cost ratio is higher than the threshold, trigger the dimensionality increase module to generate a high-dimensional tensor; Step A3-2: For the dimensions that do not meet the dimension increase conditions, keep the two-dimensional matrix form and scale the weights through broadcast operations.

4. The dynamic production scheduling optimization method based on reinforcement learning according to claim 1, characterized in that: Step B1: Implement hierarchical monitoring through eBPF monitoring extension, monitor resource consumption for different computing layers, including memory bandwidth occupancy rate and network I / O latency; Step B2: Adopt model reversible sharding and reinforcement learning decision-making, and dynamically optimize the migration decision according to the current load, historical load trend, network bandwidth, and cloud resource availability; Step B3: When the increase in the amount of computation caused by cross-domain parameter dimension increase exceeds 30% or the load exceeds the threshold, trigger the model reversible sharding module, and hot migrate the attention layer and the subsequent fully connected layers to the cloud WebAssembly sandbox, and keep the front-end convolutional layer running locally.

5. The dynamic production scheduling optimization method based on reinforcement learning according to claim 4, characterized in that: Step B4: Collaboratively optimize through the teacher-student distillation module, channel pruning module, and INT8 quantization module to ensure that the delay during the migration process does not exceed 35ms and meets the PLC-level soft real-time requirements; Remove redundant structures through the pruning module, transfer the knowledge of the teacher model to the student model through the distillation module, and reduce the model computation amount through the INT8 quantization module.

6. The dynamic production scheduling optimization method based on reinforcement learning according to claim 1, characterized in that: Design completion-on-time reward and delay penalty : ; where is an indicator function, is the process weight, is the actual completion time of the i-th process, is the expected completion time of the i-th process; Introduce global optimization items Reduce production cycle fluctuations: ; where is the standard deviation of the production cycle, reflecting scheduling volatility; is the average completion time; is the stability adjustment coefficient; Construct a multi-domain coupling reward function: ; where , , are weight coefficients, which are adaptively adjusted through the meta-learning algorithm; Perform immediate reweighting of the gradients of the new dimension in 3D / 4D form to ensure a balance among latency, energy consumption, and safety goals.

7. The dynamic production scheduling optimization method based on reinforcement learning according to claim 1, characterized in that: Step D1: Regard each ERP rule as an agent, construct a game theory model, design a revenue function including local goals and global goals, and dynamically adjust the resource competition relationship among agents through conflict weights; Step D2: Represent the ERP rules through a decision tree and transform them into a neural network model, and optimize through backpropagation; Step D3: Input real-time data using the OPC standard protocol, dynamically adjust the rule trigger conditions, and use a sliding window average or a prediction model to fill in missing data; Step D4: Use the multi-agent deep deterministic policy gradient algorithm for potential game optimization, and limit the extreme choices of agent strategies through a potential function; Step D5: Conduct multiple rounds of game training in a simulation environment, and use the OPC-UA-driven dynamic closed-loop mechanism to update parameters in real time and feedback the execution effect.

8. The dynamic production scheduling optimization method based on reinforcement learning according to claim 7, wherein The specific implementation of the potential game optimization in the said Step D4 includes: Step D4-1: Design a potential function as a regularization term to adjust the penalty intensity of cross-domain costs; Step D4-2: Update the agent behavior through policy gradients to avoid local optima; Step D4-3: Jointly calculate rewards and update policy gradients during training to ensure the authenticity of the simulation environment.

9. The dynamic production scheduling optimization method based on reinforcement learning according to claim 8, wherein, The specific implementation of the dynamic closed-loop mechanism in the said Step D5 includes: Step D5-1: Dynamically adjust agent parameters using real-time data, including device status and production progress; Step D5-2: Collect execution data to evaluate the policy effect and feedback it to the policy network to optimize parameters.

10. A dynamic production scheduling optimization ERP integration system based on reinforcement learning, which is used to implement the dynamic production scheduling optimization method based on reinforcement learning according to any one of claims 1-9, characterized in that : Intelligent decision-making module: Quantitatively encode the status of production processes, splice the status quantitative encoding and the overall progress indicator into a two-dimensional matrix as the status input, and use a convolutional neural network combined with a spatial attention mechanism to extract features from it, and output multiple priority rules as the action space; Edge lightweight module: Considering the inference speed and resource consumption of the model, use edge computing or a lightweight framework for deployment, and transfer the computing tasks to edge devices; Dynamic reward and punishment optimization module: Combine the on-time completion of processes and delay penalties to obtain a basic reward and set a reward function in combination with a global optimization term to reduce the production cycle; Game cooperation module: In the case of multiple priority rules coexisting, use a game theory model to analyze the mutual relationship and their benefits among different strategies to find the optimal solution, and add an ERP integration system to extract predefined scheduling rules from the ERP system.

Citation Information

Cited By

  • Automatic collaborative decision-making method, system and equipment for enterprise resource planning, medium and product

    CN120494762A

  • Automated collaborative decision-making method, system, device, medium and product for enterprise resource planning

    CN120494762B

  • Multi-objective collaborative scheduling optimization method and system for intelligent manufacturing workshop

    CN121010143A

  • Self-adaptive strategy pruning method and device for balancing production efficiency

    CN121030762A

  • Data acquisition and remote logic control method for multi-protocol equipment

    CN121887872A