Adaptive Scheduling Control Method for Flexible Production Lines Integrating Multi-Agent Reinforcement Learning

By mapping the equipment in the flexible production line into an agent and establishing a local communication network, dynamically updating the state space and action space, the stability and robustness problems of the existing flexible production line scheduling methods in the face of dynamic disturbances are solved, efficient adaptive scheduling control is achieved, and the anti-interference ability and production continuity of the production line are improved.

CN120103803BActive Publication Date: 2025-07-18QINSILK COM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510472965.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-18
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

When the existing flexible production line scheduling control methods face dynamic disturbances such as equipment failures and material shortages, the stability and robustness of the scheduling scheme are poor, and the calculation complexity is high, making it difficult to meet the real-time decision-making needs, and lack of an adaptive mechanism, resulting in weak recovery ability of the production line under abnormal conditions.

Method used

Using the method of integrating multi-agent reinforcement learning, multiple processing equipment in the production line are mapped into agents, and a local communication network is established. Through information sharing and collaborative decision-making among agents, the equipment status is monitored in real time, the status space and action space are dynamically updated, and the compensation scheduling scheme is generated. Combined with the output results of the multi-agent reinforcement learning model, the adaptive scheduling control of the flexible production line is realized.

Benefits of technology

It improves the intelligent adaptive scheduling capabilities of the production line in dynamic and complex environments, enhances the anti-interference capability and the continuity of the production process, reduces production interruptions and resource waste, and improves the system's response speed and scheduling quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103803B_ABST
    Figure CN120103803B_ABST
Patent Text Reader

Abstract

The present invention provides a flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning, which relates to the technical field of production line scheduling. It includes obtaining the processing equipment status and workpiece demand information in real time, constructing the state and action spaces, mapping the equipment to agents and realizing information sharing, and then selecting and executing scheduling actions. The state-action value function is updated through an experience replay pool to optimize the multi-agent model. The system monitors the equipment status during operation, dynamically updates the state and action spaces to cope with failures or material shortages, and generates a compensation scheduling plan. Finally, combined with real-time scheduling instructions, the adaptive scheduling control of the flexible production line is realized, improving production efficiency and flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to production line scheduling technology, and in particular to a flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning. Background Art

[0002] As an important part of modern manufacturing, flexible production lines have the characteristics of product diversification, small production batches, and frequent switching, which pose higher requirements for production scheduling control systems. Traditional flexible production line scheduling methods mainly include heuristic algorithms, mathematical programming, and intelligent optimization algorithms, etc. These methods have realized the automated scheduling of the production process to a certain extent. In recent years, with the development of artificial intelligence technology, reinforcement learning has received extensive attention because it can continuously optimize the decision-making process by interacting with the environment. In particular, multi-agent reinforcement learning technology can better adapt to the characteristics of distributed equipment and task collaboration in flexible production lines.

[0003] However, there are still some obvious deficiencies in existing flexible production line scheduling control methods. First, most traditional scheduling algorithms are designed based on static optimization models and cannot effectively cope with dynamic disturbances in the production process, such as sudden situations like equipment failures and material shortages, resulting in poor stability and robustness of the scheduling scheme. Second, when dealing with large-scale production systems, existing reinforcement learning methods often adopt a centralized learning framework, with high computational complexity, difficult to meet real-time decision-making requirements, and a single agent is difficult to comprehensively perceive the state information of a complex production environment, affecting the accuracy of scheduling decisions. Third, existing technologies lack an effective adaptive mechanism. When the production environment changes, they cannot quickly adjust the strategy and generate a reasonable compensation plan, making it difficult to achieve the continuous and efficient operation of the production line, especially with weak recovery ability in abnormal situations.

[0004] By introducing a flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning, the above problems can be effectively solved, realizing the intelligent adaptive scheduling of the production line in a dynamic and complex environment, and improving the overall operation efficiency and anti-interference ability of the system. Summary of the Invention

[0005] The embodiments of the present invention provide a flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention,

[0007] A flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning is provided, including:

[0008] Obtain the real-time operating status of multiple processing devices and the processing requirement information of multiple workpieces in a flexible production line, and construct a state space and an action space according to the real-time operating status and the processing requirement information;

[0009] Map the multiple processing devices to corresponding agents, establish a local communication network between the agents to realize information sharing among adjacent agents, and the agents select and execute scheduling actions according to the current state information and the local observation information of neighboring agents;

[0010] After executing the scheduling action, obtain the system immediate reward value and store it in the experience replay pool, sample data from the experience replay pool to update the state-action value function, and obtain an optimized multi-agent reinforcement learning model through repeated iterative training;

[0011] During the operation of the multi-agent reinforcement learning model, continuously monitor the operating status of the multiple processing devices. When a device failure or material shortage is identified, dynamically update the state space and action space of the affected area, and use the updated state space and action space to recalculate the scheduling strategy to obtain a compensation scheduling plan;

[0012] Combine the compensation scheduling plan with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, so as to realize the adaptive scheduling control of the flexible production line.

[0013] Constructing a state space and an action space according to the real-time operating status and the processing requirement information includes:

[0014] The real-time operating status in the flexible production line includes device status, task queue and remaining processing time, and the processing requirement information includes the processing sequence and processing time of the workpieces;

[0015] Map the real-time operating status and the processing requirement information to the feature vector space, extract the spatial features in the feature vector space through a deep neural network, and dynamically construct the state space and the action space according to the spatial features. The state space represents the real-time state distribution of the processing devices and the workpieces, and the action space represents the set of executable scheduling decisions.

[0016] Mapping the multiple processing devices to corresponding agents, establishing a local communication network between the agents to realize information sharing among adjacent agents, and the agents select and execute scheduling actions according to the current state information and the local observation information of neighboring agents includes:

[0017] Map the multiple processing devices in the flexible production line to corresponding agents, and each agent extracts features from its own state information based on the attention mechanism to obtain a state feature vector reflecting the real-time working characteristics of the processing device;

[0018] Construct an agent communication network with a dynamic topology, model the degree of association between the agents based on the agent communication network, determine the communication weight by calculating the state similarity and task relevance between the agents, and the communication weight is adaptively adjusted with the changes in the agent state and task allocation to achieve the interaction of differential information between the agents;

[0019] Construct a hybrid policy network based on the state feature vector and the differential information. The hybrid policy network combines the deterministic policy with random exploration, outputs the deterministic action distribution through the hybrid policy network, and introduces random noise to achieve policy exploration. The training objective of the hybrid policy network is jointly determined by the task completion situation and the cooperation efficiency;

[0020] Select a scheduling action according to the output result of the hybrid policy network, and store the state transition information and reward information after executing the scheduling action in the experience pool.

[0021] Construct an agent communication network with a dynamic topology, model the degree of association between the agents based on the agent communication network, and determining the communication weight by calculating the state similarity and task relevance between the agents includes:

[0022] Construct multiple agents into a directed graph network structure. The nodes in the directed graph network structure represent agents, and the edges in the directed graph network structure represent communication links. For each agent, determine its corresponding set of neighborhood nodes, obtain the device working parameter vector, task queue state vector, and resource state vector of each agent, and combine the device working parameter vector, the task queue state vector, and the resource state vector to form the state vector of the agent;

[0023] Extract features from the state vector to obtain a feature vector, calculate the inner product of the feature vectors corresponding to any two agents, and obtain the connection strength after dividing the inner product by the square root of the feature dimension and passing through the softmax activation function;

[0024] Calculate the dot product of any two state vectors divided by the product of their respective vector norms to obtain the state similarity, calculate the cosine similarity between any two agents, and use it as the task relevance. Weight the state similarity and the task relevance to obtain the communication weight.

[0025] During the operation of the multi-agent reinforcement learning model, it continuously monitors the operating states of the multiple processing devices. When a device failure or material shortage is identified, it dynamically updates the state space and action space of the affected area, and recalculates the scheduling policy using the updated state space and action space to obtain a compensation scheduling plan, including:

[0026] Collect vibration data, temperature data, and current data of the processing devices to construct a multi-dimensional sensing data matrix. Extract features from the multi-dimensional sensing data matrix to obtain spatio-temporal features. Based on the spatio-temporal features, perform time series modeling through the multi-dimensional sensing data matrix to obtain a device state feature sequence. Obtain the material inventory data and material consumption data in the material management system, perform weighting to obtain a material supply prediction result, and generate a system anomaly situation vector based on the device state feature sequence and the material supply prediction result;

[0027] The system anomaly situation vector uses a directed acyclic graph to describe the propagation path of the anomaly event. Calculate the influence probability of the anomaly event on different production units according to the propagation path, compare the influence probability with a preset probability threshold to determine the affected area, and predict the state evolution trajectory during the anomaly propagation based on the affected area through probabilistic graph reasoning;

[0028] Based on the state evolution trajectory, retrieve the historical compensation plan for a similar scenario from the compensation strategy knowledge base, and use a meta-learning network to transfer the historical compensation plan to the current anomaly scenario. The meta-learning network includes a policy network and a value network. The policy network outputs the probability distribution of the compensation action, and the value network evaluates the long-term benefits of the compensation action. Select the candidate compensation plan with the highest benefit as the final compensation scheduling plan.

[0029] Based on the state evolution trajectory, retrieving the historical compensation plan for a similar scenario from the compensation strategy knowledge base includes:

[0030] Obtain the state data of the production equipment, where the state data includes equipment state parameters, process parameters, and constraint condition parameters, and splice the state data in the feature dimension to form a dimension feature vector;

[0031] Perform time series sampling on the dimension feature vector to obtain a sampling sequence. Arrange the dimension feature vectors in the sampling sequence in chronological order and splice them in the time dimension to form a time series feature matrix. The number of rows of the time series feature matrix corresponds to the dimension of the dimension feature vector, and the number of columns of the time series feature matrix corresponds to the number of sampling time points. Characterize the time series feature of the equipment state evolution through the time series feature matrix;

[0032] Construct a temporal decay function based on the temporal feature matrix. The temporal decay function uses a negative exponential form to assign different weight coefficients to the dimensional feature vectors at different times, and realizes the temporal weighting of the historical state information through the weight coefficients.

[0033] Calculate the similarity between the temporal feature matrix and the historical temporal feature matrix stored in the compensation strategy library, calculate the feature similarity between the dimensional feature vectors at the same time, and multiply the feature similarity by the weight coefficient at the corresponding time to obtain a trajectory similarity score; sort the historical compensation strategies in the compensation strategy library according to the trajectory similarity score.

[0034] Combine the compensation scheduling plan with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, so as to realize the adaptive scheduling control of the flexible production line, including:

[0035] Square the difference between the compensation scheduling plan and the scheduling plan at the previous moment to obtain a first modulus value. Subtract the multi-agent reinforcement learning model from the scheduling plan at the previous moment item by item and take the absolute value of each item. Calculate the weighted sum of the absolute values of each item to obtain a second modulus value. Multiply the ratio of the first modulus value and the second modulus value by the temperature parameter and obtain a fusion scheduling vector through the sigmoid function.

[0036] Construct a scheduling instruction set based on the fusion scheduling vector, calculate the priority score for each scheduling instruction in the scheduling instruction set, and sort the scheduling instruction set according to the priority score to form a priority queue.

[0037] Decompose the scheduling instructions in the priority queue into an atomic operation sequence, construct a device capability matrix. The number of rows of the device capability matrix is the number of devices, and the number of columns is the number of atomic operation types. The elements in the device capability matrix represent the capability values of the corresponding devices to execute the corresponding atomic operations. Calculate the mapping probability of each atomic operation to each device based on the device capability matrix.

[0038] Allocate the atomic operation sequence to the corresponding devices according to the mapping probability, sort the atomic operations assigned to the same device according to the priority score to form a device control instruction sequence, and control each device to execute the corresponding processing operations based on the device control instruction sequence to realize the adaptive scheduling control of the flexible production line.

[0039] In the second aspect of the embodiments of the present invention,

[0040] Provide an electronic device, including:

[0041] A processor;

[0042] A memory for storing processor-executable instructions;

[0043] Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0044] In a third aspect of the embodiments of the present invention,

[0045] A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0046] The beneficial effects of this application are as follows:

[0047] By mapping multiple processing devices as agents and establishing a local communication network, information sharing and collaborative decision-making among agents are realized, effectively solving the computational bottleneck problem of traditional central control methods in complex production environments, and improving decision-making efficiency and system response speed.

[0048] An experience replay pool is used to store system reward values and update the state-action value function, enabling agents to learn from historical experience, continuously optimize decision-making strategies, improve scheduling quality and system stability, and reduce production interruptions and resource waste.

[0049] During operation, the device status is monitored in real time and the state space and action space are dynamically updated, which can quickly respond to abnormal situations such as device failures or material shortages on the production line, automatically generate a compensation scheduling plan, significantly enhance the anti-interference ability and self-adaptability of the production system, and ensure the continuity and reliability of the production process. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic flowchart of the flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning according to the embodiments of the present invention;

[0051] Figure 2 It is a schematic diagram of the dynamic communication network topology structure in the technical solution of the embodiments of the present invention;

[0052] Figure 3 It is a schematic diagram comparing the device resource utilization rates of different agent scheduling methods according to the embodiments of the present invention;

[0053] Figure 4 It is a logic block diagram of the multi-agent reinforcement learning anomaly compensation scheduling according to the embodiments of the present invention;

[0054] Figure 5 It is a schematic diagram of the dynamic change process of the integrated scheduling vector in the technical solution of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0056] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0057] Figure 1 It is a schematic flowchart of a flexible production line adaptive scheduling control method that integrates multi-agent reinforcement learning in the embodiments of the present invention, as Figure 1 shown, the method includes:

[0058] Obtain the real-time operating status of multiple processing devices in the flexible production line and the processing requirement information of multiple workpieces, and construct a state space and an action space according to the real-time operating status and the processing requirement information;

[0059] Map the multiple processing devices to corresponding agents, establish a local communication network between the agents to achieve information sharing among adjacent agents, and the agents select and execute scheduling actions according to the current state information and the local observation information of neighboring agents;

[0060] After executing the scheduling action, obtain the system immediate reward value and store it in the experience replay pool, sample data from the experience replay pool to update the state-action value function, and obtain an optimized multi-agent reinforcement learning model through repeated iterative training;

[0061] During the operation of the multi-agent reinforcement learning model, continuously monitor the operating status of the multiple processing devices. When a device failure or material shortage is identified, dynamically update the state space and action space of the affected area, and use the updated state space and action space to recalculate the scheduling strategy to obtain a compensation scheduling plan;

[0062] Combine the compensation scheduling plan with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, thereby realizing the adaptive scheduling control of the flexible production line.

[0063] In an optional implementation manner, constructing a state space and an action space according to the real-time operating status and the processing requirement information includes:

[0064] The real-time operating status in the flexible production line includes equipment status, task queue, and remaining processing time, and the processing requirement information includes the processing sequence and processing time of the workpiece.

[0065] Map the real-time operating status and the processing requirement information to the feature vector space, extract the spatial features in the feature vector space through a deep neural network, and dynamically construct the state space and the action space according to the spatial features, where the state space represents the real-time state distribution of the processing equipment and the workpiece, and the action space represents the set of executable scheduling decisions.

[0066] The acquisition of the real-time operating status is the basis of the system. The real-time operating status includes equipment status, task queue, and remaining processing time. The equipment status refers to the current working status of each device, such as whether it is idle, in processing, or in a fault state. The task queue is the list of current workpieces to be processed, containing the processing requirement information of each workpiece, such as processing sequence and processing time. The remaining processing time refers to the remaining processing duration required for the current workpiece on the device. This information can be collected in real time through sensors and the device management system to form a dynamically updated set of status information.

[0067] Mapping the real-time operating status and the processing requirement information to the feature vector space is a key step in constructing the state space and the action space. The construction of the feature vector space requires converting information such as equipment status, task queue, and remaining processing time into a unified feature representation. Specifically, the status of each device can be encoded, for example, encoding the idle state as 0, the processing state as 1, and the fault state as 2. At the same time, each workpiece in the task queue can be characterized by its processing sequence and processing time. Finally, all this information will be integrated into a feature vector to form a multi-dimensional feature space.

[0068] Extract the spatial features in the feature vector space through a deep neural network. The design of the deep neural network should include multiple layers of neurons to be able to capture the relationships between complex states. During the training process, the input feature vectors will be processed through the various layers of the network to gradually extract high-level feature representations. These feature representations can effectively reflect the real-time state distribution between the equipment and the workpiece, providing a basis for the subsequent construction of the state space and the action space.

[0069] Dynamically construct the state space and action space. The construction of the state space should take into account the real-time state distribution of the equipment and workpieces. Specifically, the state space can be divided into multiple sub-spaces, each corresponding to a specific combination of equipment state and workpiece state. For example, the state space can be divided into multiple state combinations such as "equipment is idle and there are workpieces to be processed", "equipment is busy and there are no workpieces to be processed", etc. Such a division can not only clearly represent the current operating state of the system but also provide a clear basis for subsequent scheduling decisions.

[0070] The action space is a set of schedulable decisions executable according to the current state. Each scheduling decision should correspond to a specific scheduling action, such as assigning a certain workpiece to a specific equipment for processing. The construction of the action space needs to combine the current state information to ensure that the generated scheduling decisions are feasible. For example, in the state where the equipment is idle and there are workpieces to be processed, the first workpiece in the queue can be selected and assigned to the equipment for processing. In this way, the action space can dynamically adapt to the changes in the real-time operating state.

[0071] To verify the effectiveness of the above method, it can be tested through specific data cases. For example, assume there is a flexible production line with three pieces of equipment and five workpieces. The processing time and sequence of each workpiece are known, and the state of the equipment is fed back in real time through sensors. During the operation of the system, it first obtains the real-time state information to form a feature vector. Then, it extracts features through a deep neural network and constructs the state space and action space. Finally, the system selects the optimal scheduling decision according to the current state to complete the processing of the workpieces.

[0072] In an alternative embodiment, map the multiple processing devices to corresponding agents, establish a local communication network between the agents to achieve information sharing among adjacent agents, and the agents select and execute scheduling actions according to the current state information and the local observation information of neighboring agents, including:

[0073] Map the multiple processing devices in the flexible production line to corresponding agents, and each agent extracts features from its own state information based on the attention mechanism to obtain a state feature vector reflecting the real-time working characteristics of the processing device;

[0074] Construct an agent communication network with a dynamic topology, model the association degree between the agents based on the agent communication network, determine the communication weight by calculating the state similarity and task relevance between the agents, and the communication weight is adaptively adjusted with the changes in the agent state and task assignment to achieve the interaction of differential information between the agents;

[0075] Construct a hybrid policy network based on the state feature vector and the differentiation information. The hybrid policy network combines a deterministic policy with random exploration, outputs a deterministic action distribution through the hybrid policy network, and introduces random noise to achieve policy exploration. The training objective of the hybrid policy network is jointly determined by the task completion situation and the cooperation efficiency;

[0076] Select a scheduling action according to the output result of the hybrid policy network, and store the state transition information and reward information after executing the scheduling action in the experience pool.

[0077] Multiple processing devices are mapped to corresponding agents. Each agent is responsible for monitoring and managing its corresponding processing device. The agent obtains its own state information through sensors and data acquisition modules, including the operating state of the device, the load condition, the fault information, etc. In order to extract the features of these state information, the agent adopts an attention mechanism. Specifically, the agent performs weighted processing on the collected state information to highlight important features, thereby generating a state feature vector. This vector can reflect the real-time working characteristics of the device, for example, the processing speed, energy consumption, and failure rate of the device.

[0078] Construct an agent communication network with a dynamic topology. Each agent establishes a communication connection with its neighboring agents to achieve information sharing. The degree of association between agents is quantified by establishing a model that takes into account the state similarity and task relevance between agents. The state similarity is evaluated by comparing the state feature vectors of agents, while the task relevance is evaluated based on the task types and priorities undertaken by agents. The communication weights are set according to the calculation results of these similarities and relevances, and are adaptively adjusted as the states of agents and task assignments change. This dynamic adjustment mechanism ensures that agents can optimize information exchange according to the actual situation and promote collaborative work.

[0079] Construct a hybrid policy network using the state feature vector and the differentiation information. This network combines a deterministic policy and a random exploration mechanism to improve the flexibility and effectiveness of scheduling decisions. Specifically, the hybrid policy network first generates a deterministic action distribution based on the current state feature vector to guide the agent to select the optimal scheduling action. At the same time, to avoid getting stuck in local optimal solutions, the hybrid policy network also introduces random noise to encourage the agent to explore other possible scheduling strategies. The training objective of the hybrid policy network is jointly optimized based on the task completion situation and the cooperation efficiency to ensure that the agent can achieve the best production efficiency when executing tasks.

[0080] Select specific scheduling actions based on the output results of the mixed-strategy network. For example, an agent may choose to adjust the operating speed of a processing device or reassign tasks to other devices. After executing the scheduling action, the agent records the state transition information and reward information and stores them in the experience pool. This information will be used for subsequent learning and optimization to help the agent continuously improve the accuracy and efficiency of scheduling decisions.

[0081] To verify the effectiveness of the above technical means, a specific data case can be considered. In a flexible production line, there are five processing devices, each of which is managed by an agent. Each agent collects real-time data of the device through sensors, including processing speed, fault information, etc., and generates a state feature vector. After processing, the communication network between agents is established. Agent A is connected to Agents B and C, and Agent D is connected to Agent E.

[0082] Suppose the state feature vector of Agent A shows that its device load is too high, while the device load of Agent B is relatively low. At this time, Agent A will request assistance from Agent B through the communication network. Agent B decides to transfer a part of the tasks to Agent A according to its state feature vector and task relevance. At this time, the dynamic adjustment of the communication weight ensures the timely transmission of information and the effective allocation of tasks.

[0083] The mixed-strategy network generates a deterministic action distribution based on the state feature vector. Agent A finally chooses to reduce the device operating speed by 10%. After execution, Agent A records the state transition information and the obtained rewards (such as improved production efficiency) and stores them in the experience pool for subsequent learning and optimization.

[0084] Figure 2 Schematic diagram of the dynamic communication network topology structure in the technical solution of the embodiment of the present invention:

[0085] This figure shows a state and communication relationship diagram of an intelligent agent collaboration network for processing equipment. The figure contains six processing equipment agents (A - F), and the operating status and load rate are marked inside each agent node. Among them, equipment A, B, D, and E are in the running state, with load rates of 87%, 92%, 78%, and 65% respectively; equipment C is in the standby state, with a load rate of 23%; equipment F is in the fault repair state, with a load rate of 0%. The connection lines between the agents represent the communication weights, which are divided into three categories according to the weight size: high communication weight (>0.8) is represented by a solid line, including 0.92 between A - B, 0.88 between B - D, and 0.84 between D - E; medium communication weight (0.4 - 0.8) is represented by a dashed line, including 0.65 between A - C, 0.71 between B - E, and 0.54 between E - F; low communication weight (<0.4) is represented by a dotted line, including 0.35 between C - E and 0.26 between C - F. This network structure reflects the strength of the collaboration relationship between the devices. The devices with high - load operation maintain strong communication connections, while the communication intensity between the faulty or standby devices and other devices is relatively weak.

[0086] Figure 3 Schematic diagram for comparing the equipment resource utilization rates of different intelligent agent scheduling methods in the embodiments of the present invention:

[0087] This figure shows the comparison results and their averages of the equipment utilization rates of three different methods on five processing equipment (A - E). A line chart is used in the figure. The horizontal axis represents different processing equipment, and the vertical axis represents the equipment utilization percentage. The performance of this technical solution is the best, and its utilization rate curve is always at the top. The utilization rates from equipment A to equipment E are 84%, 80%, 87%, 89%, and 82% respectively, and the average utilization rate reaches 84.4%; the traditional multi - agent method ranks second, with the corresponding utilization rate data being 70%, 72%, 68%, 65%, and 69%, and the average utilization rate is 68.8%; the rule - based method performs the worst, with the equipment utilization rates being 55%, 60%, 52%, 47%, and 57% respectively, and the average utilization rate is only 54.2%. From the data trend, this technical solution maintains a high and stable utilization rate on all equipment, reaching a maximum of 89% and a minimum of 80%, demonstrating significant performance advantages; while the other two methods not only have a low overall utilization rate level but also have large fluctuations. In particular, the utilization rate of the rule - based method drops to a low point of 47% on equipment D.

[0088] In an alternative embodiment, an intelligent agent communication network with a dynamic topology is constructed, the correlation degree between the agents is modeled based on the intelligent agent communication network, and the communication weight is determined by calculating the state similarity and task relevance between the agents, including:

[0089] Construct multiple agents into a directed graph network structure, where the nodes in the directed graph network structure represent agents, and the edges in the directed graph network structure represent communication links. For each agent, determine its corresponding set of neighbor nodes, obtain the device working parameter vector, task queue status vector, and resource status vector of each agent, and combine the device working parameter vector, the task queue status vector, and the resource status vector to form the state vector of the agent;

[0090] Extract features from the state vector to obtain a feature vector, calculate the inner product of the feature vectors corresponding to any two agents, and divide the inner product by the square root of the feature dimension and then pass it through the softmax activation function to obtain the connection strength;

[0091] Calculate the dot product of any two state vectors divided by the product of their respective vector norms to obtain the state similarity, calculate the cosine similarity between any two agents, use it as the task relevance, and weight the state similarity and the task relevance to obtain the communication weight.

[0092] In constructing a dynamic topology agent communication network, it is first necessary to construct multiple agents into a directed graph network structure. In this structure, each node represents an agent, and the edges represent the communication links between agents. To achieve this goal, it is first necessary to determine the set of neighbor nodes of each agent, that is, the agents directly connected to it.

[0093] For each agent, it is necessary to obtain its device working parameter vector, task queue status vector, and resource status vector. These vectors respectively represent the working state of the agent, the tasks to be processed currently, and the available resource situation. Combine these three vectors together to form a comprehensive state vector for subsequent analysis.

[0094] In the process of feature extraction of the state vector, principal component analysis or other dimensionality reduction techniques can be used to extract key features from it to form a feature vector. The generation of the feature vector is to better capture the behavior patterns and state characteristics of the agents.

[0095] For any two agents, calculate the inner product of their corresponding feature vectors, and divide the result by the square root of the feature dimension. This process can be understood as measuring the similarity between two agents in the feature space. By performing softmax activation processing on the inner product result, the connection strength can be obtained, and this strength reflects the potential communication ability between agents.

[0096] In the calculation of state similarity, first, the dot product operation needs to be performed on the state vectors of any two agents. Subsequently, the dot product result is divided by the product of the respective vector norms to measure their similarity. The level of state similarity directly affects the information exchange efficiency between agents.

[0097] Calculating the cosine similarity between any two agents can be regarded as another method to measure the task relevance between agents. The higher the value of the cosine similarity, the stronger the correlation between the two agents in task execution.

[0098] The calculated state similarity and task relevance are weighted to obtain the final communication weight. This weight will be used to guide the communication decisions between agents, ensuring the effective transmission of information and the reasonable allocation of resources.

[0099] Suppose there are three agents A, B, and C. The device operating parameter vector of agent A is [0.8, 0.6, 0.7], the task queue state vector is [1, 0, 1], and the resource state vector is [0.9, 0.5]. The corresponding vectors of agent B are [0.7, 0.5, 0.6], [0, 1, 0], and [0.8, 0.4], and the vectors of agent C are [0.9, 0.7, 0.8], [1, 1, 0], and [0.6, 0.5].

[0100] By combining these vectors, the state vector of agent A is [0.8, 0.6, 0.7, 1, 0, 1, 0.9, 0.5], the state vector of agent B is [0.7, 0.5, 0.6, 0, 1, 0, 0.8, 0.4], and the state vector of agent C is [0.9, 0.7, 0.8, 1, 1, 0, 0.6, 0.5].

[0101] After feature extraction, assume the obtained feature vectors are A', B', and C' respectively. Calculate the inner product of A' and B', and the result is X. Subsequently, divide X by the square root of the feature dimension to obtain the connection strength. Perform the dot product operation on the state vectors A and B to get Y, and then divide Y by the product of the respective vector norms to obtain the state similarity. Then, calculate the cosine similarity between A and B, and finally, weight the state similarity and task relevance to obtain the communication weight.

[0102] In an alternative embodiment, during the operation of the multi-agent reinforcement learning model, it continuously monitors the operating states of the multiple processing devices. When a device failure or material shortage is identified, it dynamically updates the state space and action space of the affected area, and uses the updated state space and action space to recalculate the scheduling policy to obtain a compensation scheduling plan, including:

[0103] Collect the vibration data, temperature data, and current data of the acquisition and processing equipment to construct a multi-dimensional sensing data matrix, extract spatio-temporal features from the multi-dimensional sensing data matrix, perform time-series modeling through the multi-dimensional sensing data matrix based on the spatio-temporal features to obtain an equipment status feature sequence, obtain the material inventory data and material consumption data in the material management system, perform weighting to obtain a material supply prediction result, and generate a system anomaly situation vector based on the equipment status feature sequence and the material supply prediction result;

[0104] The system anomaly situation vector uses a directed acyclic graph to describe the propagation path of abnormal events, calculates the influence probability of abnormal events on different production units according to the propagation path, compares the influence probability with a preset probability threshold to determine the affected area, and predicts the state evolution trajectory during the abnormal propagation based on the affected area through probabilistic graph reasoning;

[0105] Based on the state evolution trajectory, retrieve the historical compensation plan for similar scenarios from the compensation strategy knowledge base, and use the meta-learning network to transfer the historical compensation plan to the current abnormal scenario. The meta-learning network includes a policy network and a value network. The policy network outputs the probability distribution of compensation actions, and the value network evaluates the long-term benefits of compensation actions. Select the candidate compensation plan with the highest benefit as the final compensation scheduling plan.

[0106] Through sensors installed on multiple processing devices, the vibration data, temperature data, and current data of the devices are collected in real time. These data are integrated into a multi-dimensional sensing data matrix, and each dimension of the matrix represents a type of sensor data. Next, feature extraction technology is used to process the multi-dimensional sensing data matrix to extract spatio-temporal features. These spatio-temporal features reflect the operating state of the device at different time periods and can effectively identify the normal and abnormal states of the device.

[0107] After the spatio-temporal feature extraction is completed, the system uses time-series modeling technology to further analyze the multi-dimensional sensing data matrix to generate an equipment status feature sequence. This sequence contains the state change information of the device within a certain period of time and can provide basic data for subsequent fault detection and material supply prediction.

[0108] Perform data interaction with the material management system to obtain the material inventory data and material consumption data. By weighting these data, the system can generate a material supply prediction result. This result provides the necessary material supply background information for subsequent anomaly situation analysis.

[0109] Based on the equipment status feature sequence and the material supply prediction result, the system generates a system anomaly situation vector. This vector is used to describe the abnormal state of the current system and can reflect the severity of problems such as equipment failures and material shortages.

[0110] After the abnormal situation vector is generated, the system uses a directed acyclic graph (DAG) model to describe the propagation path of abnormal events. Through the relationship between nodes and edges, this model clarifies the possible propagation paths of abnormal events during the production process. The system calculates the influence probability of abnormal events on different production units based on this propagation path, and the calculation of the influence probability is based on the comprehensive analysis of equipment status characteristics and material supply forecasts.

[0111] The calculated influence probability is compared with a preset probability threshold to determine the affected area. The affected area refers to those production units that are greatly affected by abnormal events, and these areas need to be preferentially optimized for scheduling.

[0112] Based on the affected area, the system predicts the state evolution trajectory during the abnormal propagation process through probabilistic graph inference technology. This trajectory reflects the change trend of the states of each production unit within the affected area after the abnormal event occurs, providing a basis for formulating a compensation scheduling plan.

[0113] After obtaining the state evolution trajectory, the system retrieves historical compensation plans similar to the current abnormal scenario from the compensation strategy knowledge base. Historical compensation plans refer to the scheduling optimization measures taken for similar abnormal situations during past production processes. The system uses a meta-learning network to perform transfer learning on these historical compensation plans to adapt to the current abnormal scenario. The meta-learning network consists of a policy network and a value network. The policy network is responsible for outputting the probability distribution of compensation actions, while the value network evaluates the long-term benefits of compensation actions.

[0114] During the generation of the compensation plan, the system evaluates all candidate compensation plans and selects the plan with the highest long-term benefits as the final compensation scheduling plan. This plan can effectively alleviate production problems caused by equipment failures or material shortages, ensuring the continuity and efficiency of the production process.

[0115] Suppose that during the operation of a processing device in a manufacturing enterprise, abnormal vibrations occur, and the vibration data collected by sensors in real time is significantly higher than the normal range. Through feature extraction and time series modeling, the system determines that the device may have a fault. At the same time, the material management system shows that the inventory of raw materials required by this device is insufficient to meet production needs. The abnormal situation vector generated by the system indicates that both equipment failure and material shortage exist and affect multiple production units.

[0116] Through the directed acyclic graph model, the system identifies the propagation path of abnormal events and calculates the influence probability of different production units being affected. Suppose the influence probability of a certain production unit is higher than the preset threshold, and the system marks it as the affected area. Subsequently, the system predicts the state evolution trajectory within this area and finds that if no measures are taken, the production efficiency will decrease by 30%.

[0117] When retrieving the historical compensation scheme, the system found that there were similar cases of equipment failures and material shortages in the past. The relevant compensation schemes included adjusting the production plan, temporarily increasing material procurement, and optimizing equipment maintenance time. Through the meta-learning network, the system migrated these historical schemes to the current scenario and selected a scheme of adjusting the production plan after evaluation, which was expected to restore the production efficiency to the normal level.

[0118] Figure 4 The logic block diagram of multi-agent reinforcement learning anomaly compensation scheduling for the embodiments of the present invention:

[0119] This figure shows a complete system working flow chart, which mainly includes two parallel data acquisition and processing branches, and subsequent analysis and decision-making processes. The left branch starts from the acquisition of device sensing data, including parameters such as vibration, temperature, and current. After the extraction of multi-dimensional sensing data matrix features, time series modeling is performed to obtain the device state feature sequence. The right branch starts from the acquisition of material data, including inventory and consumption data, and performs material supply prediction through weighted processing of material data. The outputs of these two branches are jointly fed into the system anomaly situation vector generation link. Then, the process is divided into three parallel paths: the first path analyzes the anomaly propagation path through directed acyclic graph modeling; the second path calculates the influence probability to determine the change influence area, and outputs the probability graph reasoning result for state evolution trajectory prediction; the third path is connected to the compensation strategy knowledge base to retrieve historical compensation schemes, calculates the probability distribution of compensation actions through the meta-learning network (including the policy network and value network), and finally generates and executes the optimal compensation scheduling scheme through compensation action evaluation and long-term benefit calculation. This flow chart reflects a complete closed-loop control process from data acquisition, feature extraction, state modeling to decision execution, and realizes the intelligent operation management of the system through multi-dimensional data fusion and multi-level analysis.

[0120] In an alternative embodiment, retrieving the historical compensation scheme of a similar scenario from the compensation strategy knowledge base based on the state evolution trajectory includes:

[0121] Obtain the state data of the production equipment, where the state data includes equipment state parameters, process parameters, and constraint condition parameters, and splice the state data in the feature dimension to form a dimension feature vector;

[0122] Perform time series sampling on the dimension feature vector to obtain a sampling sequence, arrange the dimension feature vectors in the sampling sequence in chronological order and splice them in the time dimension to form a time series feature matrix. The number of rows of the time series feature matrix corresponds to the dimension of the dimension feature vector, and the number of columns of the time series feature matrix corresponds to the number of sampling time points. The time series feature of the device state evolution is characterized by the time series feature matrix;

[0123] Construct a temporal decay function based on the temporal feature matrix. The temporal decay function uses a negative exponential form to assign different weight coefficients to the dimensional feature vectors at different times, and realizes the temporal weighting of the historical state information through the weight coefficients.

[0124] Calculate the similarity between the temporal feature matrix and the historical temporal feature matrices stored in the compensation strategy library, calculate the feature similarity between the dimensional feature vectors at the same time, and multiply the feature similarity by the weight coefficient at the corresponding time to obtain the trajectory similarity score; sort the historical compensation strategies in the compensation strategy library according to the trajectory similarity score.

[0125] In modern manufacturing, the state monitoring of production equipment and the optimization of compensation strategies are the keys to improving production efficiency and reducing failure rates. The following is a specific implementation method for retrieving historical compensation schemes for similar scenarios from the compensation strategy knowledge base based on the state evolution trajectory.

[0126] Collect the state data of production equipment. The state data includes equipment state parameters (such as temperature, pressure, rotational speed, etc.), process parameters (such as feeding speed, processing time, etc.) and constraint condition parameters (such as maximum load, ambient temperature, etc.). These data are collected in real time by sensors and stored in a database.

[0127] Process the collected state data. First, it is necessary to splice different types of state data in the feature dimension to form a dimensional feature vector. Each element of the feature vector represents a specific state parameter, and the spliced feature vector can comprehensively reflect the current state of the equipment. For example, assuming that the temperature of a certain equipment is 70°C, the pressure is 1.2 MPa, and the rotational speed is 1500 rpm, the corresponding feature vector can be expressed as [70, 1.2, 1500].

[0128] After obtaining the dimensional feature vector, perform temporal sampling. The process of temporal sampling is to arrange the feature vectors in chronological order to form a temporal feature matrix. The number of rows of this matrix corresponds to the dimension of the feature vector, and the number of columns corresponds to the number of sampling time points. The temporal feature matrix can clearly show the evolution process of the equipment state over time. For example, assuming that the feature vectors collected at 5 different time points are [70, 1.2, 1500], [72, 1.3, 1480], [71, 1.25, 1490], [73, 1.4, 1470], [75, 1.5, 1460] respectively.

[0129] Construct a time-series decay function. The purpose of the time-series decay function is to assign different weight coefficients to the dimensional feature vectors at different times, so as to better reflect the influence of historical state information on the current state. Specifically, higher weights are assigned to more recent time points, while lower weights are assigned to more distant time points. This weight assignment can be achieved by setting a decay factor, such that the influence of states further in the past on the current decision is smaller.

[0130] After completing the construction of the time-series decay function, calculate the similarity. Compare the current time-series feature matrix with the historical time-series feature matrices stored in the compensation strategy library. The process of calculating the similarity is to compare the dimensional feature vectors at the same time and evaluate the feature similarity between them. This process can be achieved by calculating the distance or similarity metric between the feature vectors. The calculated feature similarity needs to be multiplied by the weight coefficient corresponding to that time to obtain the trajectory similarity score.

[0131] Sort the historical compensation strategies in the compensation strategy library according to the trajectory similarity scores. The higher the score, the more similar the strategy is to the current device state, and thus the more likely it is to be effective in the current situation. In this way, the most suitable strategy for the current device state can be quickly screened out from the historical compensation strategies, thereby achieving efficient compensation decisions.

[0132] In an alternative implementation, combine the compensation scheduling scheme with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, so as to realize the adaptive scheduling control of the flexible production line, including:

[0133] Square the difference between the compensation scheduling scheme and the scheduling scheme at the previous moment to obtain a first modulus value. Subtract the multi-agent reinforcement learning model from the scheduling scheme at the previous moment item by item and take the absolute value of each item. Calculate the weighted sum of the absolute values of each item to obtain a second modulus value. Multiply the ratio of the first modulus value and the second modulus value by the temperature parameter and then pass it through the sigmoid function to obtain a fusion scheduling vector;

[0134] Construct a set of scheduling instructions based on the fusion scheduling vector. Calculate the priority score for each scheduling instruction in the set of scheduling instructions. Sort the set of scheduling instructions according to the priority score to form a priority queue;

[0135] Decompose the scheduling instructions in the priority queue into an atomic operation sequence, construct a device capability matrix, where the number of rows of the device capability matrix is the number of devices, the number of columns is the number of atomic operation types, and the elements in the device capability matrix represent the capability values of the corresponding devices to execute the corresponding atomic operations. Calculate the mapping probability of each atomic operation to each device based on the device capability matrix;

[0136] Allocate the atomic operation sequence to the corresponding devices according to the mapping probability, sort the atomic operations allocated to the same device according to the priority score to form a device control instruction sequence, and control each device to execute the corresponding processing operation based on the device control instruction sequence to achieve the adaptive scheduling control of the flexible production line.

[0137] Obtain the compensation scheduling plan at the current moment and the scheduling plan at the previous moment. By performing a square operation on the difference between the two, obtain the first modulus value. This modulus value reflects the degree of change between the current scheduling plan and the previous plan, providing the basic data for adjusting the scheduling strategy.

[0138] Obtain the output result of the multi-agent reinforcement learning model, subtract it item by item from the scheduling plan at the previous moment, and calculate the absolute value of each item. These absolute values represent the deviations between the current scheduling plan and the previous plan in various aspects. Perform a weighted calculation on these absolute values to obtain the second modulus value, which is used to measure the stability and reliability of the current scheduling plan relative to the historical plan.

[0139] Calculate the ratio of the first modulus value to the second modulus value, multiply it by a temperature parameter, and then process it through the sigmoid function to finally generate a fused scheduling vector. This vector combines the compensation scheduling plan and the output result of the multi-agent reinforcement learning model, providing a balanced basis for scheduling decisions.

[0140] Based on the fused scheduling vector, construct a set of scheduling instructions. Calculate the priority score for each instruction in the set of scheduling instructions. The priority score can be based on various factors, such as task urgency, resource availability, device status, etc. Sort the priority scores to form a priority queue to ensure that high-priority scheduling instructions can be executed first.

[0141] Decompose the scheduling instructions in the priority queue into an atomic operation sequence. Each atomic operation represents the basic operation steps in the production process. Construct a device capability matrix, where the number of rows of the matrix corresponds to the number of devices, and the number of columns corresponds to the number of atomic operation types. Each element in the matrix represents the capability value of the corresponding device to execute a specific atomic operation. By evaluating the device capabilities, the efficiency and adaptability of each device in executing different operations can be determined.

[0142] After constructing the device capability matrix, calculate the mapping probability of each atomic operation to each device. The mapping probability reflects the likelihood of a specific atomic operation being executed on a specific device, taking into account factors such as the current load of the device, historical execution efficiency, and the health status of the device, etc.

[0143] According to the calculated mapping probability, allocate the atomic operation sequence to the corresponding device. For the atomic operations allocated to the same device, sort them according to the priority score to form a device control instruction sequence. This sequence clearly indicates the operation order that each device needs to execute, thereby achieving efficient production scheduling.

[0144] Based on the device control instruction sequence, control each device to execute the corresponding processing operations. By real-time monitoring the execution status of the device, ensure the flexibility and adaptability of the production line. The entire scheduling control process can dynamically respond to changes in the production environment, optimize resource allocation, and improve production efficiency.

[0145] Suppose there are three devices, namely device A, device B, and device C, in a flexible production line that can execute five atomic operations. Assume that the compensation scheduling scheme at the current moment is [2, 3, 1], and the scheduling scheme at the previous moment is [1, 2, 2]. Then the first modulus value is (2 - 1)² + (3 - 2)² + (1 - 2)² = 1 + 1 + 1 = 3. Assume that the output of the multi-agent reinforcement learning model is [3, 1, 2]. Then the second modulus value is |3 - 1| + |1 - 2| + |2 - 2| = 2 + 1 + 0 = 3. The fusion scheduling vector obtained after processing will be a decision basis that comprehensively considers current and historical data.

[0146] Figure 5 For the embodiment of the present invention, this is a schematic diagram of the dynamic change process of the fusion scheduling vector in the technical solution of the present invention:

[0147] The figure shows the dynamic change process of three indicators, namely the compensation strategy weight, the reinforcement learning weight, and the temperature parameter, within a 90-second time range of the system, which includes two key time points: a failure occurs at 20 seconds, and the failure is resolved at 50 seconds. Specifically, the temperature parameter (diamond marker) remains at about 0.75 before the failure occurs, rises rapidly to 0.98 after the failure occurs, then decreases slowly, and continues to drop to 0.75 after the failure is resolved; the compensation strategy weight (circular marker) maintains at 0.40 before the failure occurs, rises significantly to 0.90 after the failure occurs, but gradually decreases to 0.65 during the failure, and continues to decrease and finally stabilizes at about 0.40 after the failure is resolved; the reinforcement learning weight (square marker) remains at a relatively high level of 0.85 before the failure occurs, drops briefly to 0.35 after the failure occurs, then gradually rises to 0.65 during the failure, and continues to rise and finally stabilizes at about 0.85 after the failure is resolved. The changing trends of these three curves clearly reflect the dynamic adjustment strategies of each control parameter of the system during the occurrence, handling, and recovery of the failure, and embody the adaptive adjustment ability of the system.

[0148] In the second aspect of the embodiments of the present invention,

[0149] a kind of electronic device is provided, including:

[0150] a processor;

[0151] a memory for storing instructions executable by the processor;

[0152] wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0153] In the third aspect of the embodiments of the present invention,

[0154] a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0155] The present invention can be a method, a device, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are uploaded.

[0156] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A flexible production line adaptive scheduling control method integrating multi-agent reinforcement learning, characterized in that, Including: Obtain the real-time operating status of multiple processing devices in the flexible production line and the processing requirement information of multiple workpieces, and construct a state space and an action space according to the real-time operating status and the processing requirement information; Map the multiple processing devices to corresponding agents, establish a local communication network between the agents to achieve information sharing among adjacent agents, and the agents select and execute scheduling actions according to the current state information and the local observation information of neighboring agents; After executing the scheduling action, obtain the system immediate reward value and store it in the experience replay pool, sample data from the experience replay pool to update the state-action value function, and obtain an optimized multi-agent reinforcement learning model through repeated iterative training; Collect vibration data, temperature data, and current data of the processing device to construct a multi-dimensional sensing data matrix, extract spatio-temporal features from the multi-dimensional sensing data matrix, perform time series modeling on the multi-dimensional sensing data matrix based on the spatio-temporal features to obtain a device state feature sequence, obtain the material inventory data and material consumption data in the material management system, perform weighting to obtain a material supply prediction result, and generate a system anomaly situation vector based on the device state feature sequence and the material supply prediction result; The system anomaly situation vector uses a directed acyclic graph to describe the propagation path of the anomaly event, calculates the influence probability of the anomaly event on different production units according to the propagation path, compares the influence probability with a preset probability threshold to determine the affected area, and predicts the state evolution trajectory during the anomaly propagation based on the affected area through probabilistic graph reasoning; Based on the state evolution trajectory, retrieve the historical compensation plan for similar scenarios from the compensation strategy knowledge base, and use the meta-learning network to transfer the historical compensation plan to the current anomaly scenario. The meta-learning network includes a policy network and a value network. The policy network outputs the probability distribution of the compensation action, and the value network evaluates the long-term benefits of the compensation action. Select the candidate compensation plan with the highest benefit as the final compensation scheduling plan; Combine the compensation scheduling plan with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, so as to realize the adaptive scheduling control of the flexible production line.

2. The method according to claim 1, characterized in that, Constructing a state space and an action space according to the real-time operating status and the processing requirement information includes: The real-time operating status in the flexible production line includes the device status, the task queue, and the remaining processing time, and the processing requirement information includes the processing sequence and processing time of the workpiece; Map the real-time operating status and the processing requirement information to the feature vector space, extract the spatial features in the feature vector space through a deep neural network, and dynamically construct a state space and an action space according to the spatial features. The state space represents the real-time state distribution of the processing device and the workpiece, and the action space represents the set of executable scheduling decisions.

3. The method according to claim 1, wherein Map the multiple processing devices to corresponding agents. A local communication network is established among the agents to achieve information sharing among adjacent agents. The agents select and execute scheduling actions based on the current state information and the local observation information of neighboring agents, including: Map the multiple processing devices in the flexible production line to corresponding agents. Each agent extracts features from its own state information based on the attention mechanism to obtain a state feature vector reflecting the real-time working characteristics of the processing device; Construct an agent communication network with a dynamic topology. Based on the agent communication network, model the association degree among the agents. Determine the communication weight by calculating the state similarity and task relevance among the agents. The communication weight is adaptively adjusted according to the changes in the agent state and task allocation to achieve the interaction of differential information among the agents; Construct a hybrid policy network based on the state feature vector and the differential information. The hybrid policy network combines the deterministic policy with random exploration. Output a deterministic action distribution through the hybrid policy network, and introduce random noise to achieve policy exploration. The training objective of the hybrid policy network is jointly determined by the task completion situation and the cooperation efficiency; Select a scheduling action according to the output result of the hybrid policy network, and store the state transition information and reward information after executing the scheduling action in the experience pool.

4. The method according to claim 3, wherein Construct an agent communication network with a dynamic topology. Based on the agent communication network, model the association degree among the agents. Determine the communication weight by calculating the state similarity and task relevance among the agents, including: Construct a directed graph network structure with multiple agents. The nodes in the directed graph network structure represent agents, and the edges in the directed graph network structure represent communication links. Determine the corresponding neighborhood node set for each agent, obtain the device working parameter vector, task queue state vector, and resource state vector of each agent, and combine the device working parameter vector, the task queue state vector, and the resource state vector to form the state vector of the agent; Extract features from the state vector to obtain a feature vector. Calculate the inner product of the feature vectors corresponding to any two agents, and obtain the connection strength after dividing the inner product by the square root of the feature dimension and passing through the softmax activation function; Calculate the dot product of any two state vectors divided by the product of their respective vector norms to obtain the state similarity. Calculate the cosine similarity between any two agents, and use it as the task relevance. Weight the state similarity and the task relevance to obtain the communication weight.

5. The method according to claim 1, characterized in that Based on the state evolution trajectory, retrieve the historical compensation scheme for similar scenarios from the compensation policy knowledge base, including: Obtain the state data of the production equipment. The state data includes equipment state parameters, process parameters, and constraint condition parameters. Concatenate the state data in the feature dimension to form a dimension feature vector; Perform time series sampling on the dimension feature vector to obtain a sampling sequence, arrange the dimension feature vectors in the sampling sequence in chronological order and splice them in the time dimension to form a time series feature matrix. The number of rows of the time series feature matrix corresponds to the dimension of the dimension feature vector, and the number of columns of the time series feature matrix corresponds to the number of sampling time points. The time series feature of the device state evolution is characterized by the time series feature matrix; Construct a time series attenuation function based on the time series feature matrix. The time series attenuation function uses a negative exponential form to assign different weight coefficients to the dimension feature vectors at different times, and realizes the time series weighting of historical state information through the weight coefficients; Calculate the similarity between the time series feature matrix and the historical time series feature matrix stored in the compensation strategy library, calculate the feature similarity between the dimension feature vectors at the same time, and multiply the feature similarity by the weight coefficient at the corresponding time to obtain a trajectory similarity score; Sort the historical compensation strategies in the compensation strategy library according to the trajectory similarity score.

6. The method according to claim 1, characterized in that, Combine the compensation scheduling plan with the output result of the multi-agent reinforcement learning model to generate a real-time scheduling instruction sequence. The real-time scheduling instruction sequence is decomposed to form specific control instructions for each processing device and sent to the corresponding processing device, so as to realize the adaptive scheduling control of the flexible production line, including: Square the difference between the compensation scheduling plan and the scheduling plan at the previous moment to obtain a first modulus value. Subtract the multi-agent reinforcement learning model from the scheduling plan at the previous moment item by item and take the absolute value of each item. Calculate the weighted sum of the absolute values of each item to obtain a second modulus value. Multiply the ratio of the first modulus value and the second modulus value by the temperature parameter and obtain a fusion scheduling vector through the sigmoid function; Construct a scheduling instruction set based on the fusion scheduling vector, calculate the priority score for each scheduling instruction in the scheduling instruction set, and sort the scheduling instruction set according to the priority score to form a priority queue; Decompose the scheduling instructions in the priority queue into an atomic operation sequence, construct a device capability matrix, the number of rows of the device capability matrix is the number of devices, the number of columns is the number of atomic operation types, and the elements in the device capability matrix represent the ability values of the corresponding devices to execute the corresponding atomic operations. Calculate the mapping probability of each atomic operation to each device based on the device capability matrix; Allocate the atomic operation sequence to the corresponding device according to the mapping probability, sort the atomic operations allocated to the same device according to the priority score to form a device control instruction sequence, and control each device to execute the corresponding processing operation based on the device control instruction sequence to realize the adaptive scheduling control of the flexible production line.

7. An electronic device, characterized in that, Including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Wind power assembly workshop multi-objective optimization scheduling method based on reinforcement learning

    CN118690897A

  • Methods and systems for optimization of network-sensitive data collection in an industrial drilling environment

    US20190025806A1