Multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment

By constructing a multi-agent collaborative optimization deployment method, the systemic limitations in industrial manufacturing scheduling are solved, achieving low-latency, high-precision resource allocation and environmental adaptability, thereby improving the operating efficiency and stability of the production line.

CN121900170APending Publication Date: 2026-04-21BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIHANG UNIV
Filing Date
2026-01-09
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing industrial manufacturing scheduling methods have systemic limitations, including real-time computational bottlenecks, lack of scalability, lack of coordination mechanisms, and lack of physical process constraint modeling. These limitations lead to resource allocation imbalances, decreased production efficiency, and increased energy consumption in complex dynamic environments.

Method used

A multi-agent collaborative optimization deployment method is constructed. Through distributed collaborative architecture, non-cooperative game theory and Nash equilibrium solution, combined with large-scale model distributed reasoning and hierarchical task routing, real-time information sharing and strategy coordination among agents are realized, dynamic adaptive deployment of resource scheduling is achieved, and a self-improving intelligent closed loop is constructed.

Benefits of technology

It significantly improves the system's adaptability and overall efficiency in complex multi-tasking environments, achieves low-latency and high-precision resource allocation, possesses environmental adaptability and fault tolerance capabilities, and improves the operating efficiency and stability of the production line.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121900170A_ABST
    Figure CN121900170A_ABST
Patent Text Reader

Abstract

The invention provides a multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment, and relates to the technical field of production scheduling. According to the deployment method, a multi-agent collaborative architecture is constructed, traditional centralized decision making is converted into a distributed intelligent collaborative system, all agents maintain local autonomous decision making, and meanwhile, global optimization configuration of tasks and resources is achieved through a real-time information sharing and strategy coordination mechanism, so that the deployment efficiency is improved. And the adaptive capacity and the overall efficiency of the system in a complex multi-task environment are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of production scheduling technology, specifically to a multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment. Background Technology

[0002] In the current industrial manufacturing field, systems generally adopt a single-agent and centralized computing model for task scheduling and resource management. As the complexity of production tasks increases, single-agent systems exhibit many limitations when dealing with dynamic and multi-constraint manufacturing environments.

[0003] The core of using single-agent optimization algorithms to handle production scheduling problems lies in employing a single, centralized agent that makes decisions based on global production information and controls the entire production line downwards. Common implementation methods include using optimization algorithms such as genetic algorithms and simulated annealing to seek a globally optimal solution for production task arrangement and resource allocation, treating the entire manufacturing system as a centralized entity for modeling and optimization.

[0004] When using a centralized decision-making model for intelligent decision-making, a centralized inference architecture is typically adopted: this architecture deploys the trained large model on a high-performance central server or computing node, and all data from the manufacturing site must be transmitted to this central node for unified processing and inference, and finally the decision results are sent down to the execution unit.

[0005] As industrial manufacturing equipment becomes increasingly intelligent, production scheduling and optimization problems are becoming more complex, placing higher demands on the system's dynamic response and collaborative capabilities. Meanwhile, large-scale models offer greater intelligence for decision-making, but their inherent computational complexity also presents new challenges to deployment methods.

[0006] Existing industrial manufacturing scheduling methods suffer from systemic limitations, real-time computational bottlenecks, lack of scalability, lack of robust coordination mechanisms, and lack of physical process constraint modeling. Specifically, these problems manifest in the following ways: While single-agent-based optimization scheduling methods are applicable in simple production scenarios, they exhibit significant shortcomings in complex and dynamic industrial environments. Relying solely on a single decision-making entity, these methods fail to effectively utilize the collaborative potential among multiple agents, making them ill-suited to the diverse tasks and resource coordination demands of modern production lines. Consequently, when dealing with complex scheduling tasks, they often fail to achieve global system optimization, easily leading to resource allocation imbalances, decreased production efficiency, and increased energy consumption.

[0007] For example, genetic algorithms or simulated annealing algorithms used in manufacturing scheduling, while capable of finding local optima, are prone to getting trapped in local optima under multi-task and multi-constraint conditions due to the lack of information interaction and cooperation mechanisms between agents. Furthermore, they cannot dynamically adjust strategies according to environmental changes. This not only restricts production efficiency but may also lead to resource waste and response delays.

[0008] Existing large-scale model inference typically employs a centralized computing architecture, concentrating all computational tasks on a single server or node. As model size continues to increase, this centralized architecture faces ever-growing computational and storage pressures. Industrial manufacturing environments often have limited hardware resources, making it difficult to support the high computational and memory requirements of large-scale deep learning models. Furthermore, centralized inference suffers from significant latency when processing real-time data and dynamic tasks. Critical tasks in industrial manufacturing, such as scheduling and quality inspection, demand extremely high real-time response rates, but centralized deployments struggle to guarantee low-latency feedback due to communication and computational bottlenecks. Especially in large-scale parallel scenarios, it fails to fully utilize distributed computing resources, leading to low resource utilization and reduced system efficiency.

[0009] Existing large-scale model inference deployment solutions generally lack elasticity and scalability, making it difficult to adapt to changes in different hardware conditions and network environments. Traditional methods typically rely on fixed hardware configurations or single-machine deployment modes, which cannot achieve flexible scheduling and efficient inference in distributed, heterogeneous industrial scenarios.

[0010] For example, in distributed manufacturing environments with numerous nodes and varying resources, existing centralized deployment solutions often fail to deploy or suffer from reduced inference performance due to hardware differences and system instability, which severely restricts the practicality and promotion value of large models in industrial settings.

[0011] While multi-agent systems theoretically possess advantages in collaborative optimization, existing methods still have significant shortcomings in practical applications. Many systems fail to effectively integrate game theory and distributed decision-making mechanisms, resulting in insufficient information sharing and low collaboration efficiency among agents, making it difficult to achieve global system optimization. The lack of flexible and efficient collaboration mechanisms prevents multi-agent systems from fully realizing their potential in improving productivity. Existing methods often focus on optimizing local tasks, neglecting the dynamic interactions and policy coordination required by agents during the production process, leading to limited overall performance improvements in multi-task environments.

[0012] Existing optimization methods are mostly based on abstract mathematical models, lacking explicit modeling of the physical laws governing the manufacturing process (such as material deformation, heat conduction, vibration attenuation, and fluid dynamics). This leads to a disconnect between scheduling and process parameter settings and the actual physical process, affecting the feasibility and accuracy of the optimization results. Traditional methods fail to incorporate constraints from natural laws such as gravity, friction, and thermodynamic parameters into the optimization model, which may result in scheduling schemes that are physically infeasible or inefficient. Summary of the Invention

[0013] To address the shortcomings of existing technologies, this invention provides a multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment, which solves the problem of systemic limitations in industrial manufacturing scheduling methods.

[0014] To achieve the above objectives, the present invention provides the following technical solution: A multi-agent collaborative optimization deployment method based on industrial manufacturing equipment, the deployment method comprising: S1. Initialize system parameters and agent configuration Construct a parameterized model of physical laws and a multi-agent distributed collaborative architecture to realize formal modeling of physical constraints and initialization of the distributed collaborative system; S2. Multi-agent cooperative decision-making modeling and optimization We construct a utility function that integrates local performance, global metrics, and physical constraint penalties, and combine non-cooperative game theory and Nash equilibrium solutions to achieve global collaborative optimization and physical feasibility assurance for agents under autonomous decision-making. S3. Large-Model Distributed Inference and Hierarchical Task Routing For complex sub-problems or real-time triggered intelligent tasks generated during collaborative decision-making, they are dynamically routed to edge computing nodes or cloud computing nodes according to task characteristics, and the corresponding large models are called for distributed reasoning to achieve intelligent allocation and load balancing of computing tasks, thereby obtaining low-latency and high-precision distributed reasoning capabilities. S4. Dynamic Adaptive Deployment and Elastic Resource Scheduling By using a multi-dimensional evaluation of system performance and a dual-trigger adjustment mechanism, dynamic optimization of model deployment, computing resources and task scheduling is achieved, resulting in a resilient operating system with environmental adaptability and fault tolerance. S5. Closed-Loop Execution and Online Learning By leveraging feedback from actual production data and multi-agent reinforcement learning, continuous online optimization of strategies, weights, and routing parameters is achieved, resulting in a self-improving intelligent closed loop that can adaptively balance scheduling efficiency and physical constraints.

[0015] Preferably, the system parameters in S1 include: multi-agent parameters; The multi-agent parameters include: Number of agents ; All intelligent agents constitute a set Each intelligent agent Responsible for a specific manufacturing unit or resource pool; Each agent Its state space is constituted by local observation information. , This includes the queue length of its resources, equipment utilization rate, and information on currently processed workpieces; Each agent All executable decision operations constitute its action space. , Including the selection of workpieces to be processed Adjust the process parameters of the controlled equipment. Sending cooperation requests to neighboring units, etc.; The communication topology between intelligent agents is defined as a graph. , where vertex set edge set This refers to a communication link that allows direct information exchange between intelligent agents.

[0016] Preferably, the system parameters in S1 further include: large model deployment parameters; The large model deployment parameters include: A global decision-making model deployed in the cloud or on a central server is denoted as Its parameter count is ; Deployed at the edge computing nodes The lightweight model is denoted as Its parameter count is , usually satisfy ; Define performance constraints for model inference: minimum accuracy requirement and maximum allowed end-to-end delay .

[0017] Preferably, the system parameters in S1 further include: manufacturing task and resource parameters; The manufacturing tasks and resource parameters include: Workpieces to be scheduled , This represents the total number of workpieces. For the workpiece The process sequence is denoted as , For workpiece The total number of processes; Available physical manufacturing resources of the system , Total number of resources; Define three-dimensional parameters Indicates workpiece The Steps In resources The time required for processing; among which, ; ; If a certain process Unable to access resources For processing, then it is agreed Or use a maximum value express.

[0018] Preferably, the system parameters in S1 further include: physical process constraint parameters; The physical process constraint parameters include: Each manufacturing resource Having a set of physical properties Each constraint corresponds to a natural law-based limitation in the manufacturing process: Machining mechanical constraints: Maximum permissible processing force ; Maximum depth of cut ; Thermodynamic constraints: Coefficient of thermal expansion of material ; Maximum allowable temperature rise ; Dynamic constraints: Maximum permissible vibration frequency ; Maximum permissible acceleration ; Define the penalty weight vector ; for Load quality; for Unloaded mass; It is the acceleration due to gravity; for Running speed.

[0019] Preferably, S2 specifically includes: Each intelligent agent Based on its perceived local state and through communication topology Receive information from neighboring intelligent agents and make distributed collaborative decisions based on a non-cooperative game model; For each intelligent agent Define a local utility function: in, For intelligent agents The chosen action; To remove The combination of actions chosen by all other intelligent agents; After the action is executed, the intelligent agent Average workpiece waiting time for the resources under its jurisdiction; This refers to the maximum completion time of the entire manufacturing system after the action is executed; , For workpiece Completion time; For action Perform the estimated resource and energy consumption; The core scheduling target weight coefficient; This is the physical constraint penalty weight coefficient, used to balance scheduling efficiency and physical feasibility; This represents the total number of physical constraint categories. For the first A penalty function for violating physical constraints, based on the physical state of the resources under the agent's control. calculate; The initial weights satisfy: ; The weighting coefficients are dynamically adjusted through subsequent feedback learning to balance multi-objective optimization between scheduling efficiency, energy consumption, and physical feasibility; The system employs a distributed iterative algorithm to find an approximate Nash equilibrium for the game. ); At this equilibrium point, for each agent ,satisfy: .

[0020] Preferably, S3 specifically includes: For complex sub-problems arising from game-theoretic decisions in S2 or intelligent tasks triggered in real time by sensors, the large model inference module is invoked. Task modeling and classification: Abstract each decision request into a reasoning task. And label its feature vector: ; in Input the amount of data for the task; Estimate the computational complexity of the task; The maximum tolerable delay for the task; For physical process characteristic sub-vectors; Dynamic routing strategy: The system maintains a routing policy function. The decision-making logic is as follows: IF: AND ; THEN: Routing to the edge layer, by comparing each edge node Current load Select the node with the least load to execute its lightweight model. ; ELSE: will Routing to the cloud, from a global large model implement; Edge layer load balancing: Define edge nodes At any moment Load index: ; in, for exist The length of the pending task queue at any given time; for exist CPU utilization percentage at any given time; This is the upper limit of the queue length; For weight parameters ( ).

[0021] Preferably, S4 specifically includes: Continuously monitor the overall status and dynamically adjust model deployment and resource allocation strategies; System overall performance evaluation function: Define the overall system performance indicators within the evaluation period. : in, The total number of manufacturing and reasoning tasks successfully completed per unit of time; Average end-to-end processing latency for all tasks; This is an estimated total system energy consumption. This is due to the overhead caused by model migration and task rescheduling; A score is given for the consistency between the scheduling results and the physical laws. The normalized weighting coefficients are positive. Adaptive adjustment of triggering and execution: Periodic triggering: every fixed time interval Calculate the current If the decline compared to the previous cycle exceeds the threshold If so, optimization will be triggered; Event trigger: When the load of any node is detected Node failure, or network bandwidth consistently below [a certain value] Triggered immediately upon arrival; Adjustment Actions: Adjustment decisions are based on a lightweight optimizer or policy network, and the actions include: Model transfer: transferring a marginal model Migrate from an overloaded node to a lightly loaded node; Calculate unloading: Remove excessively long queues from the edge layer. Unload to the cloud for processing in real time; Policy Update: Fine-tuning the utility function weights of agents in S2 Or the routing threshold in S3 , .

[0022] Preferably, S5 specifically includes: Execution and Feedback: The collaborative scheduling instructions generated by S2 and the intelligent decisions generated by S3 are sent to the physical manufacturing system for execution; the system collects the actual production data after execution. ; Reward Calculation and Model Update: Compare the actual results with the expected results to calculate the reward signal. ; (Planned completion time / Actual completion time) + (1-Defect Rate)- (Actual energy consumption / Baseline energy consumption) - ; in, This is a completion time bonus coefficient, used to reward early or on-time completion. This is a quality incentive coefficient used to reward high yield rates; This is the energy consumption penalty coefficient, used to penalize energy consumption exceeding the benchmark. This is the penalty coefficient for violating physical constraints, used to punish behaviors that violate physical constraints. This represents the total number of physical constraint categories. For the first Points deducted for violating physical constraints; Agent policy update: utilizing Feedback is used to update the various agents through a multi-agent reinforcement learning algorithm. The decision-making strategy, and adaptively adjust its utility function weights. and penalty weight : in: To penalize the learning rate of weights; For the most recent Within the first cycle Frequency of class constraint violations; The target violation rate threshold; Routing policy update: based on Optimize the routing strategy function in S3 based on the actual completion time and accuracy at different nodes. and threshold parameters.

[0023] A multi-agent collaborative optimization deployment system based on industrial manufacturing equipment, the deployment system includes: an initialization module for system parameters and agent configuration, a multi-agent collaborative decision modeling and optimization module, a large-model distributed reasoning and task hierarchical routing module, a dynamic adaptive deployment and resource elastic scheduling module, and a closed-loop execution and online learning module; The initialization system parameter and agent configuration module is used to construct a physical law parameterized model and a multi-agent distributed collaborative architecture to realize the formal modeling of physical constraints and the initialization of the distributed collaborative system, laying a structured foundation for physical perception scheduling and multi-agent collaborative optimization. The multi-agent collaborative decision-making modeling and optimization module is used to construct a utility function that integrates local performance, global indicators and physical constraint penalties, and combine non-cooperative game theory and Nash equilibrium solutions to achieve global collaborative optimization and physical feasibility assurance for agents under autonomous decision-making. The large-scale model distributed inference and task hierarchical routing module is used to achieve intelligent allocation and load balancing of computing tasks through multi-dimensional task feature extraction and dynamic edge-cloud routing strategies, thereby obtaining low-latency and high-precision distributed inference capabilities to support scenarios such as real-time quality inspection and fault diagnosis. The dynamic adaptive deployment and resource elastic scheduling module is used to achieve dynamic optimization of model deployment, computing resources and task scheduling through multi-dimensional evaluation of system performance and dual-trigger adjustment mechanism, so as to obtain an elastic operating system with environmental adaptability and fault tolerance. The closed-loop execution and online learning module is used to continuously optimize the strategy, weights and routing parameters online through feedback from actual production data and multi-agent reinforcement learning, so as to obtain a self-improving intelligent closed loop that can adaptively balance scheduling efficiency and physical constraints.

[0024] This invention provides a multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment. Compared with existing technologies, it has the following advantages: In this invention, the deployment method transforms traditional centralized decision-making into a distributed intelligent collaboration system by constructing a multi-agent collaborative architecture. While maintaining local autonomous decision-making, each agent achieves global optimization of task and resource allocation through real-time information sharing and strategy coordination mechanisms, significantly improving the system's adaptability and overall efficiency in complex multi-task environments. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart of the deployment method in an embodiment of the present invention.

[0027] Figure 2 This is an architecture diagram of multi-agent collaborative optimization in an embodiment of the present invention.

[0028] Figure 3 This is a flowchart of large-model distributed inference and task hierarchical routing in an embodiment of the present invention. Detailed Implementation

[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0030] This application provides a multi-agent collaborative optimization deployment method and system based on industrial manufacturing equipment, which solves the problem of systemic limitations in industrial manufacturing scheduling methods.

[0031] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0032] Example: like Figures 1-3 As shown, this invention provides a multi-agent collaborative optimization deployment method based on industrial manufacturing equipment, the deployment method comprising: S1. Initialize system parameters and agent configuration Based on the physical layout and production task characteristics of the target industrial manufacturing production line, set the input parameters of the framework: Multi-Agent (MA) parameters: Number of agents ; All intelligent agents constitute a set Each intelligent agent Responsible for a specific manufacturing unit or resource pool (such as machine tool clusters, robot assembly lines, AGV fleets, and automated warehouses). Each agent Its state space is constituted by local observation information. , This includes the queue length of its resources, equipment utilization rate, and information on currently processed workpieces; Each agent All executable decision operations constitute its action space. , Including the selection of workpieces to be processed Adjust the process parameters of the controlled equipment. Sending cooperation requests to neighboring units, etc.; The communication topology between intelligent agents is defined as a graph. , where vertex set edge set This refers to a communication link that allows direct information exchange between intelligent agents; Large Model (LM) deployment parameters: A global decision-making model deployed in the cloud or on a central server is denoted as Its parameter count is ; Deployed at the edge computing nodes The lightweight model is denoted as Its parameter count is , usually satisfy ; Define performance constraints for model inference: minimum accuracy requirement and maximum allowed end-to-end delay ; Manufacturing tasks and resource parameters: Assume there is a total There are one workpiece to be scheduled, denoted as... , This represents the total number of workpieces. For the workpiece Its processing requires going through The process sequence is denoted as follows: , For workpiece The total number of processes; Assume the system has a total of The available physical manufacturing resources (such as machines and workstations) in Taiwan are denoted as , Total number of resources; Define three-dimensional parameters Indicates workpiece The Steps In resources The time required for processing; among which, ; ; If a certain process Unable to access resources For processing, then it is agreed Or use a maximum value express; Physical process constraint parameters: Assume each manufacturing resource Having a set of physical properties Each constraint corresponds to a natural law-based limitation in the manufacturing process: Machining mechanical constraints: Maximum permissible processing force (Unit: N), corresponding resources ; Maximum depth of cut (Unit: mm); Thermodynamic constraints: Coefficient of thermal expansion of material (Unit: μm / m·℃); Maximum allowable temperature rise (Unit: °C) Dynamic constraints: Maximum permissible vibration frequency (Unit: Hz); Maximum permissible acceleration (Unit: m / s²) Define the penalty weight vector Suggested initial values: (Processing force constraints); (Thermal deformation constraint); (Vibration constraint); for Load mass (unit: kg); for Unloaded mass (unit: kg); Acceleration due to gravity (unit: m / s²) 2 ); for Operating speed (unit: m / s); The physical process constraint parameters can be adjusted according to the physical sensitivity of the specific manufacturing scenario.

[0033] S2. Multi-agent cooperative decision-making modeling and optimization based on game theory Each intelligent agent Based on its perceived local state and through communication topology Receive information from neighboring intelligent agents and make distributed collaborative decisions based on a non-cooperative game model; Definition of agent utility function: for each agent Define a local utility function This is used to quantify the benefits or costs under different joint actions; Formula definition: in, For intelligent agents The chosen action; To remove The combination of actions chosen by all other intelligent agents; After the action is executed, the intelligent agent Average workpiece waiting time for the resources under its jurisdiction; Makespan is the maximum completion time of the entire manufacturing system after the action is executed. , For workpiece Completion time; For action Perform the estimated resource and energy consumption; The core scheduling target weight coefficient; This is the physical constraint penalty weight coefficient, used to balance scheduling efficiency and physical feasibility; This represents the total number of physical constraint categories. For the first A penalty function for violating physical constraints, based on the physical state of the resources under the agent's control. calculate; Weighting system design: The initial weights satisfy: ; Suggested initial setup example: ; Example of a physical constraint penalty function: Penalty for exceeding processing force limit: Penalty for excessive heat deformation: Vibration overclocking penalty: ( (for indicator functions) The weighting coefficients can be dynamically adjusted through subsequent feedback learning to balance multi-objective optimization between scheduling efficiency, energy consumption, and physical feasibility; Nash equilibrium solution and decision generation: The system employs a distributed iterative algorithm (such as optimal response dynamics) to find an approximate Nash equilibrium for the game. ); At this equilibrium point, for each agent ,satisfy: ; This condition means that, given that all other agents adopt equilibrium strategies, no single agent can unilaterally deviate from its equilibrium action. To achieve higher self-utility, the system spontaneously converges to a stable and overall better collaborative decision-making state while respecting individual autonomy.

[0034] S3. Large-Model Distributed Inference and Hierarchical Task Routing For complex sub-problems arising from game-theoretic decisions in S2 (such as process parameter optimization) or intelligent tasks triggered in real time by sensors (such as vision-based precision quality inspection and complex fault diagnosis), the large model inference module is invoked. Task modeling and classification: Abstract each decision request into a reasoning task. And label its feature vector: ; in Input the data size for the task. Estimate the computational cost for the task. For task timeliness requirements (maximum tolerable delay); This is a feature vector representing the physical process; it can include material type (e.g., steel, aluminum, composite materials) and surface roughness requirements. (Unit: μm), workpiece mass (Unit: kg), impact Transportation energy consumption, ambient temperature (Unit: °C), affecting sensor accuracy and equipment stability; Dynamic routing strategy: The system maintains a routing policy function. The decision-making logic is as follows: IF: AND ; THEN: Routing to the edge layer, by comparing each edge node Current load (See the calculation formula below), select the node with the least load to execute its lightweight model. ; ELSE: will Routing to the cloud, from a global large model Execute the inference to obtain more accurate or comprehensive results. Edge layer load balancing: Define edge nodes At any moment Load index: ; in, for exist The length of the pending task queue at any given time; for exist CPU utilization percentage at any given time; This is the upper limit of the queue length; For weight parameters ( This is used to adjust the relative importance of queue length and CPU utilization.

[0035] S4. Dynamic Adaptive Deployment and Elastic Resource Scheduling This framework continuously monitors the overall status (including computing resources, network conditions, and task flow) and dynamically adjusts model deployment and resource allocation strategies. System overall performance evaluation function: Define the overall system performance indicators within the evaluation period. : in, The total number of manufacturing and reasoning tasks successfully completed per unit of time; Average end-to-end processing latency for all tasks; This is an estimated total system energy consumption. This is due to the overhead caused by model migration and task rescheduling; Scoring is given based on the consistency between the scheduling results and physical laws, such as whether the processing path conforms to the principle of minimum energy, the material handling path conforms to the principle of shortest displacement, and the process parameters are within the physical limits of the materials; The normalized weight coefficients are positive, reflecting the priority of different optimization objectives; Adaptive adjustment of triggering and execution: Periodic triggering: every fixed time interval Calculate the current If the decline compared to the previous cycle exceeds the threshold If so, optimization will be triggered; Event trigger: When the load of any node is detected Node failure, or network bandwidth consistently below [a certain value] Triggered immediately upon arrival; Adjustment Actions: Adjustment decisions are based on a lightweight optimizer or policy network, and possible actions include: Model transfer: transferring a marginal model Migrate from an overloaded node to a lightly loaded node; Calculate unloading: Remove excessively long queues from the edge layer. Unload to the cloud for processing in real time; Policy Update: Fine-tuning the utility function weights of agents in S2 Or the routing threshold in S3 , .

[0036] S5. Closed-Loop Execution and Online Learning Execution and Feedback: The collaborative scheduling instructions generated by S2 and the intelligent decisions generated by S3 (such as quality inspection results and parameter settings) are sent to the physical manufacturing system for execution; the system collects the actual production data after execution. (e.g., actual completion time, actual energy consumption, product yield); Reward Calculation and Model Update: Compare the actual results with the expected results to calculate the reward signal. ; (Planned completion time / Actual completion time) + (1-Defect Rate)- (Actual energy consumption / Baseline energy consumption) - ; in, This is a completion time bonus coefficient, used to reward early or on-time completion. This is a quality incentive coefficient used to reward high yield rates; This is the energy consumption penalty coefficient, used to penalize energy consumption exceeding the benchmark. This is the penalty coefficient for violating physical constraints, used to punish behaviors that violate physical constraints; This represents the total number of physical constraint categories. For the first Points deduction for violations of physical constraints, for example: Exceeding processing force limit: 0.1 points will be deducted for every 10N exceeding the limit; Heat deformation exceeding tolerance: 0.05 points will be deducted for every 1μm exceeding the tolerance; Agent policy update: utilizing Feedback is provided through multi-agent reinforcement learning algorithms (such as MADDPG and IPPO) to update the various... Decision-making strategy (i.e., its decision-making strategy based on state) To action The mapping function is used to adaptively adjust its utility function weights. and penalty weight : in: Penalize the learning rate of the weights (recommended value: 0.01-0.05); For the most recent Within the first cycle Frequency of class constraint violations; The target violation rate threshold (can be set according to process requirements); Routing policy update: based on Optimize the routing policy function in S3 based on the actual completion time and accuracy at different nodes (edge / cloud). and threshold parameters; This closed-loop process enables the framework to continuously adapt to changes in the production environment and achieve continuous self-optimization.

[0037] This invention provides a multi-agent collaborative optimization deployment system based on industrial manufacturing equipment. The deployment system includes: an initialization module for system parameters and agent configuration, a multi-agent collaborative decision modeling and optimization module, a large-model distributed reasoning and task hierarchical routing module, a dynamic adaptive deployment and resource elastic scheduling module, and a closed-loop execution and online learning module. The initialization system parameter and agent configuration module is used to construct a parameterized model of physical laws and a distributed collaborative architecture of multiple agents, thereby realizing the formal modeling of physical constraints and the initialization of the distributed collaborative system, laying a structured foundation for physical perception scheduling and multi-agent collaborative optimization. The multi-agent collaborative decision-making modeling and optimization module is used to construct a utility function that integrates local performance, global indicators and physical constraint penalties, and combine non-cooperative game theory and Nash equilibrium solutions to achieve global collaborative optimization and physical feasibility assurance for agents under autonomous decision-making. The large-scale model distributed inference and task hierarchical routing module is used to achieve intelligent allocation and load balancing of computing tasks through multi-dimensional task feature extraction and dynamic edge-cloud routing strategies, thereby obtaining low-latency and high-precision distributed inference capabilities to support scenarios such as real-time quality inspection and fault diagnosis. The dynamic adaptive deployment and resource elastic scheduling module is used to achieve dynamic optimization of model deployment, computing resources and task scheduling through multi-dimensional evaluation of system performance and dual-trigger adjustment mechanism, so as to obtain an elastic operating system with environmental adaptability and fault tolerance. The closed-loop execution and online learning module is used to continuously optimize the strategy, weights and routing parameters online through feedback from actual production data and multi-agent reinforcement learning, so as to obtain a self-improving intelligent closed loop that can adaptively balance scheduling efficiency and physical constraints.

[0038] Example 1: Intelligent Workshop Scheduling and Quality Collaborative Optimization System This framework is applied in a discrete manufacturing workshop. Four agents are configured: (CNC machine tool group) (Robot assembly line) ( Logistics system) (Automatic warehouse). Deploy a lightweight scheduling model on the edge server of the workshop, and a high-precision visual inspection model on the enterprise private cloud.

[0039] Collaborative scheduling: After a new order arrives, each Based on the current load ( The workpiece path is negotiated based on the S2 game model to quickly generate the initial scheduling sequence, with an average response time of less than 5 seconds.

[0040] Real-time quality inspection: Cameras on the assembly line capture images of key components and generate a... Its characteristics =(Large, High, 300ms). According to the S3 routing policy, this task is directed to the large model in the cloud. Defect detection is performed, and the results are returned within 280ms.

[0041] Dynamic adjustment: during the period A critical machine tool experienced a sudden warning, leading to a surge in related computing requests. S4's dynamic framework detected the load on the edge server. Exceeding the limit, automatically reduce part The real-time path replanning task was temporarily offloaded to the cloud, ensuring that the delay of the core assembly line quality inspection task was not affected.

[0042] Effect: Compared to the original centralized With the system and independent quality inspection stations, the workshop's order delivery cycle was shortened by 22%, the rate of missed quality inspections was reduced by 60%, the response speed to anomalies was increased by 3 times, and the system showed good resilience in dealing with production capacity fluctuations.

[0043] Example of physical constraint fusion optimization: (The CNC machine tool group) monitors changes in cutting force in real time when scheduling machining tasks. It detects when the instantaneous cutting force of a certain process approaches the machine tool's maximum allowable force. At that time, the system automatically adjusts the feed rate to avoid penalties for exceeding force limits. With adaptive adjustment of weights, the number of force over-limit events decreased from 8 per shift initially to 1-2 after stabilization.

[0044] (Robotic assembly line) In precision assembly processes, a thermal deformation compensation mechanism is introduced. The system calculates based on changes in ambient temperature (collected in real time by temperature sensors) and the material's coefficient of thermal expansion. Automatically fine-tunes assembly position to control thermal deformation error within ±5μm, with penalty weights. It adapts to seasonal temperature changes.

[0045] ( In logistics system route planning, the physical impact of load on energy consumption is considered, and an energy consumption model is established. in, for Load mass (unit: kg); for Unloaded mass (unit: kg); Acceleration due to gravity (unit: m / s²) 2 ); for Operating speed (unit: m / s); The coefficient of rolling friction; For transportation distances, the system prioritizes low-friction routes while ensuring timeliness, reducing transportation energy consumption by approximately 18%. Physical fusion effect: After implementing physical constraint fusion optimization, the system has achieved significant improvements in the following aspects: the number of scheduling scheme adjustments due to physical infeasibility has been reduced by about 65%; the abnormal downtime caused by equipment operating beyond limits has been shortened by about 40%; the consistency of processing accuracy (CPK value) has been improved by about 1.2; the physical constraint penalty weight λ tends to stabilize after running for 1000 cycles, indicating that the system has learned the balance point between physical laws and scheduling objectives.

[0046] In summary, compared with the prior art, the present invention has the following beneficial effects: 1. In this embodiment of the invention, the deployment method transforms traditional centralized decision-making into a distributed intelligent collaboration system by constructing a multi-agent collaborative architecture. While maintaining local autonomous decision-making, each agent achieves global optimization of task and resource allocation through real-time information sharing and strategy coordination mechanisms, which significantly improves the system's adaptability and overall efficiency in complex multi-task environments.

[0047] 2. In this embodiment of the invention, the deployment method introduces game theory into the multi-agent interaction process, establishes a decision-making mechanism that takes into account both individual interests and the overall goals of the system, effectively coordinates resource competition and task conflicts among agents, achieves equilibrium optimization in a dynamic environment, and thus improves the rationality and collaborative efficiency of the overall system decision-making.

[0048] 3. In this embodiment of the invention, the deployment method uses an edge-cloud collaborative distributed inference framework to allocate the computational tasks of large-scale models to different nodes as needed. Through dynamic scheduling and allocation of computational load, the inference latency is significantly reduced, and the system's response speed and processing capability in scenarios such as real-time quality control and fault prediction are improved.

[0049] 4. In this embodiment of the invention, the deployment method, through a dynamic deployment framework with environmental awareness, can automatically optimize the model deployment strategy and computation path based on real-time hardware resources, network status and task requirements, giving the system good scalability and environmental adaptability, and ensuring efficient and stable operation in different manufacturing scenarios.

[0050] 5. In this embodiment of the invention, the deployment method integrates a multi-level scheduling method that combines global optimization and local autonomy to achieve refined management and control of production resources. Through the organic combination of agent collaboration and model reasoning, the resource load of each link is dynamically balanced, effectively avoiding resource idleness or overload, and comprehensively improving the operating efficiency and stability of the production line.

[0051] 6. In this embodiment of the invention, the deployment method is based on the efficient reasoning of large models and the rapid response of multiple agents to construct a closed-loop control system with real-time decision-making capabilities. It can adjust task allocation and process parameters in a timely manner according to changes in production status, and ensure the global optimality of decision-making through game theory, thereby significantly improving the level of manufacturing intelligence.

[0052] 7. In this embodiment of the invention, the deployment method, through a distributed architecture and a multi-agent collaboration mechanism, naturally possesses the ability to isolate faults and reconfigure functions. When some nodes or agents malfunction, the system can automatically migrate tasks and compensate for functions, ensuring continuous and stable operation and effectively improving the reliability of the system in actual industrial environments.

[0053] 8. In this embodiment of the invention, the deployment method parameterizes natural laws such as gravity, friction, thermodynamics, and materials mechanics and integrates them into the agent's utility function, constructing a physically constrained optimization model. This model not only focuses on traditional scheduling indicators (time, efficiency) but also evaluates the physical feasibility of the scheduling scheme in real time, using penalty weights. Its adaptive adjustment mechanism can learn the importance of physical constraints during operation, dynamically balance scheduling efficiency and physical feasibility, and ensure that the optimization results are highly executable and reliable in actual manufacturing environments.

[0054] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0055] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-agent collaborative optimization deployment method based on industrial manufacturing equipment, characterized in that, The deployment method includes: S1. Initialize system parameters and agent configuration Construct a parameterized model of physical laws and a multi-agent distributed collaborative architecture to realize formal modeling of physical constraints and initialization of the distributed collaborative system; S2. Multi-agent cooperative decision-making modeling and optimization We construct a utility function that integrates local performance, global metrics, and physical constraint penalties, and combine non-cooperative game theory and Nash equilibrium solutions to achieve global collaborative optimization and physical feasibility assurance for agents under autonomous decision-making. S3. Large-Model Distributed Inference and Hierarchical Task Routing For complex sub-problems or real-time triggered intelligent tasks generated during collaborative decision-making, they are dynamically routed to edge computing nodes or cloud computing nodes according to task characteristics, and the corresponding large models are called for distributed reasoning to achieve intelligent allocation and load balancing of computing tasks, thereby obtaining low-latency and high-precision distributed reasoning capabilities. S4. Dynamic Adaptive Deployment and Elastic Resource Scheduling By using a multi-dimensional evaluation of system performance and a dual-trigger adjustment mechanism, dynamic optimization of model deployment, computing resources and task scheduling is achieved, resulting in a resilient operating system with environmental adaptability and fault tolerance. S5. Closed-Loop Execution and Online Learning By leveraging feedback from actual production data and multi-agent reinforcement learning, continuous online optimization of strategies, weights, and routing parameters is achieved, resulting in a self-improving intelligent closed loop that can adaptively balance scheduling efficiency and physical constraints.

2. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, The system parameters in S1 include: multi-agent parameters; The multi-agent parameters include: Number of agents ; All intelligent agents constitute a set Each intelligent agent Responsible for a specific manufacturing unit or resource pool; Each agent Its state space is constituted by local observation information. , This includes the queue length of its managed resources, equipment utilization rate, and information on currently processed workpieces; Each agent All executable decision operations constitute its action space. , Including the selection of workpieces to be processed Adjust the process parameters of the controlled equipment. Sending cooperation requests to neighboring units, etc.; The communication topology between intelligent agents is defined as a graph. , where vertex set edge set This refers to a communication link that allows direct information exchange between intelligent agents.

3. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, The system parameters in S1 also include: large model deployment parameters; The large model deployment parameters include: A large-scale global decision-making model deployed in the cloud or on a central server is denoted as Its parameter count is ; Deployed at the edge computing nodes The lightweight model is denoted as Its parameter count is , usually satisfy ; Define performance constraints for model inference: minimum accuracy requirement and maximum allowed end-to-end delay .

4. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, The system parameters in S1 also include: manufacturing task and resource parameters; The manufacturing tasks and resource parameters include: Workpieces to be scheduled , This represents the total number of workpieces. For the workpiece The process sequence is denoted as , For workpiece The total number of processes; Available physical manufacturing resources of the system , Total number of resources; Define three-dimensional parameters Indicates workpiece The Steps In resources The time required for processing; among which, ; ; If a certain process Unable to access resources For processing, then it is agreed Or use a maximum value express.

5. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, The system parameters in S1 also include: physical process constraint parameters; The physical process constraint parameters include: Each manufacturing resource Having a set of physical properties Each constraint corresponds to a natural law-based limitation in the manufacturing process: Machining mechanical constraints: Maximum permissible processing force ; Maximum depth of cut ; Thermodynamic constraints: Coefficient of thermal expansion of material ; Maximum allowable temperature rise ; Dynamic constraints: Maximum permissible vibration frequency ; Maximum permissible acceleration ; Define the penalty weight vector ; for Load quality; for Unloaded mass; It is the acceleration due to gravity; for Running speed.

6. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, S2 specifically includes: Each intelligent agent Based on its perceived local state and through communication topology Receive information from neighboring intelligent agents and make distributed collaborative decisions based on a non-cooperative game model; For each intelligent agent Define a local utility function: ; in, For intelligent agents The selected action; To remove The combination of actions chosen by all other intelligent agents; After the action is executed, the intelligent agent Average workpiece waiting time for the resources under its jurisdiction; This refers to the maximum completion time of the entire manufacturing system after the action is executed; , For workpiece Completion time; For action Perform the estimated resource and energy consumption; The core scheduling target weight coefficient; This is the physical constraint penalty weight coefficient, used to balance scheduling efficiency and physical feasibility; This represents the total number of physical constraint categories. For the first A penalty function for violating physical constraints, based on the physical state of the resources under the agent's control. calculate; The initial weights satisfy: ; The weighting coefficients are dynamically adjusted through subsequent feedback learning to balance multi-objective optimization between scheduling efficiency, energy consumption, and physical feasibility; The system employs a distributed iterative algorithm to find an approximate Nash equilibrium for the game. ); At this equilibrium point, for each agent ,satisfy: 。 7. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, S3 specifically includes: For complex sub-problems arising from game-theoretic decisions in S2 or intelligent tasks triggered in real time by sensors, the large model inference module is invoked. Task modeling and classification: Abstract each decision request into a reasoning task. And label its feature vector: ; in Input the amount of data for the task; Estimate the computational complexity of the task; The maximum tolerable delay for the task; For physical process characteristic sub-vectors; Dynamic routing strategy: The system maintains a routing policy function. The decision-making logic is as follows: IF: AND ; THEN: Routing to the edge layer, by comparing each edge node Current load Select the node with the least load to execute its lightweight model. ; ELSE: will Routing to the cloud, from a global large model implement; Edge layer load balancing: Define edge nodes At any moment Load index: ; in, for exist The length of the pending task queue at any given time; for exist CPU utilization percentage at any given time; This is the upper limit of the queue length; For weight parameters ( ).

8. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, S4 specifically includes: Continuously monitor the overall status and dynamically adjust model deployment and resource allocation strategies; System overall performance evaluation function: Define the overall system performance indicators within the evaluation period. : ; ; in, The total number of manufacturing and reasoning tasks successfully completed per unit of time; Average end-to-end processing latency for all tasks; This is an estimated total system energy consumption. This is due to the overhead caused by model migration and task rescheduling; A score is given for the consistency between the scheduling results and the physical laws. The normalized weighting coefficients are positive. Adaptive adjustment of triggering and execution: Periodic triggering: every fixed time interval Calculate the current If the decline compared to the previous cycle exceeds the threshold If so, optimization will be triggered; Event trigger: When the load of any node is detected Node failure, or network bandwidth consistently below [a certain value] Triggered immediately upon arrival; Adjustment Actions: Adjustment decisions are based on a lightweight optimizer or policy network, and the actions include: Model transfer: transferring a marginal model Migrate from an overloaded node to a lightly loaded node; Calculate unloading: Remove excessively long queues from the edge layer. Unload to the cloud for processing in real time; Policy Update: Fine-tuning the utility function weights of agents in S2 Or the routing threshold in S3 , .

9. The multi-agent collaborative optimization deployment method based on industrial manufacturing equipment as described in claim 1, characterized in that, S5 specifically includes: Execution and Feedback: The collaborative scheduling instructions generated by S2 and the intelligent decisions generated by S3 are sent to the physical manufacturing system for execution; the system collects the actual production data after execution. ; Reward Calculation and Model Update: Compare the actual results with the expected results to calculate the reward signal. ; (Planned completion time / Actual completion time) + (1-Defect Rate)- (Actual energy consumption / Baseline energy consumption) - ; in, This is a completion time bonus coefficient, used to reward early or on-time completion. This is a quality incentive coefficient used to reward high yield rates; This is the energy consumption penalty coefficient, used to penalize energy consumption exceeding the benchmark. This is the penalty coefficient for violating physical constraints, used to punish behaviors that violate physical constraints; This represents the total number of physical constraint categories. For the first Points deducted for violating physical constraints; Agent policy update: utilizing Feedback is used to update the various agents through a multi-agent reinforcement learning algorithm. The decision-making strategy, and adaptively adjust its utility function weights. and penalty weight : ; in: To penalize the learning rate of weights; For the most recent Within the first cycle Frequency of class constraint violations; The target violation rate threshold; Routing policy update: based on Optimize the routing strategy function in S3 based on the actual completion time and accuracy at different nodes. and threshold parameters.

10. A multi-agent collaborative optimization deployment system based on industrial manufacturing equipment, characterized in that, The deployment system includes: a module for initializing system parameters and agent configuration, a module for multi-agent collaborative decision-making modeling and optimization, a module for large-scale model distributed reasoning and hierarchical task routing, a module for dynamic adaptive deployment and elastic resource scheduling, and a module for closed-loop execution and online learning. The initialization system parameter and agent configuration module is used to construct a physical law parameterized model and a multi-agent distributed collaborative architecture to realize the formal modeling of physical constraints and the initialization of the distributed collaborative system, laying a structured foundation for physical perception scheduling and multi-agent collaborative optimization. The multi-agent collaborative decision-making modeling and optimization module is used to construct a utility function that integrates local performance, global indicators and physical constraint penalties, and combine non-cooperative game theory and Nash equilibrium solutions to achieve global collaborative optimization and physical feasibility assurance for agents under autonomous decision-making. The large-scale model distributed inference and task hierarchical routing module is used to achieve intelligent allocation and load balancing of computing tasks through multi-dimensional task feature extraction and dynamic edge-cloud routing strategies, thereby obtaining low-latency and high-precision distributed inference capabilities to support scenarios such as real-time quality inspection and fault diagnosis. The dynamic adaptive deployment and resource elastic scheduling module is used to achieve dynamic optimization of model deployment, computing resources and task scheduling through multi-dimensional evaluation of system performance and dual-trigger adjustment mechanism, so as to obtain an elastic operating system with environmental adaptability and fault tolerance. The closed-loop execution and online learning module is used to continuously optimize the strategy, weights and routing parameters online through feedback from actual production data and multi-agent reinforcement learning, so as to obtain a self-improving intelligent closed loop that can adaptively balance scheduling efficiency and physical constraints.