Production control method and system based on multi-role embodied agent in industrial field
By constructing a production control method for multi-role embodied intelligent agents in the industrial field, and combining it with a causal chain reasoning consistency adaptive reinforcement correction algorithm, the problem of insufficient intelligence in industrial production control systems is solved, and multi-role collaborative decision-making and highly reliable control are realized.
Patent Information
- Application Number
- CN202511604942.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-05
- Publication Date
- 2026-02-13
- Estimated Expiration
- 2045-11-05
AI Technical Summary
The existing industrial production control systems lack sufficient intelligence, and general-purpose large language models are difficult to apply directly to the industrial field. There is a lack of effective domain knowledge integration and multi-role collaboration mechanisms.
This paper proposes a production control method based on multi-role embodied intelligent agents in the industrial field. By constructing a domain-specific language model and a multi-role intelligent control system, and combining a causal chain reasoning consistency adaptive reinforcement correction algorithm, intelligent monitoring, analysis, decision-making and control of industrial production processes can be achieved.
It has improved the intelligence level and control effect of industrial production control systems, supported multi-role collaborative decision-making, and has real-time response and high-reliability control capabilities.
Smart Images

Figure CN121050398B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of industrial intelligent control and artificial intelligence technology, and in particular to a production control method and system based on a multi-role embodied intelligent agent in the industrial field. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, large language models represented by GPT, Claude, DeepSeek, etc. have shown unprecedented capabilities in natural language processing, code generation, logical reasoning, etc. At the same time, the in-depth promotion of Industry 4.0 and intelligent manufacturing puts forward higher intelligent requirements for production control systems, and traditional control methods based on rules and mathematical models face challenges such as difficult knowledge acquisition, insufficient adaptability, high maintenance cost, etc.
[0003] Traditional industrial production control mainly relies on distributed control systems (DCS), programmable logic controllers (PLC), etc. automation equipment, combined with pre-set control algorithms and expert experience to formulate control strategies. However, existing industrial control systems have technical limitations: control strategies are based on fixed rules, making it difficult to handle complex and variable production conditions; knowledge acquisition and updating rely on manual experts, with high cost and low efficiency; the system has limited self-adaptability, and needs to redevelop and debug algorithms when facing new scenarios.
[0004] In recent years, AI technologies such as machine learning and deep learning have been introduced into industrial production control, such as neural network predictive control and reinforcement learning adaptive control. However, these methods are usually designed for specific problems and lack universality and scalability. The professional knowledge in the industrial field is highly domain-specific, and general AI models are difficult to directly apply to specific industrial production scenarios.
[0005] Chinese patent document with publication number CN113379135A discloses an intelligent production line product quality prediction method based on cloud-edge collaborative computing, mainly focusing on quality prediction and not involving real-time control decision-making in the production process. Patent document with publication number CN113219922A discloses an integrated control method based on edge-side data, still based on traditional mathematical optimization methods, lacking intelligent decision-making capabilities.
[0006] Large language models have strong language understanding, logical reasoning, knowledge integration and code generation capabilities, and can process multi-modal data, which is highly consistent with the needs of industrial production control. However, there are challenges in directly applying general large language models to industrial control: general models lack industrial domain expertise; industrial control has high real-time and reliability requirements; industrial production involves complex collaboration among multiple roles, and a single model cannot meet the requirements; and there is a lack of effective industrial knowledge injection and model optimization methods.
[0007] Embodied intelligence emphasizes the interactive perception ability of agents and the environment, which in the industrial context is reflected in intelligent systems that can perceive the production environment, understand the production state, make control decisions and perform operations. Existing research on industrial embodied intelligence mainly focuses on the control level of robots, and lacks technical solutions that deeply integrate the cognitive capabilities of large language models with industrial embodied control.
[0008] Retrieval augmented generation techniques enhance model generation capabilities by dynamically retrieving external knowledge bases, and low-rank adaptive techniques provide solutions for efficient fine-tuning of large language models. These developments lay the technical foundation for building domain-specific industrial language models.
[0009] In view of the above technical status, there is an urgent need in the field of industrial production control for a technical solution that can fully utilize the advantages of large language models, combine industrial domain knowledge, and achieve intelligent production control. This solution should effectively integrate industrial expertise, implement a multi-role collaborative intelligent control system, and have real-time response and reliable control capabilities. SUMMARY
[0010] The present application provides a production control method and system based on industrial domain multi-role embodied agents, which can quickly integrate industrial expertise, support multi-role collaborative decision-making, and have real-time, self-learning and high-reliability control capabilities, thereby improving production efficiency, product quality and operational safety.
[0011] The technical solution of the present application is as follows:
[0012] A production control method based on industrial domain multi-role embodied agents, comprising a domain role agent construction phase and an industrial embodied production control execution phase;
[0013] The domain role agent construction phase includes processing production data and technical materials of the target production industry, constructing an industry domain database and constructing a production process causal knowledge graph; decomposing industrial embodied production control tasks into monitoring analysis and control operations, fine-tuning a large language model based on the industry domain database to obtain a monitoring analysis model and a control operation model, and driving a monitoring analysis embodied agent and a control operation embodied agent with the monitoring analysis model and the control operation model, respectively;
[0014] The industrial body production control execution stage includes: monitoring and analyzing the production process causal knowledge graph based on real-time production data to retrieve production process causal knowledge graph, generating chain thinking and control targets, performing consistency discrimination on the chain thinking and causal chain, submitting the control targets to the control operation body agent, and outputting control instructions that meet the hard constraints based on the causal chain consistency-benefit joint reward strategy and closed-loop execution; the execution results are fed back for incremental fine-tuning.
[0015] Preferably, the domain role agent construction stage includes:
[0016] (1-1) Preprocessing and semantic annotation of structured process parameters, unstructured technical data and historical production control records of the target production industry, and establishing an industry domain database through hierarchical storage and semantic indexing mechanism;
[0017] According to the causal relationship of "equipment-process parameters-working conditions-quality / energy consumption indicators", a production process causal knowledge graph (Causal Knowledge Graph, C-KG) is constructed;
[0018] (1-2) Fine-tune the base large language model (Base Large Language Model, B-LLM) on the industry domain database to form a basic domain language model (Domain LLM, D-LLM) ;
[0019] (1-3) Mark the working condition characteristics of historical production data, and associate expert control targets and expert prompts under different working conditions to form training database A;
[0020] (1-4) Integrate production process simulation data and historical production data to generate optimal control strategies and feasible control schemes under different control targets to form training database B;
[0021] (1-5) Fine-tune the basic domain language model using training database A and training database B respectively to obtain a monitoring and analysis model and a control operation model ;
[0022] The monitoring and analysis model and the control operation model drive the monitoring and analysis body agent and the control operation body agent respectively.
[0023] Preferably, in the industrial body production control execution stage, the industrial body ontology combines the monitoring and analysis body agent and the control operation body agent to construct the following workflow and apply it to the production control process:
[0024] (2-1) Real-time acquisition of production state vector, alarm text and plan information, and formation of observation data after screening and filtering ;
[0025] (2-2) The monitoring and analysis embodied intelligent agent retrieves the causal chain according to the causal knowledge graph , and outputs chain thinking and control target ;
[0026] (2-3) The consistency of and is scored by the causal chain consistency discriminator, and the consistency score is obtained , when is higher than the threshold value, the control target is submitted to the control operation intelligent agent, and when is lower than the threshold value, a negative reward is given and the sampling is returned;
[0027] (2-4) The control operation embodied intelligent agent reasons the action (i.e. control set point vector) according to the control target ;
[0028] (2-5) The verified is written into the control system for execution, and the execution result is stored for online adaptive fine-tuning.
[0029] The present application aims at the technical problems of insufficient intelligence of industrial production control system in the prior art, difficulty of directly applying general large language model to industrial field, lack of effective field knowledge integration and multi-role cooperation mechanism, etc., and provides a production control method based on industrial field multi-role embodied intelligent agent, which realizes intelligent monitoring, analysis, decision and control of industrial production process by constructing field special language model and multi-role intelligent control system, and introducing causal chain reasoning consistency adaptive reinforcement correction algorithm, and can improve the intelligent level and control effect of the production control system.
[0030] The embodied intelligent agent construction stage of the field role focuses on offline data preparation, industry knowledge injection and model fine-tuning; the embodied production control execution stage focuses on online chain reasoning, causal consistency verification, reinforcement correction and closed-loop control.
[0031] The structured field knowledge includes process parameter table, equipment specification book, operation procedure, instrument loop list and other standardized technical documents; the unstructured field data includes professional literature, technical report, fault case and other non-standard information resources; the historical production data and control records include operator operation records, alarm logs, etc.
[0032] In step (1-1), the preprocessing package includes data cleaning, desensitization, time synchronization, and multi-modal alignment; the semantic annotation includes labeling of working condition categories, control targets, and operation intentions. Through preprocessing and semantic annotation, high-quality supervised data is provided for subsequent model training.
[0033] In step (1-1), the industry domain database includes:
[0034] A time series database (TS-DB) stores historical production data, operation record data, early warning data, and production response data according to time labels.
[0035] A knowledge graph database (KG-DB) uses a graph database to store entity-relation-event tuples and record process knowledge.
[0036] A vector database (Vec-DB) saves high-dimensional embedding vectors of field professional terms, drawings, reports, and images, as well as high-dimensional embedding vectors of Internet public text data, industry reports, and professional literature.
[0037] In step (1-2), external knowledge retrieval enhancement technology is used to dynamically recall relevant entries in the industry domain database based on query vectors to expand the context of the language model; through parameter efficient fine-tuning (PEFT) technology, a trainable adaptation module is inserted at the attention layer, so that the proportion of trainable parameters in fine-tuning does not exceed 2%.
[0038] The industrial embodied production control task is divided into two roles: monitoring analysis engineer (A-Agent) and control operation engineer (C-Agent). A-Agent is responsible for retrieving C-KG causal chains based on real-time production data and generating chain thinking and control targets. C-Agent receives control targets verified by a consistency discriminator and outputs control setting values that meet hard constraints based on causal chain-of-thought consistency reinforced calibration (C 3 RC) joint reward strategy.
[0039] In step (1-3), based on historical production data, working condition features are labeled and associated with corresponding expert control targets and expert prompts to construct a monitoring analysis role language model training data set.
[0040] The working condition characteristics include the numerical range, trend of change, abnormal pattern, and other state information of the production parameters, and the expert control target includes the control strategy to be taken and the control effect to be achieved for different working condition states. By associating and labeling the working condition characteristics with the expert control target, a training sample of the monitoring analysis task is formed, so that the model learns the analysis ideas and decision logic of the experts under different production conditions. The expert prompt is natural language or semi-structured guidance information written by a domain expert, which is used to constrain the domain language model The focus points, reference data range, and output format in the reasoning process are highlighted, so as to improve the accuracy and interpretability of the monitoring analysis results.
[0041] In step (1-4), the integrated production process simulation data and historical production data are used to generate optimal control strategies and feasible control schemes for different control targets, and a control operation role language model is constructed.
[0042] The production process simulation data and historical production data are of minute-level granularity; the control targets include multiple targets such as minimizing energy consumption, minimizing product quality fluctuation, maximizing production efficiency, and minimizing safety risk. The optimal control strategy is calculated by a nonlinear programming or model predictive control algorithm, representing the best control scheme in theory; the feasible control scheme is based on expert experience and meets the physical constraints and process safety constraints of the equipment, representing the actually executable control selection. Through the combination of simulation data and historical data, a wider range of production scenarios and control situations can be covered.
[0043] The core fields of the training database A include working condition characteristics, expert control targets, and expert prompts; the core fields of the training database B include target categories, optimal control quantities, feasible control quantities, safety constraints, and expert evaluations.
[0044] In step (1-5), the training database A and the training database B are used to supervise the fine-tuning (SFT) of the basic domain language model , the human feedback reinforcement learning (RLHF), and the causal chain-of-thought consistency reinforced calibration (C 3 RC) to obtain the monitoring analysis model and the control operation model .
[0045] Preferably, in step (1-5), the monitoring analysis embodied agent and the control operation embodied agent use a multi-role adaptive reinforcement correction algorithm based on causal chain-of-thought consistency for collaborative optimization, including:
[0046] defining a causal chain of length causal chain , chain consistency discrimination task, i.e., judging whether the language model generates a chain of thought (CoT) sequence is a topologically feasible path in the production process causal knowledge graph; the multi-role adaptive reinforcement correction algorithm process is:
[0047] a) Given the current observation , the monitoring analysis embodied agent retrieves the matching causal chain in the production process causal knowledge graph through graph vector indexing , the model outputs a chain of thought sequence and a control target description ;
[0048] b) The discriminator conducts consistency scoring on and ; ;
[0049] c) Construct a joint reward of causal consistency, benefit and constraint for policy update of the control operation embodied agent, where is the consistency threshold, is the improvement amount of the th benefit index, is the hard constraint penalty, is the weight coefficient;
[0050] d) Update the parameters of the model through the clipped proximal policy optimization (PPO) algorithm iteration until the expected discounted return converges; parameter update is performed offline or at a low frequency, and only the forward inference reinforcement learning agent strategy is generated in the online stage to generate agent actions (set point vectors) .
[0051] The benefit score in the reward function includes the energy consumption difference, product quality fluctuation difference and production rhythm improvement difference; the hard constraint penalty covers device physical constraints, process safety constraints and environmental constraints.
[0052] The industrial embodiment ontology is not limited to referring to the control room controller and control system, but also includes other forms of individuals, systems and organizations that have industrial data analysis, decision-making and execution capabilities.
[0053] The online adaptive fine-tuning adopts an incremental LoRA strategy based on priority experience replay, which is performed according to The size of the sample determines the sampling probability of the trajectory sample, enabling rapid absorption of new operating condition knowledge without forgetting historical strategies.
[0054] This invention also provides a production control system based on multi-role embodied intelligent agents in the industrial field, including an embodied entity unit, a controller, and a database, and performs domain role intelligent agent construction and industrial embodied production control execution according to the production control method described above.
[0055] The embodied body unit is equipped with multimodal sensors, mechanical actuators and network interfaces, which acquire real-time production data and transmit it to the controller, and execute actions according to the controller's instructions;
[0056] The controller internally runs a monitoring and analysis model and a control operation model, and integrates a causal consistency reinforcement correction algorithm. It performs intelligent reasoning based on real-time production data to obtain the optimal control parameters and issue control commands to the embodied entity.
[0057] The database includes industry-specific databases to support intelligent reasoning for the controller;
[0058] The controller interacts with the embodied body unit through an industrial communication protocol to realize a closed loop of data acquisition, intelligent reasoning, and control execution.
[0059] For a given industrial production process and data, the aforementioned production control system's domain-specific language model performs production data analysis, control decisions, and operation execution, specifically including:
[0060] Preprocess and filter redundant production data;
[0061] By monitoring and analyzing the filtered production data through the embodied intelligent agent, it is possible to determine whether the operation is abnormal and the corresponding control objectives.
[0062] Given a control objective, the control action embodied intelligent agent calculates new control setpoints, including control variables and control points;
[0063] Based on the updated control instructions, the control points are automatically modified and executed through the process logic controller or distributed control system.
[0064] After the embodied entity unit executes the instruction, it sends back an execution confirmation. The controller writes the four-tuple of "operating condition, causal chain, control quantity, and execution effect" into the database unit for subsequent online learning.
[0065] Preferably, the production control system further comprises a human-computer interaction unit, which provides a visual monitoring interface and an operation interface, including interactive devices such as a touch display screen, a voice interaction module, and an alarm indicator light, to support the monitoring of the state of the production control system and necessary manual intervention by an operator; meanwhile, the operation safety and traceability are ensured through an account permission grading and audit log mechanism.
[0066] Compared with the prior art, the present application has the following beneficial effects:
[0067] The present application deeply applies large language model technology to the field of industrial production control, constructs an intelligent language model specially for industrial control through domain knowledge injection and role-based training, and realizes the deep integration of artificial intelligence technology and industrial control.
[0068] The multi-role language model cooperation mechanism proposed in the present application realizes the intelligent decomposition and collaborative execution of complex industrial control tasks through the division of labor between the monitoring and analysis engineer role and the control operation engineer role, and significantly improves the intelligent level and decision quality of the control system.
[0069] The domain language model described in the present application combines edge feedback and expert evaluation, continuously absorbs new knowledge and control results during operation, and constantly self-corrects and evolves to form a high-specialized intelligent agent oriented to device characteristics.
[0070] The industrial embodied control system constructed in the present application organically combines the cognitive ability of the language model and the execution ability of the industrial control, and realizes a complete closed loop from data perception, intelligent analysis, decision calculation to control execution. BRIEF DESCRIPTION OF DRAWINGS
[0071] Figure 1 is a schematic diagram of the domain role intelligent agent construction stage of the production control method based on the multi-role embodied intelligent agent in the industrial field;
[0072] Figure 2 is a schematic diagram of the industrial embodied production control execution stage of the production control method based on the multi-role embodied intelligent agent in the industrial field;
[0073] Figure 3 is a schematic diagram of the hardware topology structure of the production control device based on the multi-role embodied intelligent agent in the industrial field;
[0074] Figure 4 is a schematic diagram of the software deployment architecture of the production control device based on the multi-role embodied intelligent agent in the industrial field. DETAILED DESCRIPTION
[0075] The present application will be further described in detail below in conjunction with the drawings and examples, it should be pointed out that the following examples are intended to facilitate the understanding of the present application and do not have any limiting effect on it.
[0076] A production control method based on multi-role embodied agents in the industrial field, including two main stages: (i) field role embodied agent construction stage, focusing on offline data preparation, industry knowledge injection and model fine-tuning; (ii) embodied production control execution stage, focusing on online chain reasoning, causal consistency verification, reinforcement correction and closed-loop control.
[0077] The field role embodied agent construction stage sequentially performs six sub-steps of data collection and labeling, industry field database construction, basic model fine-tuning, causal knowledge graph construction, role instruction data generation, and role model training and evaluation, as follows.
[0078] First, according to the process characteristics and control requirements of the specified production and manufacturing industry, three types of data sources are gathered: structured field knowledge, unstructured field materials, and historical production data and control operation records, to build an industry field database. Structured field knowledge includes process parameter tables, equipment specification books, operation procedures, instrument loop lists, and other standardized technical documents; unstructured field materials include professional literature, technical reports, fault cases, and other non-standard information resources; historical production data and control operation records include operator operation records, alarm logs, and other information. The collected data is cleaned, desensitized, time-synchronized, and multi-modal aligned; semantic labels such as working condition categories, control targets, and operation intentions are completed to provide high-quality supervised data for subsequent model training.
[0079] Second, an industry field database is constructed through hierarchical storage and semantic indexing mechanisms, including: an operation time series database (TS-DB) for storing cleaned control operation records by time label; a field knowledge graph database (KG-DB) for storing entity-relation-event tuples using a graph database; and a multi-modal document vector database (Vec-DB) for encoding drawings, reports, and pictures into vectors to support semantic retrieval and nearest neighbor recall.
[0080] Then, based on a multi-modal pre-training model with a parameter size greater than 50 B, a combination strategy of retrieval augmented generation (RAG) and low-rank adaptive fine-tuning (LoRA) is used for parameter fine-tuning, and a production process causal knowledge graph (C-KG) is constructed according to the causal relationship of "equipment-process parameters-working conditions-quality / energy consumption indicators", outputting a basic field language model with industry knowledge internalization and causal retrieval capability. RAG uses KG-DB, Vec-DB, and TS-DB as external knowledge indexes, and dynamically injects context according to query vectors during model reasoning steps; LoRA inserts a trainable low-rank matrix in the attention layer of the pre-training model, freezes the main parameters of the original model, and fine-tunes no more than 2% of the weights to integrate professional knowledge.
[0081] The industrial body production control task is divided into two roles of monitoring analysis agent (A-Agent) and control operation agent (C-Agent), A-Agent is responsible for retrieving C-KG causal chain based on real-time production data and generating chain thinking and control target, C-Agent receives the control target verified by a consistency discriminator, and generates control setting value according to C 3 RC joint reward policy outputs control setting value meeting the hard constraint.
[0082] Based on historical production data, the working condition characteristics are labeled and the corresponding expert control target and expert prompt are associated to construct the monitoring analysis role language model training data set. The working condition characteristics include the numerical range, trend and abnormal mode of the production parameters, and the expert control target includes the control strategy and expected control effect for different working condition states. By associating and labeling the working condition characteristics and the expert control target, the training sample of the monitoring analysis task is formed, so that the model learns the analysis ideas and decision logic of the experts under different production conditions; the expert prompt is the natural language or semi-structured guidance information written by the domain expert, which is used to constrain the focus points, reference data range and output format of R-LLM in the reasoning process, so as to improve the accuracy and explainability of the monitoring analysis result.
[0083] Integrating production process simulation data and historical production data, the optimal control strategy and feasible control scheme are generated for different control targets to construct the control operation role language model training data set. The control targets include minimizing energy consumption, minimizing product quality fluctuation, maximizing production efficiency, minimizing safety risk and other diversified targets. The optimal control strategy is calculated by nonlinear programming or model predictive control algorithm, which represents the best control scheme in theory; the feasible control scheme is based on expert experience and meets the physical constraints and process safety constraints of equipment, which represents the actually executable control selection. Through the combination of simulation data and historical data, a wider range of production scenarios and control conditions can be covered.
[0084] Finally, based on the deep language model (Deep Learning Language Model, D-LLM), two types of role language models (Role LLM, R-LLM) are fine-tuned respectively, and the corresponding A-Agent and C-Agent are constructed; through supervised fine-tuning (SFT), human feedback reinforcement learning (RLHF) and causal consistency reinforcement correction (C 3 RC), the decision accuracy, explainability and safety are improved together.
[0085] The industrial body production control execution stage realizes a perception-analysis-decision-execution closed loop for real-time production processes, including six sub-steps of data acquisition, data preprocessing, intelligent inference, instruction verification, closed loop execution, and online learning.
[0086] Firstly, process parameters, equipment status, quality indicators, energy consumption indicators, and environmental variables are obtained at a frequency of seconds to minutes through industrial communication protocols and sensor direct connection.
[0087] Secondly, data cleaning, calibration, normalization or standardization, anomaly detection and noise suppression are sequentially completed on the edge computing node, and a fixed-length multi-dimensional tensor is formed using sliding window resampling for R-LLM inference input.
[0088] Thirdly, the output of the previous step is used as the user message of A-Agent, while injecting the latest alarm text, process recipe, production plan, and other auxiliary context. A-Agent combines external knowledge base retrieval results to output three types of key information: current working condition category (normal, transition, abnormal, instability, etc.), target priority and target type (energy saving, quality improvement, yield increase, safety, etc.), and uncertainty score or confidence interval. When the uncertainty score is higher than the preset threshold, the result is automatically adopted; if it is lower than the threshold, an expert review interface is called or it is rolled back to the traditional rule base.
[0089] C-Agent first retrieves historical optimal control points and feasible control point samples, and then generates candidate control schemes using a strategy search combined with a divergence constraint algorithm: the theoretical optimal scheme is calculated by a multi-objective optimization model, and the feasible fallback scheme is based on expert experience and ensures to meet the hard constraints through strategy projection. Subsequently, risk assessment and sorting of candidate schemes are performed according to safety probability, expected return, and steady-state time. When the difference between the theoretical optimal scheme and the feasible fallback scheme exceeds the threshold, the system preferentially selects the feasible scheme with lower risk.
[0090] Then, C-Agent performs logical consistency and out-of-limit testing on the candidate instructions based on safety rule baselines; if the testing is passed, the set values are written in real time to the actuators (valves, frequency converters, mechanical arms, etc.) through DCS or PLC to form a control closed loop; if the testing fails, the system automatically switches to the last version of feasible set values and generates an abnormal event record.
[0091] Finally, after control execution, the system continuously monitors key indicators such as energy consumption, product quality, production rhythm, and safety risk; when the deviation of any indicator exceeds the set tolerance interval, it immediately rolls back to the last feasible set value and records the event. The newly generated data pairs (working condition, control target, execution effect, expert feedback) are automatically written into the training question and answer database, and the system performs incremental fine-tuning at a set period (e.g., daily) to realize continuous learning.
[0092] The application also provides a production control device based on a multi-role embodied intelligent agent in an industrial field, which comprises a data acquisition unit, a data processing unit, a language model reasoning unit, a control execution unit, a data storage unit and a human-machine interaction (HMI) unit.
[0093] The data acquisition unit is responsible for real-time acquisition of industrial production process data, including industrial Ethernet interface, serial communication interface, analog input module and other hardware interfaces, supporting connection with various industrial devices. The data processing unit comprises an edge computing processor and a data preprocessing algorithm module, which is used for completing time synchronization, exception elimination, data compression and filtering operations.
[0094] The language model reasoning unit comprises a monitoring and analysis engineer role language model and a control operation engineer role language model, and is responsible for performing intelligent analysis and decision calculation tasks. The control execution unit is responsible for sending control instructions to a distributed control system, so as to ensure that the control decisions can be accurately and timely executed to production equipment. The data storage unit comprises a real-time database, a domain knowledge base, a training database and other multi-level storage architectures, and supports efficient storage and retrieval of different types of data.
[0095] The human-machine interaction unit provides a visual monitoring interface and an operation interface, and comprises interactive devices such as a touch display screen, a voice interaction module and an alarm indicator, supports monitoring of system states and necessary manual intervention by an operator, and at the same time guarantees operation safety and traceability through account permission grading and audit log mechanisms.
[0096] The control flow of the device adopts complete closed-loop control logic: a) the data acquisition unit acquires real-time production data and transmits the data to the data processing unit; b) the data processing unit completes preprocessing and quality checking and sends the results to the language model reasoning unit; c) the monitoring and analysis role language model judges the working condition state and gives a control target; d) the control operation role language model calculates optimal control parameters; e) the control execution unit issues control instructions to a distributed control system and drives an execution mechanism; f) the human-machine interaction unit displays control states and system running information in real time, and pushes an alarm when an exception occurs; g) execution results and feedback data are written into the data storage unit, providing support for subsequent online learning and model incremental updating.
[0097] In the embodiment of the application, an automobile water pump production line is selected as a scene, the production line is composed of a rough machining area, a finishing area, an assembly area and a test packaging area, typical processes involve shell machining, bearing press fitting, impeller fastening, sealing detection and other multiple processes, and the production line has the characteristics of multiple equipment types, complex process parameters and multi-dimensional quality control indicators.
[0098] For example, as shown in FIG. 1, the production line comprises a rough machining area 1, a finishing area 2, an assembly area 3 and a test packaging area 4. Figure 1As shown, the domain role agent construction phase of the field language model driven industrial embodied production control method of the present application is divided into six steps: data collection and cleaning, knowledge base construction, basic model injection, role instruction data generation, role model fine-tuning, and offline evaluation.
[0099] Step S101: Data collection and cleaning. The embodiment deploys a data mirroring service in the manufacturing execution system, equipment state monitoring system, and quality traceability system of the production line, and pulls historical raw data at regular intervals; at the same time, it obtains shell machining work instructions, assembly process cards, equipment maintenance manuals, and other PDF / Word documents from the process engineering department. The text is formatted, watermarked, and standardized by scripts, and the process data is time-aligned, interpolated for missing values, and outliers are removed to form structured CSV and JSON files.
[0100] Step S102: Knowledge base construction. The cleaned data is written into three types of storage according to the hierarchy: first, a time series database TS-DB is used to store process parameters, equipment vibration signals, and energy consumption curves; second, a knowledge graph database KG-DB is used to establish relationship nodes in the form of "equipment-process-defect-cause-effect rule"; third, a vector database Vec-DB is used to obtain high-dimensional vectors by encoding drawings, pictures, and reports, facilitating subsequent approximate nearest neighbor search.
[0101] Step S103: Basic model injection. The embodiment selects a multi-modal large language model with a parameter size of 65B as the basic large language model (B-LLM). LoRA is inserted into the multi-layer attention module to insert a low-rank trainable matrix, freeze 98% of the original model parameters, and only train 2% of the adjustable weights. The RAG framework is used to configure the database retrieval tool chain.
[0102] Step S104: Role instruction data generation. According to the responsibilities of the monitoring and analysis role A-Agent and the control operation role C-Agent, multiple rounds of dialogues are generated from historical data pairs and simulation data using automatic scripts. Each piece of data includes five types of fields: role, input context, expected output, expert annotation, and quality label.
[0103] Step S105: Role model fine-tuning. For A-Agent, SFT is performed first, and then multiple process experts score the model reasoning results, and the PPO (Proximal Policy Optimization) algorithm in RLHF is used to further improve the robustness of the strategy. For C-Agent, a multi-objective optimization knowledge distillation link is added after SFT, which compresses the optimal trajectory generated by the heuristic MPC controller into an instruction template and distills it to the strategy head of C-Agent.
[0104] Step S106: Offline evaluation. Construct multiple test sets respectively to evaluate the A-Agent's working condition identification accuracy and target matching rate, and the C-Agent's hard constraint violation rate and optimality loss rate. If the indicators do not meet the control threshold, return to S104 to resample or supplement expert annotations.
[0105] As shown in FIG. 8, the system realizes minute-level closed loop through the perception-analysis-decision-execution-feedback link during online operation. Figure 2
[0106] Step E201: Real-time data acquisition. The edge acquisition terminal integrates multiple bus protocols such as OPC-UA and Profinet, with a sampling period of 5-60 s; the collected signals include shell drill spindle power, impeller pressure-in pressure curve, shaft end temperature rise, qualified rate cumulative value, CO2 emission, etc.
[0107] Step E202: Data preprocessing. The edge node GPU performs batch normalization, anomaly detection and denoising; then resamples according to a 100s sliding window to generate a fixed-dimensional tensor input buffer area.
[0108] Step E203: A-Agent intelligent reasoning. The tensor and the latest alarm text are sent to the A-Agent together; the model first calls Vec-DB to recall the relevant SOP (Standard Operating Procedure), near three-shift fault cases, etc. according to the similarity, and then generates a JSON format output.
[0109] When the confidence is greater than 0.75, the system automatically enters the decision-making process; otherwise, the platform prompts "please confirm the on-duty engineer" in the visual board pop-up window and allows manual override.
[0110] Step E204: C-Agent intelligent reasoning. C-Agent first queries KG-DB to obtain relevant hard constraints, and then retrieves similar scenarios in the historical optimal control point set to generate candidate solutions. Then apply a multi-objective sorting function to output a sorted list. If the optimal difference of multiple candidate instructions is less than 5% and the hard constraints are all met, select the more energy-efficient solution.
[0111] Step E205: Instruction verification. The generated set value is checked against the rule baseline. If a conflict is found, the C-Agent automatically falls back to the candidate safety solution. All instructions that pass the check are packaged into MQTT messages and sent to the DCS network; the PLC returns an execution confirmation after execution.
[0112] Step E206: Closed-loop execution. The actuator (servo press, temperature control valve, frequency converter, etc.) adjusts its action according to the set value; the feedback signal is synchronized into the TS-DB. The system calculates the control deviation at a fixed interval, and considers it effective when the deviation is less than the threshold for three consecutive times.
[0113] Step E207: Online learning. After the production shift ends, the system extracts the (working condition → control target → control quantity → index trend) four-tuple and adds it to the training question-answer database; the off-peak period automatically triggers incremental LoRA fine-tuning, continuously evolving the domain model.
[0114] Figure 3 The overall hardware topology of the device of the embodiment is shown, wherein: the IT area includes a language model inference unit and a data storage unit; the OT area includes an edge computing gateway, a data acquisition unit, a DCS / PLC control system, and a field actuator; the HMI area includes a multi-modal interaction terminal and an operator console. Each node is connected through a gigabit Ethernet and a redundant fiber ring network, and a bidirectional whitelist communication strategy is configured between the OT and IT areas.
[0115] Figure 4 A software deployment architecture diagram is provided for the device of the embodiment. The edge operating system is responsible for driving layer data acquisition and preprocessing; the model inference service is deployed using a cloud cluster, and the A-Agent and C-Agent run in container form; the knowledge database provides KG-DB, Vec-DB, and TS-DB services; the process orchestration service is responsible for queue scheduling, load balancing, and health checking.
[0116] In addition, the control execution process of the scheme of the present application is double-protected by the rule baseline and the safety fallback mechanism. Even when the A-Agent reasoning confidence is insufficient or the C-Agent solution conflicts with the hard constraints, it can be switched to a conservative solution or manually taken over in time, ensuring the safety of production and the integrity of important equipment.
[0117] The above embodiments describe the technical solutions and advantages of the present application in detail. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the present application. Any modification, supplement, and equivalent replacement within the principle range of the present application should be included in the protection scope of the present application.
Claims
1. A production control method based on multi-role embodied intelligent agents in the industrial field, characterized in that, This includes the domain role intelligent agent construction phase and the industrial embodied production control execution phase; The domain role intelligent agent construction phase includes: processing production data and technical information of the target production industry, constructing an industry domain database and a causal knowledge graph of the production process; decomposing industrial embodied production control tasks into monitoring analysis and control operations; fine-tuning the large language model based on the industry domain database to obtain the monitoring analysis model and the control operation model; and driving the monitoring analysis embodied intelligent agent and the control operation embodied intelligent agent, respectively. The execution phase of industrial embodied production control includes: monitoring and analyzing the embodied intelligent agent to generate chain thinking and control objectives based on real-time production data by retrieving the causal knowledge graph of the production process; performing consistency judgment on the chain thinking and the causal chain; submitting the control objectives to the control operation embodied intelligent agent after passing the consistency judgment and benefit joint reward strategy; the control operation embodied intelligent agent outputs control instructions that meet hard constraints based on the consistency of the causal chain and benefit joint reward strategy and executes them in a closed loop; the execution results are fed back for incremental fine-tuning.
2. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 1, characterized in that, The domain role intelligent agent construction phase includes: (1-1) Preprocess and semantically annotate the structured process parameters, unstructured technical data and historical production control records of the target production industry, and establish an industry domain database through hierarchical storage and semantic indexing mechanism; Construct a causal knowledge graph of the production process based on the causal relationships of "equipment-process parameters-operating conditions-quality / energy consumption indicators"; (1-2) In the basic large language model The above-mentioned industry domain database is used for fine-tuning to form a basic domain language model. ; (1-3) Mark the working condition characteristics of historical production data and associate them with expert control objectives and expert prompts under different working conditions to form training database A; (1-4) Integrate production process simulation data and historical production data to generate optimal control strategies and feasible control schemes under different control objectives, forming a training database B; (1-5) Apply training database A and training database B respectively to the basic domain language model Fine-tuning was performed to obtain the monitoring and analysis model. and control operation model ; Monitoring and analysis model and control operation model They respectively drive the monitoring and analysis embodied intelligent agent and the control and operation embodied intelligent agent.
3. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 1, characterized in that, In the industrial embodied production control execution phase, the industrial embodied entity, combined with the monitoring and analysis embodied intelligence and the control and operation embodied intelligence, constructs the following workflow and applies it to the production control process: (2-1) Real-time acquisition of production status vectors, alarm texts, and planning information, followed by filtering to form observation data. ; (2-2) Monitoring and analysis of embodied intelligent agents combined with the aforementioned causal knowledge graph to retrieve causal chains Output chain thinking With control objectives ; (2-3) Using a causal chain consistency discriminant to... and Perform consistency scoring to obtain a consistency score. ,when When the value exceeds the threshold, the control target will be activated. Submitted to the control operation agent, when When the sample size falls below a threshold, a negative reward is given and resampling is initiated. (2-4) Control operation embodied intelligent agent according to control target Reasoning action ; (2-5) Verification passed The data is written into the control system for execution, and the execution results are stored for online adaptive fine-tuning.
4. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 2, characterized in that, In step (1-1), the industry-specific database includes: A time-series database stores historical production data, operation record data, early warning data, and production response data according to time tags; Knowledge graph databases use graph databases to store entity-relationship-event tuples to record process knowledge; The vector database stores high-dimensional embedded vectors of domain-specific terms, drawings, reports, and images, as well as high-dimensional embedded vectors of publicly available text data, industry reports, and professional documents from the internet.
5. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 2, characterized in that, In steps (1-2), external knowledge retrieval enhancement technology is used to dynamically recall relevant entries in the industry domain database based on the query vector to expand the context of the language model; and a trainable adaptation module is inserted into the attention layer through parameter efficiency fine-tuning technology so that the proportion of fine-tuned trainable parameters does not exceed 2%.
6. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 2, characterized in that, In steps (1-3), the operating condition characteristics include the numerical range, trend of change, and abnormal patterns of production parameters; the expert control objectives include the control strategies to be adopted for different operating conditions and the expected control effects; and the expert prompts are natural language or semi-structured guidance information written by domain experts.
7. The production control method based on a multi-role embodied intelligent agent in the industrial field according to claim 2, characterized in that, In steps (1-4), the control objectives include minimizing energy consumption, minimizing product quality fluctuations, maximizing production efficiency, and minimizing safety risks; the optimal control strategy is calculated through nonlinear programming or model predictive control algorithms; and the feasible control scheme is based on expert experience and meets equipment physical constraints and process safety constraints.
8. A production control system based on a multi-role embodied intelligent agent in the industrial field, characterized in that, Including an embodied entity unit, a controller, and a database, the production control method according to any one of claims 1-7 is used to construct domain role intelligent agents and execute industrial embodied production control. The embodied body unit is equipped with multimodal sensors, mechanical actuators and network interfaces, which acquire real-time production data and transmit it to the controller, and execute actions according to the controller's instructions; The controller internally runs a monitoring and analysis model and a control operation model, and integrates a causal consistency reinforcement correction algorithm. It performs intelligent reasoning based on real-time production data to obtain the optimal control parameters and issue control commands to the embodied entity. The database includes industry-specific databases to support intelligent reasoning for the controller; The controller interacts with the embodied body unit through an industrial communication protocol to realize a closed loop of data acquisition, intelligent reasoning, and control execution.
9. The production control system based on a multi-role embodied intelligent agent in the industrial field according to claim 8, characterized in that, It also includes a human-machine interaction unit, providing a visual monitoring interface and operation interface, including a touch screen, voice interaction module, and alarm indicator lights, supporting operators to monitor the status of the production control system and make necessary manual interventions; at the same time, it ensures operational security and traceability through account permission hierarchy and audit log mechanism.
Citation Information
Patent Citations
Integrated control method based on edge side data under endogenous uncertainty
CN113219922A
Intelligent production line product quality low-delay integrated prediction method and system based on cloud-edge cooperative computing
CN113379135A
Intelligent management system and method based on multi-agent cooperation
CN120746475A
Large model fine tuning method based on causal graph and thinking chain enhancement and related device
CN120781920A