A hand production line digital twin AI decision feedback communication system

CN122883902APending Publication Date: 2026-10-09GUANGDONG WEISS ARTS & CRAFTS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611340651.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-01
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0003]然而,现有基于数字孪生的决策系统在面对生产条件快速变化时,其核心的智能决策能力常显不足,特别是在应对新材料、新工艺的引入时,具体的技术问题在于,当前系统中的人工智能模型通常在历史数据集上训练完成后即固化为静态模块,当产线开始试制一款采用特殊新材质如光敏树脂的手办原型时,材料表面特性如附着力、粗糙度、亲漆性等关键物理属性的变化,导致实际喷涂过程中的流体动力学与成膜机理不同于以往,这使得基于旧有数据训练的静态AI模型所作出的决策迅速失准,无法有效补偿新材质带来的工艺偏差,直接导致试产次品率上升,虽然系统能够感知到质量劣化,但现有技术架构下,从感知异常、触发模型更新到新策略生效的反馈闭环存在显著延迟,其根本原因在于,传统模型缺乏在线实时演化的能力,难以在不停机的前提下安全、高效地吸收少量实时生产数据并完成自我更新;同时,模型训练与在线推理环节通常被割裂设计,缺乏一个能实现低延迟感知、决策与策略下发的一体化管道,因此,如何使数字孪生中的AI具备如熟练技师般的即时学习与适应能力,在少量样本下快速调整策略并实现近乎实时的反馈控制,成为制约该技术在高柔性、高混合度生产中广泛应用的核心瓶颈

Benefits of technology

[0049]本发明的有益效果是:通过实时捕获产线缺陷状态,在数字空间中快速构建并优化控制策略,经鲁棒性验证与迁移风险评估后,以渐进、安全的方式将优化策略融合至物理产线,并在执行过程中利用实时数据持续修正模型,从而在无需停机、不产生实际物料浪费的前提下,实现对手办涂装等精密工艺问题的自主、快速、闭环优化与自适应控制,显著提升了产线应对新材料、新工艺的敏捷性与品质一致性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122883902A_ABST
    Figure CN122883902A_ABST
Patent Text Reader

Abstract

The present application relates to a hand production line digital twin AI decision feedback communication system, in particular to the field of hand production line, through real-time capture of production line defect state, quickly build and optimize control strategy in digital space, after robustness verification and migration risk assessment, in a gradual and safe way, the optimized strategy is fused to the physical production line, and the model is continuously corrected by using real-time data in the execution process, so as to realize the self, fast, closed loop optimization and adaptive control of the precise process problems such as hand coating without shutdown and actual material waste, which significantly improves the agility and quality consistency of the production line in response to new materials and new processes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of figurine production lines, and more specifically, to a digital twin AI decision feedback communication system for figurine production lines. Background Technology

[0002] With the deep integration of intelligent manufacturing and personalized customization, the production of high-end consumer goods such as high-precision figurines places extreme demands on process quality and efficiency. In the coating production line of such products, digital twin technology is regarded as the key to achieving precision control. A typical application scenario is that the production line is equipped with a variety of data acquisition terminals, including high-resolution 3D vision sensors, precision airbrush actuators, and various environmental sensors, and is connected to the manufacturing execution system in real time. These terminals continuously generate multi-source heterogeneous data streams covering the microscopic morphology of the paint surface, equipment operating parameters, environmental conditions, and production order information. The digital twin system aggregates and merges this data in real time to build a dynamic model in virtual space that is synchronously mapped with the physical production line. Based on this model, the embedded artificial intelligence decision-making module analyzes the spraying process in real time. Its core objective is to identify micron-level defects and generate compensation strategies in real time, such as dynamically adjusting parameters such as the movement trajectory of the airbrush, the paint pressure, and the flow rate, to achieve closed-loop correction of process deviations and ensure the consistency and high quality of product surface treatment.

[0003] However, existing digital twin-based decision-making systems often fall short in their core intelligent decision-making capabilities when faced with rapidly changing production conditions, especially when dealing with the introduction of new materials and processes. The specific technical problem lies in the fact that current AI models are typically trained on historical datasets and then solidified into static modules. When a production line begins trial production of a prototype using a special new material such as photosensitive resin, changes in key physical properties of the material's surface characteristics, such as adhesion, roughness, and paint affinity, lead to differences in the fluid dynamics and film-forming mechanisms during the actual spraying process. This causes the decisions made by the static AI model trained on existing data to quickly become inaccurate, failing to effectively compensate for the process deviations introduced by the new material, directly resulting in an increased defect rate in trial production. Although the system can detect quality degradation, under the existing technical architecture, there is a significant delay in the feedback loop from detecting anomalies and triggering model updates to the new strategy taking effect. The fundamental reason is that traditional models lack the ability to evolve online in real time, making it difficult to safely and efficiently absorb small amounts of real-time production data and complete self-updates without downtime. At the same time, model training and online inference are usually designed separately, lacking an integrated pipeline that can achieve low-latency perception, decision-making, and strategy delivery. Therefore, how to enable the AI ​​in the digital twin to have the instant learning and adaptation capabilities of a skilled technician, quickly adjust strategies with a small number of samples, and achieve near real-time feedback control has become the core bottleneck restricting the widespread application of this technology in highly flexible and highly mixed production. Summary of the Invention

[0004] This invention addresses the technical problems existing in the prior art by providing a digital twin AI decision feedback communication system for a figurine production line. The system solves the problems mentioned in the background by setting up a state capture and simulation construction module, a virtual exploration and strategy distillation module, a strategy portability verification module, and a security fusion control module.

[0005] The technical solution of this invention to solve the above-mentioned technical problems is as follows: Specifically, it includes: a state capture and simulation construction module, a virtual exploration and policy distillation module, a policy portability verification module, and a security fusion control module connected in sequence, wherein;

[0006] State capture and simulation construction module: In response to events where the real-time quality data of the physical production line does not meet the preset threshold, it synchronously captures the real-time data of all sensors and actuators in the physical production line, forming a state dataset containing the complete instantaneous state of the production line. In the digital twin environment, it uses the state dataset as the initial state, loads the current production line control strategy as the initial strategy, and dynamically generates a local reward function based on the real-time defect pattern, thereby constructing an initial virtual environment for reinforcement learning training.

[0007] Virtual Exploration and Policy Distillation Module: In the initial virtual environment, the control agent conducts high-speed exploration based on the initial policy and combined with a directional perturbation mechanism, and collects the resulting interaction sequences as experience trajectory data. At the same time, using policy distillation technology, a lightweight set of optimization policy candidates is learned from the experience trajectory data.

[0008] Strategy portability verification module: Each candidate strategy in the optimization strategy candidate set is placed in a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation, and the strategy performance degradation boundary prediction method is applied to quantitatively evaluate the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment. Based on the robustness evaluation and the expected performance loss, a final optimization strategy is selected.

[0009] Safety Fusion Control Module: The final optimized strategy is weighted and fused with the current production line control strategy to generate fusion control instructions and send them to the actuators of the physical production line. At the same time, based on the difference between the actual state collected after the physical production line executes the fusion control instructions and the predicted state predicted by the digital twin environment according to the previous state and the instructions, the dynamic fusion weight coefficient of the weighted fusion is dynamically adjusted and the model parameters of the digital twin model are updated.

[0010] In a preferred embodiment, the process of synchronously capturing real-time data from all sensors and actuators in the physical production line to form a state dataset containing the complete instantaneous state of the production line in the state capture and simulation construction module is as follows:

[0011] When a high-resolution 3D vision sensor deployed at the end of the production line detects that the root mean square error of the paint film thickness uniformity of three consecutive workpieces exceeds a preset threshold, an event is triggered. After triggering, a synchronous acquisition command is sent to all data sources on the production line. The data sources include the high-resolution 3D vision sensor, the drive controller that controls the precision inkjet, the sensor that monitors environmental parameters, the servo controller that drives the robot axis movement, and the database interface connected to the manufacturing execution system.

[0012] The collected real-time data is integrated and processed to form a state dataset, which consists of the following five types of information:

[0013] The first type of information is the three-dimensional point cloud and defect coordinate annotations on the surface of the workpiece currently in production, obtained through a high-resolution three-dimensional vision sensor.

[0014] The second type of information is the spatial position and orientation data of all robotic arm joint angles and end effectors fed back by the robot servo controller.

[0015] The third type of information consists of real-time working air pressure, paint flow rate, and drive voltage parameters fed back by the precision inkjet drive controller.

[0016] The fourth type of information consists of real-time temperature, humidity, and airborne dust particle concentration readings fed back by environmental sensors;

[0017] The fifth type of information consists of all process specifications and parameters specified in the current production order, obtained through the Manufacturing Execution System (MES) interface.

[0018] In a preferred embodiment, the process of constructing an initial virtual environment for reinforcement learning training specifically includes:

[0019] Using the instantaneous state of the production line represented by the state dataset as the initial state, the digital twin model is driven to construct a simulation environment instance in the virtual space that is consistent with the physical production line state. At the same time, the current production line control policy being used is read from the production line control server, and the current production line control policy is used as the initial policy of the agent. The parameters of the initial policy are then loaded into the agent's policy network.

[0020] The analysis focuses on the defect coordinate annotations in the first type of information within the state dataset to identify the currently dominant defect type. Based on the identified dominant defect type, multiple quality evaluation index components most relevant to the dominant defect type are selected from a predefined reward function component library. For each selected quality evaluation index component, a corresponding virtual sensor is set up in the simulation environment instance of the digital twin model to calculate the index value of the quality evaluation index in any given subsequent state, and a corresponding target setting value is preset. The calculation process of the local reward function is as follows:

[0021] A1. For each selected quality evaluation index component, calculate the difference between the index value obtained by the virtual sensor in a given subsequent state and the target setting value corresponding to the index component; A2. Input each difference obtained in A1 into a preset nonlinear mapping function for processing. The nonlinear mapping function is used to map the input value to a smooth output value; A3. Assign a weight dynamically adjusted according to the dominant defect type to each quality evaluation index component. Multiply the output value of each index component after processing by the nonlinear mapping function in A2 by the weight assigned to the index component to obtain the weighted result of the index component; A4. Sum the weighted results of all quality evaluation index components to obtain a preliminary reward item; A5. A6. Calculate a reward focus factor between zero and one, the magnitude of which depends on the severity of the dominant defect type; A7. Multiply the preliminary reward component obtained through A4 with the reward focus factor obtained through A5 to obtain a first product focused on correcting the dominant defect; A8. Calculate the difference between the number 1 and the reward focus factor, and multiply this difference with a basic reward term used to maintain basic production efficiency and energy consumption levels to obtain a second product used to balance overall performance, where the basic reward term is calculated based on the state and actions taken by the agent in the simulation environment instance; A9. Add the first product obtained through A6 to the second product obtained through A7, and the sum is the value of the local reward function for a given state, action, and subsequent state.

[0022] In digital twin software, an interactive reinforcement learning training environment instance is instantiated using a state dataset as the initial state, an initial policy loaded into the policy network as the basis for the agent's behavior, and logic for dynamically calculating local reward functions from A1 to A8 as the environmental feedback mechanism. This training environment instance, the initial policy as the basis for behavior, the state dataset as the source of the initial state, and the dominant defect type information used to guide the calculation of the local reward function are collectively encapsulated to form a complete initial virtual environment.

[0023] In a preferred embodiment, in the virtual exploration and strategy distillation module, the process of controlling the agent to conduct high-speed exploration based on an initial strategy and combined with a directional perturbation mechanism, and to collect experience trajectory data, specifically involves:

[0024] The agent obtains an initial policy as the basis for behavior from the initial virtual environment and configures the parameters of the initial policy as the initial parameters of the agent's policy network. In the simulation environment instance of the initial virtual environment, the agent explores the interaction according to a hybrid action selection strategy. The hybrid action selection strategy selects the action directly calculated by the current agent's policy network based on the environmental state presented by the simulation environment instance with a preset high probability, and at the same time selects to execute a directional perturbation action with a preset low probability.

[0025] During the exploration process, the agent executes a hybrid action selection strategy in the simulation environment instance at a speed far exceeding the real-time operating clock of the physical production line, and interacts with the simulation environment instance in multiple rounds. It continuously records the state of the simulation environment instance at each moment, the actions taken by the agent, the immediate rewards fed back by the simulation environment instance, and the state of the next simulation environment instance after the interaction. All these interaction records arranged in chronological order together constitute the experience trajectory data.

[0026] In a preferred embodiment, the process of generating the directional disturbance action is as follows: A knowledge base dynamically maintained based on historical data of process and defect correlation accumulated during long-term production line operation is queried. This knowledge base is recorded in matrix form, with rows corresponding to different defect types and columns corresponding to different action dimensions of production line control. The matrix element values ​​represent the historical correlation strength of a specified defect type on a specified action dimension adjustment. For the dominant defect type corresponding to the current optimization task, a corresponding row value is extracted from the knowledge base as a sensitivity vector. A random direction noise vector with the same dimension as the action space is generated. The sensitivity vector and the random direction noise vector are multiplied element-wise to obtain a weighted noise vector.

[0027] The weighted noise vector is multiplied by a preset perturbation amplitude coefficient to obtain a directional perturbation value. The directional perturbation value is then vector-added with the original action value calculated by the agent's policy network based on the environmental state presented by the current simulation environment instance. The result is the directional perturbation action.

[0028] In a preferred embodiment, the process of learning and generating a lightweight set of optimization policy candidates from empirical trajectory data using policy distillation techniques specifically involves:

[0029] First, based on empirical trajectory data, a state value evaluation network and an advantage function estimator are trained to evaluate the merits of state-action pairs, where a state-action pair refers to the combination of the environmental state recorded in the empirical trajectory data and the action taken by the agent. Then, using the trained advantage function estimator, the advantage value of all state-action pairs in the empirical trajectory data is evaluated. The advantage value is used to quantify the expected improvement of taking a specified action in a specified state relative to the average level, and state-action pairs with advantage values ​​higher than a set threshold are selected to form a high-quality subset of empirical data.

[0030] Then, several student policy networks with simpler structures than the agent policy network are initialized; these student policy networks are trained using policy distillation techniques, with the training objective being to minimize a composite loss function, which is calculated as the sum of a first weighted loss and a second weighted loss.

[0031] By optimizing the composite loss function, each student policy network learns to imitate high-quality decision-making patterns. After training, each student policy network is independently and quickly evaluated in the initial virtual environment to obtain evaluation metrics including average cumulative reward and critical defect correction rate. Finally, all student policy networks are ranked according to the evaluation metrics, and several student policy networks with the best performance are selected. Their network parameters, evaluation metrics, and performance scores are packaged together to form a lightweight optimization policy candidate set.

[0032] In a preferred embodiment, the specific process of placing each candidate policy in the optimization policy candidate set into a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation in the policy portability verification module is as follows:

[0033] Based on the initial state of the production line and the digital twin model represented by the state dataset obtained from the initial virtual environment, a high-fidelity verification simulation environment is constructed. Three types of non-ideal factor models learned from historical data of the physical production line are injected into the high-fidelity verification simulation environment. The first type is a sensor noise model, used to add random disturbances to the readings of virtual sensors that conform to the actual noise statistical characteristics of their physical counterparts. The second type is an actuator delay and response deviation model, used to simulate the dynamic characteristics of response lag and nonlinearity of actuators after receiving commands. The third type is a random environmental disturbance model, used to simulate unpredictable environmental disturbances in the production workshop, such as subtle airflow fluctuations.

[0034] Each candidate strategy in the optimization strategy candidate set is loaded into the high-fidelity verification simulation environment, and multiple independent Monte Carlo simulation runs are performed starting from the initial state of the production line.

[0035] Record and calculate the average performance index obtained by each candidate strategy across all simulation rounds. The average performance index includes the average cumulative reward or defect correction success rate. This average performance index serves as a numerical value to quantify the robustness evaluation result of the candidate strategy and is called the robustness evaluation score.

[0036] In a preferred embodiment, the specific process of quantitatively evaluating the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment is as follows:

[0037] For each candidate policy in the optimization policy candidate set, run it in a simulation environment instance without injecting non-ideal factors, and collect the state access probability distribution induced by the policy decision behavior; for each accessed state, calculate the policy sensitivity of the candidate policy at that state. The policy sensitivity is defined by calculating the Jacobian matrix of the policy function of the candidate policy at the input state and obtaining the Frobenius norm of this Jacobian matrix.

[0038] Meanwhile, using a confidence estimator trained in conjunction with the digital twin model to evaluate the model's predictive uncertainty, the confidence estimate of the digital twin model at each access state and the state-action pair consisting of the action output by the candidate policy in that state is evaluated.

[0039] Next, for each candidate strategy, a value is calculated to quantify its expected performance loss, called the performance degradation prediction value.

[0040] Based on the robustness evaluation score and performance degradation prediction value of each candidate strategy, a comprehensive decision is made to select the final optimization strategy. Specifically: First, the robustness evaluation scores of all candidate strategies in the optimization strategy candidate set are normalized so that all scores are mapped to a numerical range of zero to one; the performance degradation prediction values ​​of all candidate strategies in the optimization strategy candidate set are normalized so that all prediction values ​​are mapped to a numerical range of zero to one; then, a comprehensive decision score is calculated for each candidate strategy. The calculation process is as follows: multiply the first weight coefficient by the normalized robustness evaluation score of the candidate strategy to obtain the first weighting term; multiply the second weight coefficient by the normalized performance degradation prediction value of the candidate strategy to obtain the second weighting term; subtract the second weighting term from the first weighting term, and the difference is the comprehensive decision score of the candidate strategy; finally, compare the comprehensive decision scores of all candidate strategies in the optimization strategy candidate set, and select the candidate strategy with the highest comprehensive decision score as the final optimization strategy.

[0041] In a preferred embodiment, the process of dynamically adjusting the dynamic fusion weight coefficients in the security fusion control module specifically involves:

[0042] A dynamically changing fusion weight coefficient is set, with an initial value of zero, indicating complete reliance on the current production line control strategy. In each control cycle, the final optimized strategy and the current production line control strategy are loaded and run separately. Based on the actual state, which is consistent with the structure of the state dataset formed by real-time collection and processing of various sensors on the physical production line, the actions suggested by the final optimized strategy and the actions suggested by the current production line control strategy are obtained. The actions suggested by the final optimized strategy are recorded as the final optimized actions, and the actions suggested by the current production line control strategy are recorded as the current production line actions. According to the dynamic fusion weight coefficient, the final optimized actions and the current production line actions are weighted and fused to generate fused control instructions.

[0043] At the same time as issuing the fusion control command to the actuator, using a digital twin model, based on the actual physical production line state collected at the end of the previous control cycle and the fusion control command to be issued, the predicted state expected to be reached after execution is simulated and calculated. After the physical production line completes execution, the actual new state is collected through a sensor network. This actual new state is a data set similar to the state dataset, reflecting the instantaneous state of the production line after execution. The difference between the predicted state and the actual new state in the dimension representing the key process quality indicators of the production line is calculated to obtain the instantaneous prediction error. The instantaneous prediction error is compared with a preset safety error threshold. Based on the comparison result and combined with the preset adjustment step size rule, the value of the dynamic fusion weight coefficient is dynamically updated.

[0044] When the dynamic fusion weight coefficient reaches a value of one and remains stable in subsequent consecutive control cycles, or when the production line control strategy changes due to external instructions, the weighted fusion process is considered complete. After the fusion process is completed, the final optimized strategy is set as the new current production line control strategy, and the state of the safety fusion control module is reset to prepare for the fusion deployment of the next optimized strategy.

[0045] In a preferred embodiment, the process of updating the model parameters of the digital twin model specifically includes:

[0046] During the dynamic adjustment of the weighted fusion weights, the complete data sequence generated in each control cycle, including the previous state, the issued fusion control commands, the collected actual new state, and the calculated real-time prediction error, is continuously stored in a first-in-first-out data buffer with a fixed capacity. When the data buffer is full, or when the average value of the real-time prediction errors of multiple consecutive control cycles stored in the buffer exceeds a preset trigger threshold, the online parameter update of the digital twin model is initiated.

[0047] Online parameter updates are performed by minimizing a composite loss function;

[0048] By optimizing the composite loss function, online, incremental fine-tuning of the model parameters of the digital twin model can be achieved, thereby reducing the difference between model predictions and physical reality.

[0049] The beneficial effects of this invention are: by capturing the defect status of the production line in real time, the control strategy is quickly constructed and optimized in the digital space. After robustness verification and migration risk assessment, the optimization strategy is integrated into the physical production line in a progressive and safe manner. During the execution process, the model is continuously corrected using real-time data. Thus, without stopping the machine or generating actual material waste, autonomous, rapid, closed-loop optimization and adaptive control of precision process problems such as figurine painting are achieved, which significantly improves the agility and quality consistency of the production line in dealing with new materials and new processes. Attached Figure Description

[0050] Figure 1 This is a flowchart of the method of the present invention;

[0051] Figure 2 This is a block diagram of the system structure of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0053] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0054] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0055] Example 1

[0056] This embodiment provides, for example Figure 1-2 The illustrated digital twin AI decision feedback communication system for a figurine production line includes, in particular, a state capture and simulation construction module, a virtual exploration and strategy distillation module, a strategy portability verification module, and a security fusion control module connected in sequence.

[0057] State capture and simulation construction module: In response to events where the real-time quality data of the physical production line does not meet the preset threshold, it synchronously captures the real-time data of all sensors and actuators in the physical production line, forming a state dataset containing the complete instantaneous state of the production line. In the digital twin environment, it uses the state dataset as the initial state, loads the current production line control strategy as the initial strategy, and dynamically generates a local reward function based on the real-time defect pattern, thereby constructing an initial virtual environment for reinforcement learning training.

[0058] Virtual Exploration and Policy Distillation Module: In the initial virtual environment, the control agent conducts high-speed exploration based on the initial policy and combined with a directional perturbation mechanism, and collects the resulting interaction sequences as experience trajectory data. At the same time, using policy distillation technology, a lightweight set of optimization policy candidates is learned from the experience trajectory data.

[0059] Strategy portability verification module: Each candidate strategy in the optimization strategy candidate set is placed in a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation, and the strategy performance degradation boundary prediction method is applied to quantitatively evaluate the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment. Based on the robustness evaluation and the expected performance loss, a final optimization strategy is selected.

[0060] The safety fusion control module: It performs weighted fusion of the final optimization strategy and the current production line control strategy, generates fusion control commands and issues them to the actuators of the physical production line. At the same time, based on the difference between the actual state collected after the physical production line executes the fusion control commands and the predicted state predicted by the digital twin environment according to the previous state and the commands, it dynamically adjusts the dynamic fusion weight coefficients of the weighted fusion and updates the model parameters of the digital twin model.

[0061] In this embodiment, it is specifically necessary to explain that the process of synchronously capturing real-time data from all sensors and actuators in the physical production line and forming a state dataset containing the complete instantaneous state of the production line in the state capture and simulation construction module is as follows:

[0062] An event is triggered when a high-resolution 3D vision sensor deployed at the end of the production line detects that the root mean square error (RMSE) of the paint film thickness uniformity of three consecutive workpieces exceeds a preset threshold. The preset threshold can be set according to the process requirements of different products. For example, for high-end figurine painting, the threshold can be set to ensure that the RMSE of the paint film thickness uniformity does not exceed ±5 micrometers. The calculation method for the RMSE of the paint film thickness uniformity is as follows: First, the high-resolution 3D vision sensor acquires 3D point cloud data of a specified measurement area on the workpiece surface and extracts the paint film thickness value of each measurement point. Next, the arithmetic mean of the paint film thickness values ​​of all measurement points is calculated. Then, the difference between the thickness value of each measurement point and the average value is calculated. Finally, the square root of the squares of these differences is taken, and the result is the RMSE of the paint film thickness uniformity of the workpiece surface. The system continuously calculates and determines whether the error values ​​of three consecutive workpieces exceed a preset threshold to confirm whether an optimization event is triggered. Once triggered, a synchronous acquisition command is sent to all data sources on the production line. These data sources include high-resolution 3D vision sensors, drive controllers that control precision inkjet pens, sensors that monitor environmental parameters, servo controllers that drive robot axis movements, and database interfaces connected to the manufacturing execution system. The synchronous acquisition command can be implemented through precise hardware synchronization signals or based on a precision network clock protocol to ensure that the timestamp deviation of all data is within milliseconds, thereby guaranteeing the time alignment accuracy of the status dataset.

[0063] The collected real-time data is integrated and processed to form a state dataset, which consists of the following five types of information:

[0064] The first type of information is the three-dimensional point cloud and defect coordinate annotation on the surface of the workpiece currently in production, obtained through a high-resolution three-dimensional vision sensor. The high-resolution three-dimensional vision sensor can be a laser scanner or a structured light scanner, with a point cloud resolution of up to 0.01 mm, which can accurately characterize the micro-undulations and defect contours of the paint surface.

[0065] The second type of information consists of the position and orientation data of all robotic arm joint angles and end effectors in space, fed back by the robot servo controller. The position and orientation data are usually represented by a six-dimensional vector, including three-dimensional coordinates and three Euler angles, with an accuracy of micrometers and milliradians.

[0066] The third type of information consists of real-time working air pressure, paint flow rate, and drive voltage parameters fed back by the precision inkjet drive controller; the monitoring accuracy of real-time working air pressure can reach 0.1 Pascal, and the monitoring accuracy of paint flow rate can reach 0.1 ml / min, providing a data basis for precise process control;

[0067] The fourth type of information is the real-time temperature, humidity, and airborne dust particle concentration readings fed back by environmental sensors; the environmental sensors must meet the cleanroom monitoring standards, such as temperature monitoring accuracy ±0.5℃, humidity monitoring accuracy ±2%RH, and dust particle counting conforming to ISO standards.

[0068] The fifth type of information is all the process specifications parameters specified in the current production order obtained through the manufacturing execution system interface. The process specifications parameters include at least the paint type, total coating thickness, coating sequence, drying time, and special effect requirements. All five types of information together constitute a complete, time-aligned snapshot of the production line's instantaneous state. The completeness and high accuracy of this snapshot lay a reliable data foundation for building a high-fidelity virtual simulation environment in the digital twin environment, ensuring the consistency between virtual optimization and physical reality.

[0069] The process of constructing an initial virtual environment for reinforcement learning training is as follows:

[0070] Using the instantaneous state of the production line represented by the state dataset as the initial state, a digital twin model is driven to construct a simulation environment instance in virtual space that is consistent with the physical production line state. The digital twin model is a high-fidelity model calibrated based on physical laws (such as fluid mechanics and dynamics) and actual equipment parameters, capable of simulating the entire process of paint atomization, transport, deposition, and film formation during spraying. Simultaneously, the current production line control strategy being used is read from the production line control server and used as the initial strategy of the agent. The parameters of the initial strategy are then loaded into the agent's policy network. The policy network is typically a deep neural network, and loading its parameters means using the existing control logic as the starting point for the reinforcement learning agent's exploration, rather than random initialization. This significantly accelerates the convergence speed of the subsequent optimization process.

[0071] The analysis of the first type of information in the state dataset, including defect coordinate annotations, identifies the currently dominant defect type. For example, by clustering analysis of the distribution and morphological characteristics of defect points, defects can be identified as typical types such as "orange peel," "sag," "pinholes," or "insufficient coverage." Based on the identified dominant defect type, multiple quality evaluation index components most relevant to the dominant defect type are selected from a predefined reward function component library. For example, if the dominant defect is "orange peel," then "surface roughness" and "gloss uniformity" are selected as core quality evaluation index components; if it is "sag," then "coating thickness distribution variance" and "edge sharpness" are selected as core components. For each selected quality evaluation index component, a corresponding virtual sensor is set in the simulation environment instance of the digital twin model to calculate the index value of the quality evaluation index under any given subsequent state, and a corresponding target setting value is preset. The virtual sensor is implemented by embedding an algorithm module consistent with the measurement principle of physical sensors in the digital twin model. For example, the virtual roughness meter calculates the arithmetic mean roughness by analyzing the height map of the virtual workpiece surface. The calculation process of the local reward function is as follows:

[0072] A1. For each selected quality evaluation index component, calculate the difference between the index calculation value obtained by the virtual sensor in a given subsequent state and the target setting value corresponding to the index component; A2. Input each difference obtained through A1 into a preset nonlinear mapping function for processing. The nonlinear mapping function is used to map the input value to a smooth output value. The preferred nonlinear mapping function is a hyperbolic tangent function, which can saturate large differences into a bounded range, preventing drastic fluctuations in the reward signal, thereby improving the stability of reinforcement learning training; A3. Assign a weight to each quality evaluation index component that is dynamically adjusted according to the dominant defect type, and multiply the output value of each index component after processing by the nonlinear mapping function in A2 by... The weighted result of the indicator components is obtained by assigning weights to the indicator components. The principle of dynamic adjustment of weights is to assign higher weights, such as 0.7, to indicators related to the dominant defect, and lower weights, such as 0.3, to other related indicators, so that the reward function strongly guides the agent to correct the main defect in the early stage of training. A4. The weighted results of all quality evaluation indicator components are summed to obtain a preliminary reward item. A5. A reward focus factor between zero and one is dynamically calculated, and its value depends on the severity of the dominant defect type. The severity is quantified by the weighted sum of the two indicators of "defect point cloud density" (i.e., the number of defect points per unit area) and "average depth deviation" of the type of defect identified in the current state dataset.The initial value of the reward focus factor is set to 0.9, and it decreases linearly to a lower limit (e.g., 0.3) according to a preset decay rate (e.g., 0.1 per N rounds) as the number of training rounds increases, thus achieving a smooth transition from focused error correction to comprehensive optimization; A6, multiply the preliminary reward component obtained through A4 with the reward focus factor obtained through A5 to obtain the first product focused on correcting the dominant defect; A7, calculate the difference between the number 1 and the reward focus factor, and multiply this difference with a basic reward term used to maintain basic production efficiency and energy consumption levels to obtain the second product used to balance comprehensive performance, where the basic reward term is calculated based on the state and actions taken by the agent in the simulation environment instance; the basic reward term consists of a linear combination of a production efficiency reward component and an energy consumption penalty component; the production efficiency reward component is related to the agent's completion of a single task The simulation time for spraying is inversely proportional to the energy consumption penalty component; the energy consumption penalty component is directly proportional to the product of the average paint flow rate and average working air pressure controlled by the agent during the action cycle (representing instantaneous power); the specific value of the basic reward item is obtained by subtracting the energy consumption penalty component from the production efficiency reward component and then multiplying it by a normalization coefficient; A8, add the first product obtained through A6 and the second product obtained through A7, and the sum is the value of the local reward function for a given state, action, and subsequent state; this calculation method creates a dynamically adaptive reward signal, which can strongly drive the agent to explore strategies to solve the current prominent quality problems in the early stage of training, and gradually guide the strategy to evolve towards a Pareto optimal solution that balances multiple objectives such as quality, efficiency, and energy consumption in the later stage, effectively avoiding optimization direction deviation or convergence difficulties that may be caused by a single static reward function.

[0073] In digital twin software, an interactive reinforcement learning training environment instance is instantiated using a state dataset as the initial state, an initial policy loaded into the policy network as the basis for the agent's behavior, and logic for dynamically calculating local reward functions from A1 to A8 as the environmental feedback mechanism. This training environment instance, the initial policy as the basis for behavior, the state dataset as the source of the initial state, and the dominant defect type information used to guide the calculation of the local reward function are collectively encapsulated to form a complete initial virtual environment. This initial virtual environment, as a self-contained, goal-oriented training task unit, is directly passed to the subsequent virtual exploration and policy distillation modules, ensuring lossless transfer and efficient execution of the optimization task context.

[0074] In this embodiment, it is specifically necessary to explain that in the virtual exploration and strategy distillation module, the process by which the controlling agent conducts high-speed exploration based on the initial strategy and combines it with a directional perturbation mechanism, and collects experience trajectory data, is as follows:

[0075] An initial policy is obtained from the initial virtual environment to form the basis of the behavior, and the parameters of the initial policy are configured as the initial parameters of the agent's policy network. In the simulation environment instance of the initial virtual environment, the agent explores the interaction according to a hybrid action selection strategy. The hybrid action selection strategy selects an action directly calculated by the current agent's policy network based on the environmental state presented by the simulation environment instance with a preset, higher probability, and simultaneously selects to execute a directional perturbation action with a preset, lower probability. The higher probability is usually set to 0.95, and the lower probability, i.e., the exploration rate, is usually set to 0.05, so as to achieve a balance between the reliability of utilizing existing strategies and the possibility of exploring new strategies.

[0076] During the exploration process, the agent executes a hybrid action selection strategy in the simulation environment instance at a speed far exceeding the real-time clock speed of the physical production line, and interacts with the simulation environment instance in multiple rounds. The running speed of the simulation environment can be more than 1000 times faster than the real-time clock of the physical production line, thereby accumulating a large amount of exploration data in a short period of time. It continuously records the state of the simulation environment instance at each moment, the actions taken by the agent, the immediate rewards fed back by the simulation environment instance, and the state of the next simulation environment instance after the interaction. All these interaction records arranged in chronological order constitute the experience trajectory data. The experience trajectory data is usually stored in a circular buffer in the form of a four-tuple sequence of (state, action, reward, next state) for subsequent learning.

[0077] The process of generating targeted disturbance actions is as follows: A knowledge base dynamically maintained based on historical data of process and defect correlations accumulated during long-term production line operation is queried. This knowledge base records the correlation strength between defect types and action dimensions in matrix form, with rows corresponding to different defect types and columns corresponding to different action dimensions of production line control. Matrix element values ​​represent the historical correlation strength of a specified defect type on a specified action dimension adjustment. The matrix element values ​​of the knowledge base can be continuously updated through online learning. For example, when a certain action adjustment is verified to effectively improve a certain type of defect, the correlation strength value at the corresponding position will increase accordingly. For the dominant defect type corresponding to the current optimization task, a corresponding row of values ​​is extracted from the defect type and action dimension correlation strength knowledge base as a sensitivity vector. Each element value of the sensitivity vector represents the effect of the corresponding action dimension adjustment on the correction... The sensitivity of the current dominant defect type is assessed. For example, for the "orange peel" defect, the sensitivity of the spraying air pressure dimension may be positive and relatively large, indicating that appropriately increasing the air pressure may help improve the orange peel effect. However, for the "sagging" defect, the sensitivity of this dimension may be negative, indicating that the air pressure needs to be reduced. A random direction noise vector with the same dimension as the action space is generated. The random direction noise vector is usually sampled from a multidimensional normal distribution with zero mean and an identity matrix. The sensitivity vector and the random direction noise vector are multiplied element-wise to obtain a weighted noise vector. This element-wise multiplication operation modulates the amplitude of the noise disturbance in different action dimensions with the sensitivity vector, thereby guiding the purely random exploration to the dimension direction that is more sensitive to the current defect, improving the efficiency and targeting of the exploration.

[0078] The weighted noise vector is multiplied by a preset perturbation amplitude coefficient to obtain a directional perturbation value. The perturbation amplitude coefficient is an adjustable hyperparameter used to control the step size of the exploration. It is usually set to a small value, such as 0.1, based on the actual range of the action value. The directional perturbation value is then vector-added with the original action value calculated by the agent's policy network based on the environmental state presented by the current simulation environment instance. The result is the directional perturbation action. This generation method ensures that the perturbation is a biased and directional exploration around the action recommended by the current policy, rather than a completely random walk.

[0079] The process of learning and generating a lightweight set of optimization policy candidates from empirical trajectory data using policy distillation techniques is as follows:

[0080] First, based on empirical trajectory data, a state value evaluation network and an advantage function estimator are trained to evaluate the merits of state-action pairs. A state-action pair refers to the combination of the environmental state recorded in the empirical trajectory data and the action taken by the agent. The advantage function estimator is trained using a temporal difference learning method, with the training objective of minimizing the mean squared error of the difference between the predicted advantage value and the actual reward discount, and between the predicted advantage value and the estimated state value. Next, using the trained advantage function estimator, the advantage values ​​of all state-action pairs in the empirical trajectory data are evaluated. The advantage value quantifies the expected improvement of taking a specified action in a given state relative to the average level. State-action pairs with advantage values ​​higher than a set threshold are selected to form a high-quality subset of empirical data. The set threshold can be a fixed quantile of all calculated advantage values, for example, selecting state-action pairs with the top 20% advantage values.

[0081] Then, several student policy networks with simpler structures than the agent policy network are initialized. For example, the original agent policy network might be a deep neural network with 5 hidden layers, while the student policy networks can be designed to contain only 2 to 3 hidden layers with fewer neurons per layer to reduce computational complexity and storage overhead. These student policy networks are trained using policy distillation techniques, with the training objective being to minimize a composite loss function, calculated as the sum of a first weighted loss and a second weighted loss. The first weighted loss is obtained by multiplying the first loss weight coefficients by a probability distribution difference metric, preferably the Kullback-Leibler divergence, which quantifies the difference between the action probability distribution output by the student policy network in a given environmental state and a target action probability distribution. The target action probability distribution is derived from the statistical characteristics of high-dominance state-action pairs selected from a subset of high-quality empirical data or the output of a complex teacher policy network trained on that subset. Specifically, in a preferred embodiment, the target action probability distribution is determined by selecting all high-dominance values ​​in the same or similar environmental states from a subset of high-quality empirical data. The action values ​​are estimated using Gaussian kernel density estimation, and the resulting probability density function is the probability distribution of the target action in that state. The second weighted loss is obtained by multiplying the weight coefficients of the second loss by a feature representation difference metric. The feature representation difference metric is calculated by the difference between the cosine similarity values ​​of the hidden layer feature vectors of the Digital One and Student Policy networks in a given environment state and the hidden layer feature vectors of the State Value Evaluation network in the same environment state. This difference measures the semantic dissimilarity of the feature representations within the two networks. The cosine similarity value is the product of the two feature vectors and its... The ratio of the product of modulo lengths, the closer its value is to 1, the more similar the features are. By minimizing the difference in feature representation, the student policy network is forced to learn deep features related to state value judgments, rather than just imitating surface actions. This helps to improve the generalization ability and robustness of the policy. The first loss weight coefficient and the second loss weight coefficient are both preset positive numbers, and their sum is one. They are used to balance the relative importance of probability distribution differences and feature representation differences in the composite loss function. For example, the first loss weight coefficient can be set to 0.7 and the second loss weight coefficient can be set to 0.3 to emphasize action imitation, while being supplemented by feature-level guidance.

[0082] By optimizing the composite loss function, each student policy network learns to imitate high-quality decision-making patterns. After training, each student policy network undergoes independent and rapid performance evaluation in the initial virtual environment to obtain evaluation metrics including average cumulative reward and critical defect correction rate. Rapid performance evaluation is achieved by running a fixed number of rounds (e.g., 100 rounds) in the simulation environment and calculating the average performance metrics. Finally, all student policy networks are ranked according to the evaluation metrics, and several student policy networks with the best performance are selected. Their network parameters, evaluation metrics, and performance scores are packaged together to form a lightweight optimization policy candidate set. Each candidate policy in the optimization policy candidate set includes, in addition to network parameters, its corresponding evaluation metrics and scores, providing a basis for subsequent verification and selection.

[0083] In this embodiment, the specific process of placing each candidate policy in the optimization policy candidate set into a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation in the policy transferability verification module is as follows:

[0084] Based on the initial state of the production line and its digital twin model, represented by a state dataset obtained from the initial virtual environment, a high-fidelity verification simulation environment is constructed. This high-fidelity verification simulation environment, building upon the basic digital twin model, increases the simulation fidelity for non-ideal physical factors. Three types of non-ideal factor models learned from historical physical production line data are injected into the high-fidelity verification simulation environment. The first type is a sensor noise model, used to add random perturbations to the virtual sensor readings that conform to the actual noise statistical characteristics of their physical counterparts. For example, to simulate the point cloud noise of a high-resolution 3D vision sensor, a perturbation conforming to specific... The first type of model uses Gaussian white noise with mean and variance, or specific spatially correlated noise caused by motion fuzziness. The second type of model is the actuator delay and response deviation model, which is used to simulate the dynamic characteristics of the actuator's response lag and nonlinearity after receiving a command. For example, the flow control of a precision airbrush can be modeled as a dynamic system with dead time and first-order inertial elements, and its parameters are obtained by system identification of historical step response data. The third type of model is the random environmental disturbance model, which is used to simulate unpredictable environmental disturbances in the production workshop, such as subtle airflow fluctuations. This model can be simulated by introducing a low-intensity random force or torque into the dynamic equations of the simulation environment.

[0085] Each candidate strategy in the optimization strategy candidate set is loaded into a high-fidelity verification simulation environment, and multiple independent Monte Carlo simulations are run starting from the initial state of the production line. The number of simulation rounds is usually set between 50 and 200 to ensure the statistical stability of the evaluation results. In each simulation round, in addition to applying the three types of non-ideal factor models, a small, bounded adversarial perturbation is intentionally superimposed on the environmental state perceived by the agent to test the fault tolerance and stability of the strategy in the face of perception bias. The adversarial perturbation can be a random perturbation with bounded amplitude (e.g., not exceeding 5% of the normal measurement value) added to each dimension of the state vector, or a small adversarial sample generated by the gradient-based fast gradient sign method, which aims to detect the vulnerability near the policy decision boundary.

[0086] The average performance index of each candidate strategy across all simulation rounds is recorded and calculated. The average performance index includes the average cumulative reward or defect correction success rate. This average performance index serves as a numerical value to quantify the robustness evaluation result of the candidate strategy, and is called the robustness evaluation score. Through this process, strategies that not only perform well under ideal conditions, but also maintain stable performance in a near-realistic complex environment with noise, delay, and interference are selected in the virtual environment. This greatly increases the likelihood of the selected strategies being successfully applied on the actual physical production line.

[0087] The specific process for quantitatively evaluating the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment is as follows:

[0088] For each candidate policy in the optimization policy candidate set, it is run in a simulation environment instance without the injection of non-ideal factors, and the state access probability distribution induced by the policy decision behavior is collected. This is usually achieved by running the policy a sufficient number of steps in the simulation and counting the frequency of different state regions visited. For each visited state, the policy sensitivity of the candidate policy at that state is calculated. The policy sensitivity is defined by calculating the Jacobian matrix of the policy function of the candidate policy at the input state and obtaining the Frobenius norm of this Jacobian matrix. This norm value is used to comprehensively measure the overall drastic change of all action components of the policy output when the input state is slightly perturbed. The Jacobian matrix can be efficiently calculated using automatic differentiation techniques. A high Frobenius norm means that the policy's decision output at that state point is very sensitive to changes in the input, that is, the policy behavior is "fragile". Slight perception errors may lead to completely different control commands, which is very risky in actual transfer.

[0089] Simultaneously, a confidence estimator, trained in conjunction with the digital twin model, is used to assess the model's predictive uncertainty. This estimator evaluates the confidence level of the digital twin model at each access state and the state-action pair formed by the action output by the candidate policy in that state. The confidence level quantifies the uncertainty of the digital twin model's prediction of the dynamics of transitioning from the current state-action pair to the next state; a higher value indicates a less reliable prediction. The confidence estimator can be implemented using ensemble learning, Bayesian neural networks, or direct prediction variance methods. It is trained synchronously during the digital twin model training phase and can identify regions where the model has low prediction confidence due to insufficient training data or complex dynamics.

[0090] Next, for each candidate policy, a value used to quantify its expected performance loss is calculated, called the performance degradation prediction. The performance degradation prediction equals a scaling factor multiplied by an expected value. The scaling factor is a positive constant that needs to be pre-calibrated using historical migration experimental data. The scaling factor can be obtained during the development phase by deploying a small number of policies from simulation to a physical prototype, recording their actual performance degradation, and then performing a linear regression fit with the calculated expected value. The expected value represents the weighted average of the following product under the state access probability distribution induced by the candidate policy: this product represents the product of the policy sensitivity of the candidate policy at the state and the confidence estimate of the digital twin model at the corresponding state and the action output by the candidate policy at that state; this product... The core idea is that if a policy exhibits high sensitivity (high policy sensitivity) in a region of high model uncertainty (high confidence estimate), then this is a potential "performance collapse" risk point. The expected value calculation averages these risk points over the entire behavioral trajectory of the policy to obtain an overall risk measure. The performance degradation prediction value is a non-negative quantitative indicator. The larger the value, the greater the potential decline in the core performance indicators when the candidate policy migrates from the current digital twin environment to the target physical environment, i.e., the higher the risk of migration failure. This method does not require knowledge of the real physical model; it only utilizes the uncertainty information of the simulation model itself and the local behavioral characteristics of the policy to make a forward-looking and quantitative prediction of migration performance loss that is difficult to measure directly.

[0091] Based on the robustness evaluation score and performance degradation prediction value of each candidate strategy, a comprehensive decision is made to select the final optimization strategy. Specifically: First, the robustness evaluation scores of all candidate strategies in the optimization strategy candidate set are normalized so that all scores are mapped to a numerical range of zero to one. The normalization can use the min-max normalization method. The performance degradation prediction values ​​of all candidate strategies in the optimization strategy candidate set are normalized so that all prediction values ​​are mapped to a numerical range of zero to one. Since a smaller performance degradation prediction value is better, after normalization, its complement (1 - normalized value) can be taken as the "migration safety" score, so that a larger value indicates higher safety. Then, a comprehensive decision score is calculated for each candidate strategy. The calculation process is as follows: multiply the first weight coefficient by the normalized robustness evaluation score of the candidate strategy to obtain the first weighting term; multiply the second weight coefficient by the normalized performance degradation prediction value of the candidate strategy to obtain the second weighting term. The difference between the first and second weighted terms is the comprehensive decision score of the candidate strategy. Both the first and second weighting coefficients are preset constants greater than zero, used to adjust the relative importance of robustness assessment and migration risk prediction in the comprehensive decision. In practical applications, the first weighting coefficient can be set to 0.6 and the second weighting coefficient to 0.4, emphasizing both strategy robustness and sufficient vigilance regarding migration risk. Finally, the comprehensive decision scores of all candidate strategies in the optimization strategy candidate set are compared, and the candidate strategy with the highest comprehensive decision score is selected as the final optimized strategy. This multi-criteria decision-making mechanism systematically balances the strategy's "performance in the virtual environment" and "potential risks of migration to the physical environment," avoiding the problem of "simulation overfitting" and failure in actual deployment of strategies selected based solely on a single indicator (such as the highest reward in the virtual environment). This scientifically and reliably selects the optimized strategy most likely to succeed on the real production line.

[0092] In this embodiment, it is specifically necessary to explain that the process of dynamically adjusting the dynamic fusion weight coefficients in the security fusion control module is as follows:

[0093] A dynamically changing fusion weight coefficient is set, with an initial value of zero, indicating complete reliance on the current production line control strategy. In each control cycle, the final optimized strategy and the current production line control strategy are loaded and run separately. Based on the actual state, which is consistent with the state dataset structure and is collected and processed in real time by various sensors on the physical production line, the actions suggested by the final optimized strategy and the actions suggested by the current production line control strategy are obtained. The actual state information typically includes point clouds scanned in real time by high-resolution 3D vision sensors, robot joint angles, inkjet pen working parameters, environmental readings, etc., and its data structure is the same as the state dataset, providing consistent input for decision-making. The actions suggested by the final optimized strategy are recorded as the final optimized actions, and the actions suggested by the current production line control strategy are recorded as the current production line actions. Based on the dynamic fusion weight... The dynamic fusion weighting coefficient is used to weight and fuse the final optimized action with the current production line action to generate a fused control command. The specific calculation process is as follows: subtract the dynamic fusion weighting coefficient from the number one to obtain the weight of the current production line action; use the dynamic fusion weighting coefficient as the weight of the final optimized action; multiply the weight of the current production line action by the current production line action to obtain the first product; multiply the weight of the final optimized action by the final optimized action to obtain the second product; add the first product and the second product, and the sum is the fused control command issued to the physical production line actuator in the current control cycle. This linear weighted fusion method ensures a smooth transition of control commands. When the dynamic fusion weighting coefficient changes from 0 to 1, the control is smoothly transferred from the old strategy to the new strategy, avoiding the impact of command jumps on precision equipment.

[0094] At the same moment the fusion control command is issued to the actuator, a digital twin model is used to simulate and calculate the predicted state to be reached after execution, based on the actual physical production line state collected at the end of the previous control cycle and the fusion control command to be issued. After the physical production line completes execution, the actual new state is collected through a sensor network. This actual new state is a data set similar to the state dataset, reflecting the instantaneous state of the production line after execution. The difference between the predicted state and the actual new state in terms of the dimensions characterizing key process quality indicators of the production line is calculated to obtain the instantaneous prediction error. Key process quality indicators may include paint film thickness uniformity, coating coverage integrity, and spraying path. For accuracy, the difference can be measured by the weighted sum of squares of the differences between the predicted and actual values ​​of these indicators. The instantaneous prediction error is compared to a preset safety error threshold. Based on the comparison result and a preset adjustment step size rule, the value of the dynamic fusion weight coefficient is dynamically updated. Specifically, a dynamically changing dynamic fusion weight coefficient is set, with an initial value of zero, indicating complete reliance on the current production line control strategy. The update of the dynamic fusion weight coefficient follows these rules: a downward adjustment step size (e.g., 0.2) and an upward adjustment step size smaller than this value (e.g., 0.05), and a stable observation window (e.g., 5-10 control cycles); in each control cycle:

[0095] If the real-time prediction error is greater than the safety error threshold (which can be set according to process requirements, such as 20% of the allowable tolerance of the paint film thickness), the dynamic fusion weight coefficient is reduced by the down-adjustment step size value, and the result is compared with zero, taking the larger one as the new value; this achieves a rapid and safe revert to the current production line control strategy.

[0096] If the real-time prediction error is less than or equal to the safety error threshold, and this condition has been met for the number of periods specified by the stable observation window, then the dynamic fusion weight coefficient is increased by the step size value, and the result is compared with one, taking the smaller one as the new value; this achieves a gradual and cautious tilt toward the final optimization strategy.

[0097] In this rule, the step size for downward adjustment is greater than the step size for upward adjustment, which reflects the safety priority principle of "quick retreat and slow advance" during the control strategy switching process;

[0098] When the dynamic fusion weight coefficient reaches a value of one and remains stable in subsequent consecutive control cycles, or when the production line control strategy changes due to external instructions, the weighted fusion process is considered complete. After the fusion process is completed, the final optimized strategy is set as the new current production line control strategy, and the state of the safety fusion control module is reset to prepare for the next optimization strategy fusion deployment. This process realizes seamless, safe, and smooth switching of the optimization strategy on the real production line. It is data-driven throughout and requires no manual intervention, which greatly improves the system's autonomy and reliability.

[0099] The process of updating the model parameters of a digital twin model is as follows:

[0100] During the dynamic adjustment of the weighted fusion weights, the complete data sequence generated in each control cycle, including the previous state, the issued fusion control commands, the collected actual new state, and the calculated real-time prediction error, is continuously stored in a first-in-first-out (FIFO) data buffer with a fixed capacity. The data buffer capacity can be set according to the update frequency and computing resources, for example, it can store data from the most recent 100-500 control cycles. When the data buffer is full, or the average real-time prediction error of multiple consecutive control cycles stored in the buffer exceeds a preset trigger threshold, the online parameter update of the digital twin model is initiated.

[0101] Online parameter updates are performed by minimizing a composite loss function, the value of which is the sum of a first loss and a second loss. The first loss is the state prediction error term, calculated as follows: for each set of data stored in the data buffer, the actual new state recorded in that set of data is calculated, and the predicted state output after inputting the previous state and fused control commands from the same set of data into the digital twin model to be updated is calculated, and the square of the difference between the two is summed. This term drives the model to learn to accurately predict the dynamic behavior of the physical system. The second loss is the error consistency penalty term, calculated as follows: first, a penalty term weight coefficient greater than zero is preset for the digital twin model; this coefficient is used to balance the contributions of the two losses, for example, it can be set to 0.1; then, for each set of data stored in the data buffer, the following calculation is performed: take the instantaneous prediction error value corresponding to that set of data, multiply it by the digital twin model's prediction error value for the current set of data. Given the inputs of the "previous state" and the "fusion control command," an intermediate product is obtained by multiplying the square of the magnitude of the gradient vector of the output "predicted state" relative to the input "fusion control command." This calculation means that for data points with large historical prediction errors (instantaneous prediction error values), the square of the magnitude of the model's parameter update gradient (reflecting the degree of influence of parameter changes on the output) at that point will be amplified. This creates a "focused correction" mechanism: the model is forced to make greater adjustments to its parameters when updating those "state-action" regions that were previously inaccurately predicted. Finally, the second loss sum is obtained by multiplying the penalty term weight coefficient by the sum of the intermediate products corresponding to all data groups. By minimizing this composite loss function, the model not only pursues overall prediction accuracy but also focuses on correcting its performance in historically high error regions, thereby more efficiently narrowing the gap between the model and physical reality, especially correcting those systematic model biases that may lead to fusion control risks.

[0102] By optimizing the composite loss function, the model parameters of the digital twin model can be fine-tuned online and incrementally, thereby reducing the difference between the model prediction and the physical reality. This online model update based on real-time production data enables the digital twin model to continuously adapt to slow changes such as equipment wear and environmental drift, maintain its prediction fidelity, and provide a long-term reliable foundation for subsequent strategy optimization and security integration.

[0103] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0104] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0105] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0106] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0107] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0108] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0109] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A digital twin AI decision feedback communication system for a figurine production line, characterized in that, Specifically, it includes: The system consists of a state capture and simulation construction module, a virtual exploration and policy distillation module, a policy portability verification module, and a security fusion control module, which are connected in sequence. State capture and simulation construction module: In response to events where the real-time quality data of the physical production line does not meet the preset threshold, it synchronously captures the real-time data of all sensors and actuators in the physical production line, forming a state dataset containing the complete instantaneous state of the production line. In the digital twin environment, it uses the state dataset as the initial state, loads the current production line control strategy as the initial strategy, and dynamically generates a local reward function based on the real-time defect pattern, thereby constructing an initial virtual environment for reinforcement learning training. Virtual Exploration and Policy Distillation Module: In the initial virtual environment, the control agent conducts high-speed exploration based on the initial policy and combined with a directional perturbation mechanism, and collects the resulting interaction sequences as experience trajectory data. At the same time, using policy distillation technology, a lightweight set of optimization policy candidates is learned from the experience trajectory data. Strategy portability verification module: Each candidate strategy in the optimization strategy candidate set is placed in a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation, and the strategy performance degradation boundary prediction method is applied to quantitatively evaluate the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment. Based on the robustness evaluation and the expected performance loss, a final optimization strategy is selected. The safety fusion control module: It performs weighted fusion of the final optimization strategy and the current production line control strategy, generates fusion control commands and issues them to the actuators of the physical production line. At the same time, based on the difference between the actual state collected after the physical production line executes the fusion control commands and the predicted state predicted by the digital twin environment according to the previous state and the commands, it dynamically adjusts the dynamic fusion weight coefficients of the weighted fusion and updates the model parameters of the digital twin model.

2. The digital twin AI decision feedback communication system for a figurine production line according to claim 1, characterized in that: In the state capture and simulation construction module, the process of synchronously capturing real-time data from all sensors and actuators in the physical production line to form a state dataset containing the complete instantaneous state of the production line is as follows: When a high-resolution 3D vision sensor deployed at the end of the production line detects that the root mean square error of the paint film thickness uniformity of three consecutive workpieces exceeds a preset threshold, an event is triggered. After triggering, a synchronous acquisition command is sent to all data sources on the production line. The data sources include the high-resolution 3D vision sensor, the drive controller that controls the precision inkjet, the sensor that monitors environmental parameters, the servo controller that drives the robot axis movement, and the database interface connected to the manufacturing execution system. The collected real-time data is integrated and processed to form a state dataset, which consists of the following five types of information: The first type of information is the three-dimensional point cloud and defect coordinate annotations on the surface of the workpiece currently in production, obtained through a high-resolution three-dimensional vision sensor. The second type of information is the spatial position and orientation data of all robotic arm joint angles and end effectors fed back by the robot servo controller. The third type of information consists of real-time working air pressure, paint flow rate, and drive voltage parameters fed back by the precision inkjet drive controller. The fourth type of information consists of real-time temperature, humidity, and airborne dust particle concentration readings fed back by environmental sensors; The fifth type of information consists of all process specifications and parameters specified in the current production order, obtained through the Manufacturing Execution System (MES) interface.

3. The digital twin AI decision feedback communication system for a figurine production line according to claim 2, characterized in that: The process of constructing an initial virtual environment for reinforcement learning training is as follows: Using the instantaneous state of the production line represented by the state dataset as the initial state, the digital twin model is driven to construct a simulation environment instance in the virtual space that is consistent with the physical production line state. At the same time, the current production line control policy being used is read from the production line control server, and the current production line control policy is used as the initial policy of the agent. The parameters of the initial policy are then loaded into the agent's policy network. The analysis focuses on the defect coordinate annotations in the first type of information within the state dataset to identify the currently dominant defect type. Based on the identified dominant defect type, multiple quality evaluation index components most relevant to the dominant defect type are selected from a predefined reward function component library. For each selected quality evaluation index component, a corresponding virtual sensor is set up in the simulation environment instance of the digital twin model to calculate the index value of the quality evaluation index in any given subsequent state, and a corresponding target setting value is preset. The calculation process of the local reward function is as follows: A1. For each selected quality evaluation index component, calculate the difference between the index calculation value obtained by the virtual sensor in a given subsequent state and the target setting value corresponding to the index component. A2. For each difference obtained through A1, input it into a preset non-linear mapping function for processing. The non-linear mapping function is used to map the input value into a smooth output value. A3. Assign a weight that is dynamically adjusted according to the dominant defect type to each quality evaluation index component. Multiply the output value of each index component in A2 after processing by the nonlinear mapping function by the weight assigned to the index component to obtain the weighted result of the index component. A4. Sum the weighted results of all quality evaluation indicator components to obtain a preliminary reward item; A5. Dynamically calculate a reward focus factor between zero and one, the magnitude of which depends on the severity of the dominant defect type; A6. Multiply the preliminary reward items obtained through A4 with the reward focus factor obtained through A5 to obtain the first product focusing on the correction of the dominant defect; A7. Calculate the difference between the number 1 and the reward focus factor. Multiply this difference by a basic reward term used to maintain basic production efficiency and energy consumption levels to obtain a second product used to balance overall performance. The basic reward term is calculated based on the state and actions taken by the agent in the simulation environment instance. A8. Add the first product obtained through A6 to the second product obtained through A7. The sum is the value of the local reward function for a given state, action, and subsequent state. In digital twin software, an interactive reinforcement learning training environment instance is instantiated by using the state dataset as the initial state, the initial policy loaded into the policy network as the basis for the agent's behavior, and the logic of dynamically calculating the local reward function according to A1 to A8 as the environmental feedback mechanism. This training environment instance, the initial policy that forms the basis of behavior, the state dataset that serves as the source of the initial state, and the dominant defect type information used to guide the calculation of the local reward function are collectively encapsulated to form a complete initial virtual environment.

4. The digital twin AI decision feedback communication system for a figurine production line according to claim 3, characterized in that: In the virtual exploration and strategy distillation module, the process of controlling the agent to conduct high-speed exploration based on an initial strategy and combined with a directional perturbation mechanism, and to collect experience trajectory data, is as follows: The agent obtains an initial policy as the basis for behavior from the initial virtual environment and configures the parameters of the initial policy as the initial parameters of the agent's policy network. In the simulation environment instance of the initial virtual environment, the agent explores the interaction according to a hybrid action selection strategy. The hybrid action selection strategy selects the action directly calculated by the current agent's policy network based on the environmental state presented by the simulation environment instance with a preset high probability, and at the same time selects to execute a directional perturbation action with a preset low probability. During the exploration process, the agent executes a hybrid action selection strategy in the simulation environment instance at a speed far exceeding the real-time operating clock of the physical production line, and interacts with the simulation environment instance in multiple rounds. It continuously records the state of the simulation environment instance at each moment, the actions taken by the agent, the immediate rewards fed back by the simulation environment instance, and the state of the next simulation environment instance after the interaction. All these interaction records arranged in chronological order together constitute the experience trajectory data.

5. The digital twin AI decision feedback communication system for a figurine production line according to claim 4, characterized in that: The process of generating the directional disturbance action is as follows: A knowledge base dynamically maintained based on historical data of process and defect correlation accumulated during long-term production line operation is queried, which records the correlation strength between defect types and action dimensions in matrix form. Rows correspond to different defect types, and columns correspond to different action dimensions controlled by the production line. Matrix element values ​​represent the historical correlation strength of a specified defect type on a specified action dimension adjustment. For the dominant defect type corresponding to the current optimization task, a corresponding row value is extracted from the knowledge base as a sensitivity vector. A random direction noise vector with the same dimension as the action space is generated. The sensitivity vector and the random direction noise vector are multiplied element-wise to obtain a weighted noise vector. The weighted noise vector is multiplied by a preset perturbation amplitude coefficient to obtain a directional perturbation value. The directional perturbation value is then vector-added with the original action value calculated by the agent's policy network based on the environmental state presented by the current simulation environment instance. The result is the directional perturbation action.

6. The digital twin AI decision feedback communication system for a figurine production line according to claim 5, characterized in that: The process of learning and generating a lightweight set of optimization policy candidates from empirical trajectory data using policy distillation techniques is as follows: First, based on empirical trajectory data, a state value evaluation network and an advantage function estimator are trained to evaluate the merits of state-action pairs, where a state-action pair refers to the combination of the environmental state recorded in the empirical trajectory data and the action taken by the agent. Then, using the trained advantage function estimator, the advantage value of all state-action pairs in the empirical trajectory data is evaluated. The advantage value is used to quantify the expected improvement of taking a specified action in a specified state relative to the average level, and state-action pairs with advantage values ​​higher than a set threshold are selected to form a high-quality subset of empirical data. Then, several student policy networks with simpler structures than the agent policy network are initialized; these student policy networks are trained using policy distillation techniques, with the training objective being to minimize a composite loss function, which is calculated as the sum of a first weighted loss and a second weighted loss. By optimizing the composite loss function, each student policy network learns to imitate high-quality decision-making patterns. After training, each student policy network is independently and quickly evaluated in the initial virtual environment to obtain evaluation metrics including average cumulative reward and critical defect correction rate. Finally, all student policy networks are ranked according to the evaluation metrics, and several student policy networks with the best performance are selected. Their network parameters, evaluation metrics and performance scores are packaged together to form a lightweight optimization policy candidate set.

7. The digital twin AI decision feedback communication system for a figurine production line according to claim 6, characterized in that: In the policy portability verification module, the specific process of placing each candidate policy in the optimization policy candidate set into a high-fidelity verification simulation environment with injected noisy and delay models for robustness evaluation is as follows: Based on the initial state of the production line and the digital twin model represented by the state dataset obtained from the initial virtual environment, a high-fidelity verification simulation environment is constructed. Three types of non-ideal factor models learned from historical data of the physical production line are injected into the high-fidelity verification simulation environment. The first type is a sensor noise model, used to add random disturbances to the readings of virtual sensors that conform to the actual noise statistical characteristics of their physical counterparts. The second type is an actuator delay and response deviation model, used to simulate the dynamic characteristics of response lag and nonlinearity of actuators after receiving commands. The third type is a random environmental disturbance model, used to simulate unpredictable environmental disturbances in the production workshop, such as subtle airflow fluctuations. Each candidate strategy in the optimization strategy candidate set is loaded into the high-fidelity verification simulation environment, and multiple independent Monte Carlo simulation runs are performed starting from the initial state of the production line. Record and calculate the average performance index obtained by each candidate strategy across all simulation rounds. The average performance index includes the average cumulative reward or defect correction success rate. This average performance index serves as a numerical value to quantify the robustness evaluation result of the candidate strategy and is called the robustness evaluation score.

8. The digital twin AI decision feedback communication system for a figurine production line according to claim 7, characterized in that: The specific process for quantitatively evaluating the expected performance loss of each candidate strategy when migrating from the virtual environment to the physical environment is as follows: For each candidate policy in the optimization policy candidate set, run it in a simulation environment instance without injecting non-ideal factors, and collect the state access probability distribution induced by the policy decision behavior; for each accessed state, calculate the policy sensitivity of the candidate policy at that state. The policy sensitivity is defined by calculating the Jacobian matrix of the policy function of the candidate policy at the input state and obtaining the Frobenius norm of this Jacobian matrix. Meanwhile, using a confidence estimator trained in conjunction with the digital twin model to evaluate the model's predictive uncertainty, the confidence estimate of the digital twin model at each access state and the state-action pair consisting of the action output by the candidate policy in that state is evaluated. Next, for each candidate strategy, a value is calculated to quantify its expected performance loss, called the performance degradation prediction value. Based on the robustness evaluation score and performance degradation prediction value of each candidate strategy, a comprehensive decision is made to select the final optimization strategy. Specifically, the robustness evaluation scores of all candidate strategies in the optimization strategy candidate set are normalized so that all scores are mapped to a numerical range of zero to one; the performance degradation prediction values ​​of all candidate strategies in the optimization strategy candidate set are normalized so that all prediction values ​​are mapped to a numerical range of zero to one; then, a comprehensive decision score is calculated for each candidate strategy. The calculation process is as follows: the first weighting coefficient is multiplied by the normalized robustness evaluation score of the candidate strategy to obtain the first weighting term. Multiply the second weighting coefficient by the normalized performance degradation prediction value of the candidate strategy to obtain the second weighting term; subtract the second weighting term from the first weighting term, and the difference is the comprehensive decision score of the candidate strategy; finally, compare the comprehensive decision scores of all candidate strategies in the optimization strategy candidate set, select the candidate strategy with the highest comprehensive decision score, and determine it as the final optimization strategy.

9. The digital twin AI decision feedback communication system for a figurine production line according to claim 8, characterized in that: In the aforementioned security fusion control module, the process of dynamically adjusting the dynamic fusion weight coefficients for weighted fusion is as follows: Set a dynamically changing dynamic fusion weight coefficient with an initial value of zero, indicating complete reliance on the current production line control strategy; in each control cycle, load and run the final optimization strategy and the current production line control strategy respectively; based on the actual state that is consistent with the state dataset structure formed by real-time collection and processing of various sensors on the physical production line, obtain the actions suggested by the final optimization strategy and the actions suggested by the current production line control strategy. The action suggested by the final optimization strategy is recorded as the final optimization action, and the action suggested by the current production line control strategy is recorded as the current production line action; Based on the dynamic fusion weight coefficient, the final optimized action and the current production line action are weighted and fused to generate fusion control instructions; At the same moment the fusion control command is issued to the actuator, a digital twin model is used to simulate and calculate the predicted state to be reached after execution, based on the actual physical production line state collected at the end of the previous control cycle and the fusion control command to be issued. After the physical production line is completed, the actual new state is collected through the sensor network. This actual new state is a data set that is similar to the state dataset and reflects the instantaneous state of the production line after execution. The difference between the predicted state and the actual new state in the dimension characterizing the key process quality indicators of the production line is calculated to obtain the instantaneous prediction error. The instantaneous prediction error is compared with a preset safety error threshold, and the value of the dynamic fusion weight coefficient is dynamically updated based on the comparison result and the preset adjustment step size rule. When the dynamic fusion weight coefficient reaches a value of one and remains stable in subsequent consecutive control cycles, or when the production line control strategy changes due to external instructions, the weighted fusion process is considered complete. After the fusion process is completed, the final optimized strategy is set as the new current production line control strategy, and the state of the safety fusion control module is reset to prepare for the fusion deployment of the next optimized strategy.

10. The digital twin AI decision feedback communication system for a figurine production line according to claim 9, characterized in that: The process of updating the model parameters of the digital twin model is as follows: During the dynamic adjustment of the weighted fusion weights, the complete data sequence generated in each control cycle, including the previous state, the issued fusion control commands, the collected actual new state, and the calculated real-time prediction error, is continuously stored in a first-in-first-out data buffer with a fixed capacity. When the data buffer is full, or when the average value of the real-time prediction errors of multiple consecutive control cycles stored in the buffer exceeds a preset trigger threshold, the online parameter update of the digital twin model is initiated. Online parameter updates are performed by minimizing a composite loss function; By optimizing the composite loss function, online, incremental fine-tuning of the model parameters of the digital twin model can be achieved, thereby reducing the difference between model predictions and physical reality.