Virtual simulation automatic production line control method
By constructing a digital twin model and a reinforcement learning agent, real-time control of the virtual simulation system is achieved, solving the problem of real-time synchronization between virtual simulation and physical production line. This enables equipment fault prediction and production optimization, improving response speed and autonomous decision-making capabilities.
Patent Information
- Application Number
- CN202511550432.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-28
- Publication Date
- 2026-02-03
AI Technical Summary
Existing virtual simulation technology cannot achieve deep integration of virtual and real worlds, lacks online learning and autonomous decision-making capabilities, cannot reflect equipment wear and material changes in real time, and relies on manual intervention and fixed rules, resulting in slow response speed and difficulty in finding dynamic optimal solutions among multiple objectives.
A digital twin model is constructed, data is collected in real time through a sensor network, a reinforcement learning agent is deployed to make decisions, and an immersive interactive interface is provided to realize the real-time control and autonomous optimization of the virtual simulation system.
It achieves real-time synchronization between the virtual simulation system and the physical production line, can predict equipment failures and production bottlenecks, dynamically adjust production, find the balance point of multiple objectives, has adaptive learning capabilities, and improves response speed and decision-making efficiency.
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of virtual simulation production line, and particularly relates to a virtual simulation automatic production line control method. BACKGROUND
[0002] Virtual simulation technology has been widely used in modern industrial production line design and planning. The traditional virtual simulation method mainly has the following limitations:
[0003] 1. The existing technology mostly uses virtual simulation as an offline verification tool in the production line design stage. Once the production line is put into actual operation, the virtual model is disconnected from the physical entity, and cannot reflect the real-time state of equipment wear and tear, material changes, and the like, and cannot dynamically guide production based on real-time data.
[0004] 2. The traditional simulation system lacks advanced AI algorithms and cannot perform predictive analysis on potential faults and production bottlenecks. It also cannot generate and execute optimal adjustment strategies autonomously after a disturbance occurs.
[0005] 3. The optimization of the existing scheme mostly relies on manual intervention and logic based on fixed rules, which is slow in response and difficult to find a dynamic optimal solution among multiple objectives.
[0006] 4. The interaction between the operator and the simulation system is mostly passive monitoring and manual parameter setting, and the system cannot actively provide decision-making suggestions or immersive collaborative work training.
[0007] Therefore, there is a need for a simulation control system that can realize deep integration of virtual and real, has online learning and autonomous decision-making capabilities. SUMMARY
[0008] To solve the above problems, the application provides a virtual simulation automatic production line control method.
[0009] To achieve the above functions, the technical scheme adopted by the application is as follows:
[0010] A virtual simulation automatic production line control method, comprising the following steps:
[0011] (I) Building a digital twin model: building a dynamic digital twin model corresponding to the physical production line;
[0012] (II) Real-time data perception and mapping: collecting the running data of the physical production line in real time through a sensor network, and driving the digital twin model to realize synchronous running;
[0013] (III) Reinforcement learning decision: in the environment of the digital twin model, a reinforcement learning agent is deployed, which learns and makes decisions according to the state space composed of the production line state, the action space composed of the control instruction, and the multi-objective reward function, and outputs the optimal control strategy;
[0014] (IV) Strategy issuing and execution: the optimal control strategy is issued to the actuators of the physical production line to realize real-time regulation and control of the production line;
[0015] (V) Self-evolution: the result data after strategy execution is fed back to the digital twin model and the reinforcement learning agent for model correction and continuous learning of the agent, and the equipment life prediction model and performance degradation model in the digital twin model are adaptively updated.
[0016] Further, the state space includes at least one of real-time state of equipment, work-in-process quantity, material queue length, order urgency, and environmental parameters.
[0017] Further, the action space includes at least one of scheduling decision, equipment operating parameter adjustment, maintenance instruction triggering, and energy management strategy.
[0018] Further, the multi-objective reward function is a weighted function related to production efficiency, equipment energy consumption, product quality, order delivery timeliness, and equipment health degree.
[0019] Further, the reinforcement learning decision step specifically includes: using the digital twin model as a simulation environment to perform large-scale offline training and online fine-tuning of the reinforcement learning agent.
[0020] Further, the method further includes a human-machine collaborative interaction step: providing an immersive interaction interface through which an operator intervenes, reviews or modifies the control strategy output by the reinforcement learning agent.
[0021] Further, the immersive interaction interface is a virtual reality, augmented reality or mixed reality interface.
[0022] Further, in the self-evolution step, the feedback data is used to adaptively update the equipment life prediction model and performance degradation model in the digital twin model
[0023] An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the steps in the method when executing the program.
[0024] A computer-readable storage medium having a computer program stored thereon, wherein the program is executed by a processor to implement the steps in the method.
[0025] The present application takes the beneficial effects as follows:
[0026] 1. Breaks the barrier between virtual and physical worlds, making the simulation system change from a design tool to a real-time control brain that can dynamically adjust the production line at the millisecond level.
[0027] 2. Based on historical and real-time data, the digital twin model can predict equipment failures and production bottlenecks in virtual space, and the reinforcement learning agent can develop countermeasures in advance to achieve predictive maintenance and self-healing of production disturbances.
[0028] 3. Through the well-designed multi-objective reward function, the system can find a dynamic balance point among efficiency, cost, energy consumption and other mutually contradictory indicators, and achieve global and long-term optimization of the production system.
[0029] 4. The system can continuously learn and optimize as production data, equipment status and product types evolve, without the need for manual reprogramming, and has strong adaptive ability. DETAILED DESCRIPTION
[0030] The technical solutions of the present application will be described below in a clear and complete manner. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0031] Embodiment 1
[0032] The present application is implemented with an automobile engine assembly line as an example:
[0033] (I) Construct a digital twin model: use high-precision three-dimensional modeling and physics engines to construct a digital twin body of an engine assembly line containing six robot assembly stations, three AGVs and a roller conveyor, the model integrates the kinematics and dynamics attributes of the robots;
[0034] (II) Real-time data perception and mapping: install vibration sensors and current sensors on each robot, install GPS modules and power meters on AGVs, and install RFID readers at key stations to track engine identity, all data is synchronized in real time to the digital twin body through the industrial Internet gateway;
[0035] (III) Reinforcement learning decision-making:
[0036] State space: queue length in front of each station, average joint temperature of the robot, remaining battery power of the AGV, and remaining production time of the current order;
[0037] Action space: assign AGV to go to the Xth station, adjust the robot operating mode to "standard
[0038] “Predictive Maintenance” flag for a specific robot;
[0039] Reward function R: R = (number of engines completed today * 10) - (total energy consumption * 0.1) - (number of delayed orders * 50) + (number of failures successfully predicted and avoided * 100);
[0040] Offline training in the digital twin for 3 months using the Proximal Policy Optimization algorithm,
[0041] Subsequent online operation and continuous fine-tuning;
[0042] (IV) Strategy issuance and execution: when the system predicts that the vibration trend of the No. 2 assembly robot is abnormal, the agent decides to “maintain it after 30 minutes and dynamically allocate its tasks to robots 1 and 3”, and the instruction is automatically executed through the PLC control system;
[0043] (V) Self-evolution: after each maintenance, the actual wear data of the robot is used to update the remaining life prediction algorithm in its digital twin model.
[0044] Example 2
[0045] Embodiment 2
[0046] The present application is implemented with a multi-species, small-batch electronic product mounting production line as an example:
[0047] (I) Constructing a digital twin model: a flexible production line digital twin is constructed, including a high-speed chip mounter, an optical detection device, and a reflow soldering furnace, focusing on simulating the influence of the temperature field of the hot air reflow soldering furnace on the soldering quality;
[0048] (II) Real-time data perception and mapping: thermocouples are deployed in the 10 temperature zones of the reflow soldering furnace to collect temperature data in real time, and the optical detection device associates the image and position information of solder joint quality defects (such as virtual welding and bridge) to the corresponding PCB model in the digital twin in real time;
[0049] (III) Reinforcement learning decision-making:
[0050] State space: real-time temperature of each temperature zone, current PCB model, mounting precision data, and environmental humidity;
[0051] Action space: fine-tuning the set temperature of each temperature zone of the reflow soldering furnace and adjusting the conveyor belt speed;
[0052] Reward function R: R = (number of good products * 5) - (number of defective products * 20) - (furnace energy consumption * 0.05);
[0054] When the system is running, the operator wears AR glasses. When the agent suggests that the temperature of the 5th temperature zone be raised by 5℃ to improve the welding quality of a new product, the suggestion is superimposed on the AR view of the physical equipment in the form of a virtual label.
[0055] The operator can approve, reject or modify the parameter (for example, change the raise to 3℃) according to his own experience through gesture interaction. The operator's decision will be fed back to the reinforcement learning agent as a new training sample to accelerate its learning process.
[0056] (4) Strategy issuance and execution: The temperature control strategy confirmed or modified by the operator is automatically issued to the temperature control system of the reflow soldering furnace for execution.
[0057] (5) Self-evolution: The system specially records the conditions under which the operator modifies the AI's suggestion, thereby learning the implicit knowledge of human experts and making subsequent decision suggestions more and more in line with expert expectations.
[0058] The above describes the present application and its embodiments, which are not limiting. In general, if a person skilled in the art is inspired by it, without departing from the purpose of the present application, similar structural ways and embodiments can be designed without creative design, which should belong to the protection scope of the present application.
Claims
1. A virtual simulation automated production line control method, characterized in that, Includes the following steps: (i) Constructing a digital twin model: Constructing a dynamic digital twin model corresponding to the physical production line; (ii) Real-time data perception and mapping: Collecting the operation data of the physical production line in real time through a sensor network and driving the digital twin model to achieve synchronous operation; (III) Reinforcement Learning Decision Making: Deploying reinforcement learning agents in the environment of a digital twin model. The intelligent agent learns and makes decisions based on the state space consisting of the production line status, the action space consisting of control commands, and the multi-objective reward function, and outputs the optimal control strategy; (iv) Strategy distribution and execution: The optimal control strategy is distributed to the actuators of the physical production line to realize real-time control of the production line; (v) Self-evolution: Feedback the result data after the strategy is executed to the digital twin model and reinforcement learning agent for model correction and continuous learning of the agent, and adaptively update the equipment life prediction model and performance degradation model in the digital twin model.
2. A virtual simulation automated production line control method according to claim 1, characterized in that, The state space includes at least one of the following: real-time equipment status, work-in-process quantity, material queue length, order urgency, and environmental parameters.
3. A virtual simulation automated production line control method according to claim 1, characterized in that, The action space includes at least one of scheduling decisions, equipment operating parameter adjustments, maintenance command triggering, and energy management strategies.
4. A virtual simulation automated production line control method according to claim 1, characterized in that, The multi-objective reward function is a weighted function of production efficiency, equipment energy consumption, product quality, order delivery timeliness, and equipment health.
5. A virtual simulation automated production line control method according to claim 1, characterized in that, The reinforcement learning decision-making steps specifically include: using the digital twin model as a simulation environment to perform large-scale offline training and online fine-tuning on the reinforcement learning agent.
6. A virtual simulation automated production line control method according to claim 1, characterized in that, The method also includes a human-computer collaborative interaction step: providing an immersive interactive interface through which the operator intervenes, reviews, or modifies the control strategy output by the reinforcement learning agent.
7. A virtual simulation automated production line control method according to claim 6, characterized in that, The immersive interactive interface is a virtual reality, augmented reality, or mixed reality interface.
8. A virtual simulation automated production line control method according to claim 1, characterized in that, In the self-evolution step, feedback data is used to adaptively update the equipment life prediction model and performance degradation model in the digital twin model.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method as described in any one of claims 1 to 8.