Optical crystal production dynamic optimization control system based on defect monitoring
By constructing a dynamic optimization control system based on defect monitoring, the problems of defect response lag, model accuracy drift, and data silos in optical crystal production were solved, achieving high precision, stabilization, and continuous optimization of the optical crystal production process, and improving production consistency and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 宁波翌波光电科技有限公司
- Filing Date
- 2026-03-06
- Publication Date
- 2026-05-08
AI Technical Summary
Existing optical crystal production control systems have shortcomings in terms of defect response closed-loop performance, autonomous decision-making adaptability, long-term operational stability, and cross-domain data collaboration capabilities, making it difficult to achieve high precision, stability, and continuous optimization.
A dynamic optimization control system based on defect monitoring is adopted, which combines a defect monitoring module, a reinforcement learning decision-making module, an in-situ repair execution module, a process optimization module, a data feedback module, a precision correction module, and a federated learning collaboration module to construct an inner real-time closed loop and an outer cross-domain collaborative closed loop, thereby realizing the coordinated linkage between defect response and process adjustment.
It achieves closed-loop dynamic optimization of the entire production process, improves decision-making autonomy and long-term operational accuracy and stability, enhances large-scale production and cross-scenario adaptability, and takes into account the comprehensive optimization effect of crystal quality and production efficiency.
Smart Images

Figure CN121806513B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial process control and intelligent manufacturing technology, and in particular to a dynamic optimization control system for optical crystal production based on defect monitoring. Background Technology
[0002] Optical crystals are widely used in high-end technology fields such as lasers, nonlinear optical devices, precision optoelectronic systems, and semiconductor manufacturing. The integrity of their internal structure and the consistency of their growth directly affect the stability, reliability, and long-term service life of the devices. During crystal growth and subsequent processing, various types of defects are easily generated inside and on the surface of the crystal due to multiple factors such as fluctuations in raw materials, changes in equipment operating status, and environmental disturbances. These defects often have evolutionary and cumulative characteristics, which places higher demands on the dynamic control capabilities of the production process.
[0003] With the expansion of production scale and the continuous improvement of crystal performance indicators, optical crystal manufacturing has gradually evolved from traditional experience-based process adjustment to data-driven and intelligent control. However, from the perspective of the overall production process, the existing production control system still faces many technical bottlenecks in defect detection, decision response, and cross-process collaboration, making it difficult to meet the requirements of high precision, continuous operation, and adaptive optimization.
[0004] On the one hand, the generation and evolution of defects in crystal production are often highly coupled with process parameters, equipment status, and environmental conditions. If defect monitoring results cannot be effectively linked with subsequent control decisions, it can easily lead to delayed control responses or a single treatment method, making it difficult to balance local defect suppression with overall process stability. In complex production scenarios, relying solely on static rules or manually set parameters makes it difficult to achieve differentiated responses to different defect characteristics, and it is also not conducive to forming a continuous feedback closed-loop optimization mechanism.
[0005] On the other hand, the multi-source uncertainties in the production process have obvious stage-specific and hidden characteristics, such as fluctuations in raw material quality, equipment performance degradation, and environmental disturbances. The degree of impact of these factors on crystal quality varies significantly at different production stages. If the control model lacks dynamic correction capabilities, decision bias and decreased control accuracy are likely to occur during long-term operation, affecting production consistency and yield.
[0006] Furthermore, under conditions of multi-batch or large-scale production, the data samples available to a single production unit are limited, and model training and optimization are easily constrained by insufficient samples. At the same time, the data generated by different production units, upstream and downstream processes, and supply chain links have significant collaborative value, but due to data security and privacy protection requirements, relevant information is difficult to share directly, resulting in cross-domain data not being able to fully participate in overall optimization decisions and limiting the global collaborative capabilities of the production system.
[0007] In summary, existing optical crystal production control systems still have significant shortcomings in terms of defect response closed-loop performance, autonomous decision-making adaptability, long-term operational stability, and cross-domain data collaboration capabilities. There is an urgent need for a dynamic optimization control mechanism that can integrate defect monitoring, intelligent decision-making, dynamic correction, and multi-source data collaboration to achieve high precision, stability, and continuous optimization in the optical crystal production process. Summary of the Invention
[0008] To address the shortcomings of existing technologies, the present invention aims to provide a dynamic optimization control system for optical crystal production based on defect monitoring. This system integrates an inner real-time closed loop with an outer cross-domain collaborative closed loop to simultaneously address issues such as defect response lag, model accuracy drift, and data silos, thereby achieving adaptive optimization and global collaborative control of the production process.
[0009] To achieve the above objectives, the present invention provides the following technical solution: a dynamic optimization control system for optical crystal production based on defect monitoring, comprising:
[0010] The defect monitoring module is used to collect production data, extract features, and output defect feature vectors.
[0011] The reinforcement learning decision module, connected to the defect monitoring module, is used to receive the defect feature vector and production status parameters, and output decision instructions including in-situ repair actions and / or process optimization actions through the reinforcement learning model.
[0012] The in-situ repair execution module, connected to the reinforcement learning decision module, is used to execute the in-situ repair action and collect repair data;
[0013] The process optimization module is connected to the reinforcement learning decision module and is used to execute the process optimization actions and collect process data.
[0014] The data feedback module is connected to the in-situ repair execution module and the process optimization module, respectively, and is used to store the repair data and process data and feed them back to the reinforcement learning decision module.
[0015] The accuracy correction module, connected to both the defect monitoring module and the reinforcement learning decision module, includes:
[0016] The variable acquisition unit is used to collect data on raw material purity, equipment vibration, environmental data, and growth dynamics to form the acquired data.
[0017] The correction factor unit, connected to the variable acquisition unit, is used to generate correction factors based on the acquired data and historical data and send them to the reinforcement learning decision module.
[0018] The federated learning collaboration module, which connects the reinforcement learning decision module and the accuracy correction module respectively, includes:
[0019] The data acquisition unit is used to acquire cross-domain data from external systems and the supply chain.
[0020] A collaborative optimization unit, connected to the data acquisition unit, is used to perform privacy aggregation on the cross-domain data through federated learning, generate global optimization parameters, and send them to the reinforcement learning decision module.
[0021] The reinforcement learning decision module dynamically corrects the decision instructions based on the correction factor and the global optimization parameters.
[0022] Furthermore, the reinforcement learning decision-making module includes:
[0023] A state space construction unit is used to construct a state space containing defect feature vectors and production system state parameters, wherein the production system state parameters include at least crystal growth rate, furnace temperature distribution, raw material ratio, energy consumption, and production efficiency.
[0024] An action space construction unit is used to construct an action space that includes in-situ repair actions and process optimization actions. The in-situ repair actions include at least laser power, laser action time, and local temperature field gradient. The process optimization actions include at least growth rate adjustment, temperature adjustment, and raw material ratio adjustment.
[0025] The reward function construction unit is used to construct a composite reward function, which calculates a comprehensive reward value by weighting the defect repair quality, production efficiency and energy consumption indicators based on preset quality weight coefficient, efficiency weight coefficient and energy consumption weight coefficient.
[0026] The training and decision-making unit is connected to the state space construction unit, action space construction unit, and reward function construction unit, respectively, and is used to perform the training and real-time decision-making tasks of the reinforcement learning model.
[0027] Furthermore, the training and decision-making unit includes:
[0028] The model initialization subunit is used to load the historical training sample set to pre-train the reinforcement learning model.
[0029] The real-time decision-making subunit is used to receive real-time state parameters and output the optimal decision instruction based on the pre-trained model.
[0030] The parameter update subunit is used to iteratively update the network parameters of the reinforcement learning model based on feedback data from the data feedback module and correction factors from the accuracy correction module.
[0031] When outputting decision instructions, the real-time decision subunit synchronously receives and applies the correction factor sent by the accuracy correction module and the global optimization parameters sent by the federated learning collaboration module to dynamically correct the decision instructions.
[0032] Furthermore, the defect monitoring module includes:
[0033] A multimodal sensing unit is used to collect multi-source monitoring data of the optical crystal production process, wherein the multi-source monitoring data includes at least spectral data, ultrasonic data and infrared thermal imaging data;
[0034] A noise filtering unit, connected to the multimodal sensing unit, is used to process the multi-source monitoring data using a wavelet packet transform algorithm to filter out electromagnetic interference and equipment noise.
[0035] The feature extraction unit, connected to the noise filtering unit, is used to extract the core feature parameters of the defect based on the filtered multi-source monitoring data, and generate a defect feature vector containing the defect type, defect location, defect size and defect growth rate.
[0036] The defect feature vector is output to the reinforcement learning decision module at a sampling frequency not lower than a first preset threshold.
[0037] Furthermore, the in-situ repair execution module includes:
[0038] The positioning control unit is used to control the three-dimensional motion platform to position the focus of the repair device to the center of the defect according to the defect location information in the decision instruction, with a positioning accuracy not lower than the first positioning accuracy threshold.
[0039] The laser execution unit, connected to the positioning control unit, is used to control a high-precision pulsed laser to perform pulsed laser repair operations according to the laser power and laser action time parameters in the decision command;
[0040] The temperature field execution unit, connected to the positioning control unit, is used to control the partition temperature control device to adjust the local temperature field of the repair area according to the local temperature field gradient parameters in the decision instruction.
[0041] The repair device includes the high-precision pulsed laser and the zoned temperature control device. The laser execution unit and the temperature field execution unit synchronously collect temperature data and stress data of the repair area when performing the repair operation, and send the temperature data and stress data as repair process data to the data feedback module.
[0042] Furthermore, the process optimization module includes:
[0043] The parameter receiving unit is used to receive the process optimization actions contained in the decision instruction, wherein the process optimization actions include at least the crystal growth rate adjustment amount, the temperature adjustment amount, and the raw material ratio adjustment amount;
[0044] A control signal generation unit, connected to the parameter receiving unit, is used to convert the process optimization action into a corresponding control signal based on an adaptive PID compensation algorithm.
[0045] An execution adjustment unit, connected to the control signal generation unit, is used to drive the crystal growth furnace control system, raw material supply adjustment device, and energy consumption monitoring instrument to perform process parameter adjustments according to the control signal.
[0046] The process status acquisition unit is connected to the execution adjustment unit and is used to collect the adjusted growth rate, temperature distribution, raw material ratio and energy consumption data in real time, and send them to the data feedback module as process status data.
[0047] The adjustment accuracy of the raw material supply adjustment device is not lower than the first accuracy threshold, and the temperature control accuracy of the crystal growth furnace control system is not lower than the second accuracy threshold.
[0048] Furthermore, the data feedback module includes:
[0049] The data receiving unit is used to receive and integrate the repair process data and the process status data;
[0050] A data cleaning unit, connected to the data receiving unit, is used to perform outlier detection and filtering on the integrated data using a data cleaning algorithm.
[0051] A data storage unit, connected to the data cleaning unit, is used to store the cleaned data in a distributed database;
[0052] The sample encapsulation unit, connected to the data storage unit, is used to extract data from the distributed database at a preset period and encapsulate it into a training sample set containing state parameters, action instructions and reward values.
[0053] The training sample set is fed back to the reinforcement learning decision module for iterative updates of its model parameters.
[0054] Furthermore, the correction factor unit in the accuracy correction module includes:
[0055] The data preprocessing subunit is used to preprocess the raw material purity data, equipment vibration data, environmental data and growth dynamic data from the variable acquisition unit.
[0056] The weight allocation subunit, connected to the data preprocessing subunit, is used to allocate real-time weights to various types of dynamic data based on the adaptive mutual information entropy-random forest algorithm.
[0057] An anomaly handling subunit, connected to the weight allocation subunit, is used to identify and block abnormal data based on preset statistical criteria, and adjust the weight of the blocked data accordingly.
[0058] The correction factor calculation subunit is connected to the weight allocation subunit, the anomaly handling subunit, and the data feedback module, respectively. It is used to calculate the repair parameter correction factor and the process parameter correction factor based on the weighted dynamic data and the historical training sample set provided by the data feedback module through an online gradient descent algorithm.
[0059] The correction factor calculation subunit sends the generated correction factor to the reinforcement learning decision module with a response time not exceeding a preset delay threshold.
[0060] Furthermore, the data acquisition unit in the federated learning collaboration module includes:
[0061] The process data acquisition subunit is used to acquire raw material particle size data from the raw material pretreatment process, surface roughness data from the processing process, and optical performance data from the testing process.
[0062] The supply chain data interface subunit is used to obtain raw material batch quality data and equipment maintenance record data from the supply chain database and equipment maintenance management system via API interface;
[0063] The data acquisition unit integrates the collected raw material particle size data, surface roughness data, optical performance data, raw material batch quality data, and equipment maintenance record data into cross-domain correlated data and sends it to the collaborative optimization unit.
[0064] Furthermore, the collaborative optimization unit in the federated learning collaborative module includes:
[0065] The parameter aggregation subunit is used to perform homomorphic encryption processing on the cross-domain associated data based on the hierarchical federated learning architecture, and to aggregate model parameters from different production batches and processes to generate global optimization parameters.
[0066] The reward function adjustment subunit is used to extract supply chain features from the cross-domain associated data and generate an association weight vector, and to dynamically adjust the composite reward function of the reinforcement learning decision module using the association weight vector;
[0067] The cross-domain processing subunit is used to identify abnormal features in the cross-domain associated data through the isolated forest algorithm, locate the root cause of the anomaly based on the blockchain traceability node, and trigger alternative optimization schemes.
[0068] Specifically, the parameter aggregation subunit sends the global optimization parameters to the reinforcement learning decision module, and the reward function adjustment subunit sends the dynamically adjusted composite reward function to the reinforcement learning decision module.
[0069] The beneficial effects of this invention are:
[0070] 1. Achieve closed-loop dynamic optimization of the entire production process: By organically coupling defect monitoring, reinforcement learning decision-making, in-situ repair, process optimization, data feedback and dynamic precision correction, a continuous closed-loop control mechanism is constructed to achieve synergistic linkage between defect response and process adjustment, overcoming the problems of separation between repair and optimization and imbalance between local control and overall performance in the existing production process.
[0071] 2. Enhance decision-making autonomy and long-term operational accuracy and stability: The reinforcement learning decision-making module makes adaptive decisions based on defect feature vectors and production status parameters, eliminating the need for manual preset of fixed parameter rules; combined with the accuracy correction module, it dynamically corrects multi-source variables, effectively suppressing decision deviations caused by changes in raw materials, equipment and environment, and improving the prediction accuracy and control stability of the system in long-term operation.
[0072] 3. Enhance large-scale production and cross-scenario adaptability: Through the federated learning collaboration mechanism, collaborative optimization of data across batches, processes and domains can be achieved without exposing the original data, alleviating the problem of insufficient data in a single production unit, improving the model's generalization ability and multi-scenario adaptability, and making it suitable for optical crystal production of different scales and working conditions.
[0073] 4. Comprehensive optimization effect that balances crystal quality and production efficiency: Reinforcement learning decision-making covers both in-situ repair actions and process optimization actions, enabling defect control and indicators such as production efficiency and energy consumption to be optimized in a coordinated manner, avoiding the decline in overall production performance caused by adjusting only around a single quality target.
[0074] 5. The system has a clear structure and good engineering feasibility: the interfaces between the functional modules are clear and the hierarchy is clear. It can be integrated and deployed on the basis of the existing optical crystal production system, reducing the complexity of system modification and having good engineering application prospects. Attached Figure Description
[0075] Figure 1 This is a schematic diagram of the structure of the dynamic optimization control system for optical crystal production based on defect monitoring in this invention;
[0076] Figure 2This is a schematic diagram of the training and decision-making unit in this invention;
[0077] Figure 3 This is a schematic diagram of the structure of the correction factor unit in this invention;
[0078] Figure 4 This is a schematic diagram of the data acquisition unit in this invention;
[0079] Figure 5 This is a schematic diagram of the structure of the collaborative optimization unit in this invention.
[0080] Figure labels: 1. Defect monitoring module; 11. Multimodal sensing unit; 12. Noise filtering unit; 13. Feature extraction unit; 2. Reinforcement learning decision-making module; 21. State space construction unit; 22. Action space construction unit; 23. Reward function construction unit; 24. Training and decision-making unit; 241. Model initialization subunit; 242. Real-time decision-making subunit; 243. Parameter update subunit; 3. In-situ repair execution module; 31. Positioning control unit; 32. Laser execution unit; 33. Temperature field execution unit; 4. Process optimization module; 41. Parameter receiving unit; 42. Control signal generation unit; 43. Execution adjustment unit; 44. Process status acquisition unit. 5. Data Collection Unit; 51. Data Receiving Unit; 52. Data Cleaning Unit; 53. Data Storage Unit; 54. Sample Packaging Unit; 6. Precision Correction Module; 61. Variable Acquisition Unit; 62. Correction Factor Unit; 621. Data Preprocessing Subunit; 622. Weight Allocation Subunit; 623. Anomaly Handling Subunit; 624. Correction Factor Calculation Subunit; 7. Federated Learning Collaboration Module; 71. Data Acquisition Unit; 711. Process Data Acquisition Subunit; 712. Supply Chain Data Interface Subunit; 72. Collaborative Optimization Unit; 721. Parameter Aggregation Subunit; 722. Reward Function Adjustment Subunit; 723. Cross-Domain Processing Subunit. Detailed Implementation
[0081] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Identical components are denoted by the same reference numerals. It should be noted that the terms "front," "rear," "left," "right," "upper," and "lower" used in the following description refer to directions in the accompanying drawings, and the terms "bottom surface," "top surface," "inner," and "outer" refer to directions toward or away from the geometric center of a specific component, respectively.
[0082] like Figure 1 As shown in the figure, this embodiment provides a dynamic optimization control system for optical crystal production based on defect monitoring.
[0083] I. Application Scenarios of the Implementation Examples;
[0084] This embodiment takes the sapphire crystal blister production process as the application object. It utilizes defect monitoring, reinforcement learning decision-making, in-situ repair, process optimization, data feedback, accuracy correction and cross-domain collaboration mechanisms to build a dynamic optimization control system for the continuous production process of sapphire crystals, so as to realize real-time response to crystal defects, in-situ repair and collaborative optimization of process parameters.
[0085] II. Overall System Composition;
[0086] The dynamic optimization control system in this embodiment includes the following functional modules, and achieves data interconnection through industrial Ethernet:
[0087] Defect monitoring module 1 is used to collect production data and extract features, and output defect feature vectors;
[0088] The reinforcement learning decision module 2 is connected to the defect monitoring module 1. It is used to receive defect feature vectors and production status parameters, and output decision instructions including in-situ repair actions and / or process optimization actions through the reinforcement learning model.
[0089] The in-situ repair execution module 3 is connected to the reinforcement learning decision module 2 and is used to execute in-situ repair actions and collect repair data according to decision instructions.
[0090] Process optimization module 4, connected to reinforcement learning decision module 2, is used to execute process optimization actions and collect process data;
[0091] The data feedback module 5 is connected to the in-situ repair execution module 3 and the process optimization module 4 respectively, and is used to store repair data and process data and feed them back to the reinforcement learning decision module 2.
[0092] The accuracy correction module 6, connected to the defect monitoring module 1, the reinforcement learning decision-making module 2, and the data feedback module 5, includes:
[0093] The variable acquisition unit 61 is used to collect data on raw material purity, equipment vibration, environmental data, and growth dynamics.
[0094] Correction factor unit 62 is used to generate correction factors based on collected data and historical data and send them to reinforcement learning decision module 2;
[0095] The federated learning collaboration module 7 connects to the reinforcement learning decision-making module 2, the accuracy correction module 6, and the data feedback module 5, and includes:
[0096] Data acquisition unit 71 is used to acquire cross-domain data from external systems and the supply chain;
[0097] The collaborative optimization unit 72 is used to perform privacy aggregation on cross-domain data through federated learning, generate global optimization parameters, and send them to the reinforcement learning decision module 2.
[0098] The reinforcement learning decision module 2 dynamically corrects the decision instructions based on the correction factors from the accuracy correction module 6 and the global optimization parameters from the federated learning collaboration module 7, forming a two-layer optimization system with real-time closed loop and cross-domain collaborative closed loop.
[0099] III. Hardware Configuration and Operating Environment;
[0100] (a) Defect monitoring module 1;
[0101] This module is used to collect production data during the sapphire crystal growth process and extract features, outputting defect feature vectors. Its hardware and operating environment include:
[0102] Spectral sensor: Ocean Optics HR4000;
[0103] Ultrasonic testing equipment: Olympus EPOCH 650;
[0104] Infrared thermal imager: FLIR A655sc;
[0105] Data sampling frequency: 10Hz.
[0106] This module performs multimodal detection of lattice distortion, microcracks, and thermal anomalies within the crystal, extracts defect feature vectors, and uses them as one of the inputs to the reinforcement learning decision module 2.
[0107] (ii) Reinforcement Learning Decision Module 2;
[0108] This module receives defect feature vectors and production status parameters, and outputs decision instructions that include in-situ repair actions and / or process optimization actions. Its hardware and operating environment include:
[0109] Industrial servers:
[0110] CPU: Intel Xeon Gold 6330;
[0111] GPU: NVIDIA A100;
[0112] Memory: 6GB;
[0113] Software environment: TensorFlow 2.8.
[0114] (III) In-situ repair execution module 3;
[0115] This module is used to execute in-situ repair actions based on the decision instructions output by reinforcement learning decision module 2, and to collect repair data. Its hardware and operating environment include:
[0116] Pulsed laser: SPI Lasers redPOWER 100;
[0117] 3D motion platform: Yaskawa MP2300 (positioning error ≤5μm);
[0118] Zone temperature control device: Eurotherm 3504.
[0119] In-situ repair procedures include local laser annealing, micro-area stress release, and fine adjustment of the temperature field.
[0120] (iv) Process optimization module 4;
[0121] This module executes the process optimization actions output by reinforcement learning decision module 2 and collects corresponding process data. Its hardware and operating environment include:
[0122] Sapphire growth furnace: KYKY-XD-800;
[0123] Raw material supply regulating device;
[0124] Energy consumption monitoring instrument: Schneider PM8000.
[0125] Process optimization actions include adjusting drawing speed, correcting temperature gradient, and updating energy consumption control strategies.
[0126] (v) Data Feedback Module 5;
[0127] This module stores repair data and process data, and feeds the data back to reinforcement learning decision module 2. Its hardware and operating environment include:
[0128] Database: MongoDB 5.0 distributed database;
[0129] Storage capacity: 10TB;
[0130] Communication network: Profinet industrial Ethernet;
[0131] Network latency: ≤10ms.
[0132] (vi) Precision correction module 6;
[0133] This module is used to dynamically correct the prediction accuracy in the reinforcement learning decision-making process, including:
[0134] The variable acquisition unit 61, its hardware and operating environment include:
[0135] Raw material purity tester: Thermo Scientific iCE 3500;
[0136] Equipment vibration sensor: SZ-6-6A;
[0137] Environment and edge computing device: HUAWEI AR502H.
[0138] The variables collected include: raw material purity, equipment vibration amplitude, ambient temperature and humidity, and crystal growth dynamics data.
[0139] Correction factor unit 62, initial variable weight settings include:
[0140] Early growth stage: raw material purity 0.4, equipment vibration 0.2, environmental data 0.2, growth dynamics 0.2, correction factor constraint range: 0.9~1.1.
[0141] The generated correction factors are sent to reinforcement learning decision module 2 in real time to correct decision instructions.
[0142] (vii) Federated Learning Collaboration Module 7;
[0143] This module is used to achieve cross-domain collaborative optimization without sharing the original data. Its hardware and operating environment include:
[0144] Edge node server: Dell PowerEdge R750 (1 unit per batch / process);
[0145] Federated learning framework: FedML;
[0146] Blockchain system: Hyperledger Fabric;
[0147] Federated learning aggregation cycle: 50ms;
[0148] Local model update magnitude threshold: ≤15%.
[0149] Among them: the data acquisition unit 71 is used to acquire cross-domain data from external systems and supply chains;
[0150] The collaborative optimization unit 72 performs privacy-preserving aggregation on cross-domain data through federated learning, generates global optimization parameters, and sends them to the reinforcement learning decision module 2.
[0151] IV. Working principle of Example 1;
[0152] In the sapphire crystal production process using the bubble-forming method, the system operates according to the following working principle:
[0153] Defect perception stage: Defect monitoring module 1 synchronously collects spectral, ultrasonic and infrared thermographic data at a sampling frequency of 10Hz, extracts defect features such as lattice distortion and microcracks, forms defect feature vectors and sends them to reinforcement learning decision module 2 in real time.
[0154] Intelligent decision-making stage: Reinforcement learning decision module 2 receives defect feature vectors and current production status parameters, and at the same time receives correction factors from accuracy correction module 6 and global optimization parameters generated by federated learning collaboration module 7.
[0155] Based on a comprehensive consideration of local defect states and global optimization objectives, the reinforcement learning model outputs decision instructions that include in-situ repair actions and / or process optimization actions.
[0156] Execution and feedback phase: In-situ repair execution module 3 performs local laser repair and temperature control adjustment according to the decision instructions; process optimization module 4 adjusts the sapphire growth process parameters simultaneously.
[0157] The repair data and process data generated during the execution process are stored through the data feedback module 5 and fed back to the reinforcement learning decision module 2 in real time, forming the first closed loop.
[0158] Dynamic accuracy correction and cross-domain collaboration stage: Accuracy correction module 6 generates dynamic correction factors based on multi-source variables to continuously correct reinforcement learning decision biases;
[0159] The Federated Learning Collaboration Module 7 periodically aggregates data across batches, processes, and the supply chain to generate global optimization parameters, which are used to improve the generalization ability and stability of the decision-making model, forming a second-layer collaborative closed loop.
[0160] V. Implementation Results and Technical Achievements;
[0161] Under the above system configuration and working principle, this embodiment was tested for 30 days, producing 100 sapphire crystals, and achieving the following technical results:
[0162] The defect repair success rate reached 96.8%, with a lattice distortion repair rate of 98.2% and a microcrack repair rate of 95.3%.
[0163] Production efficiency increased by 12.5%, reaching the preset efficiency threshold of 10%;
[0164] Energy consumption per unit decreased by 8.3%, which is 9% lower than the preset energy consumption threshold.
[0165] The predicted defect repair success rate error is controlled within ±0.7%;
[0166] Under small-batch production conditions, the defect rate was reduced by 7.2%, and the cross-process parameter mismatch rate was reduced to 1.5%.
[0167] Under operating conditions with temperature and humidity fluctuations of ±5%, the defect rate fluctuation is ≤2.7%, and the long-term operating accuracy decay is only 0.4%.
[0168] VI. Summary of Implementation Examples;
[0169] This embodiment demonstrates that the defect monitoring-based dynamic optimization control system for optical crystal production described in this invention can achieve deep integration of defect monitoring, intelligent decision-making, in-situ repair, process optimization, dynamic correction, and cross-domain collaboration under complex production conditions, significantly improving the quality consistency, efficiency, and stability of sapphire crystal production, and possessing good engineering feasibility and industrial application value.
[0170] Example 2 is the second embodiment of the present invention.
[0171] I. Technical Positioning and Differences of Example 2;
[0172] Based on the overall system structure and operating environment in Example 1, this second embodiment further details the internal algorithm implementation mechanism of the reinforcement learning decision module 2. It focuses on clarifying the specific implementation methods of state space construction, action space construction, composite reward function design, training and real-time decision-making logic, and dynamic correction mechanism. It is applicable to sapphire crystal bubble growth method and other optical crystal production scenarios with similar growth kinetics.
[0173] II. The overall structure of reinforcement learning decision-making module 2;
[0174] Reinforcement learning decision-making module 2 includes the following functional units:
[0175] The state space construction unit 21 is used to construct a state space containing defect feature vectors and production system state parameters. The production system state parameters include at least crystal growth rate, furnace temperature distribution, raw material ratio, energy consumption and production efficiency.
[0176] Action space construction unit 22 is used to construct an action space that includes in-situ repair actions and process optimization actions. The in-situ repair actions include at least laser power, laser action time and local temperature field gradient, and the process optimization actions include at least growth rate adjustment amount, temperature adjustment amount and raw material ratio adjustment amount.
[0177] The reward function construction unit 23 is used to construct a composite reward function. The composite reward function calculates a comprehensive reward value by weighting the defect repair quality, production efficiency and energy consumption indicators based on preset quality weight coefficient, efficiency weight coefficient and energy consumption weight coefficient.
[0178] The training and decision-making unit 24 is connected to the state space construction unit 21, the action space construction unit 22, and the reward function construction unit 23, respectively, and is used to perform the training and real-time decision-making tasks of the reinforcement learning model.
[0179] Reference Figure 2 The training and decision-making unit 24 includes:
[0180] Model initialization subunit 241 is used to load the historical training sample set to pre-train the reinforcement learning model;
[0181] Real-time decision subunit 242 is used to receive real-time state parameters and output the optimal decision instruction based on the pre-trained model;
[0182] The parameter update subunit 243 is used to iteratively update the network parameters of the reinforcement learning model based on the feedback data from the data feedback module 5 and the correction factor from the accuracy correction module 6.
[0183] Among them, when the real-time decision-making subunit 242 outputs decision instructions, it synchronously receives and applies the correction factor sent by the accuracy correction module 6 and the global optimization parameters sent by the federated learning collaboration module 7 to dynamically correct the decision instructions.
[0184] The reinforcement learning model adopts a deep reinforcement learning structure that integrates DQN and PPO. DQN is used for discrete action value estimation, and PPO is used for continuous action policy optimization to improve stability and convergence speed in a multi-dimensional continuous action space.
[0185] The structural configuration of the reinforcement learning model includes:
[0186] Input layer dimensions: 15 (4-dimensional defect features + 11-dimensional production status parameters);
[0187] Hidden layer structure: 3 layers (128 / 64 / 32 neurons);
[0188] Output layer dimension: 9;
[0189] In-situ repair action A1: 3D;
[0190] Process optimization action A2: 6 dimensions;
[0191] Learning rate: 0.001;
[0192] Number of training iterations: 5000;
[0193] Convergence criterion: Reward value ≥ 80 points.
[0194] III. Methods for constructing the state space;
[0195] The state space is composed of defect feature vectors and production system state parameters, and is uniformly represented as a high-dimensional state vector:
[0196] The defect feature vector includes the degree of lattice distortion, microcrack density, uniformity of defect spatial distribution, and thermal stress fluctuation index.
[0197] The production system status parameters include at least: crystal growth rate, furnace zone temperature distribution, raw material ratio, energy consumption, and production efficiency.
[0198] The state space construction unit 21 performs time synchronization and normalization processing on the above multi-source heterogeneous data, and uses it as the input state of the reinforcement learning model to reflect the comprehensive operating state of the current crystal growth process.
[0199] IV. Methods for constructing the action space;
[0200] The action space consists of a combined action space composed of in-situ repair actions and process optimization actions:
[0201] In-situ repair procedures include at least: laser power, laser treatment time, and local temperature gradient.
[0202] Process optimization actions include at least: adjustments to the growth rate, adjustments to the temperature of each zone of the furnace, and adjustments to the raw material ratio.
[0203] The action space construction unit 22 maps the above action parameters into continuous action vectors, enabling the reinforcement learning model to output local repair and global process adjustment instructions simultaneously within the same decision cycle, thereby achieving collaborative decision-making for defect repair and process optimization.
[0204] V. The construction logic of the composite reward function;
[0205] Reward function construction unit 23 constructs a composite reward function to quantitatively evaluate the decision results of the reinforcement learning model. Its basic form is as follows:
[0206] ;
[0207] in, The output value of the composite reward function represents the comprehensive reward obtained by the reinforcement learning model within the current decision-making cycle;
[0208] , , These are the preset first weight coefficient, second weight coefficient, and third weight coefficient, respectively;
[0209] This parameter maps the defect repair success rate and is obtained by proportionally normalizing the actual repair success rate. When the repair success rate is not less than 95%, it takes a positive value greater than 1.
[0210] This is a scalar measure of the current defect severity, calculated from the defect feature vector.
[0211] This serves as the baseline value for target defect control and a preset reference value for the system.
[0212] This is the defect tolerance scaling factor, used to adjust the sensitivity of the compensation function;
[0213] The production efficiency improvement index is calculated as the ratio of real-time efficiency to baseline efficiency.
[0214] The percentage increase in efficiency represents the relative change in efficiency within the current period.
[0215] The modulus parameter of the first-kind complete elliptic integral is used to characterize the nonlinear characteristics of efficiency as a function of process parameters.
[0216] The order parameter of the Bessel function is used to describe the oscillation characteristics of energy consumption caused by process adjustments;
[0217] Energy intensity per unit time, derived from energy consumption monitoring data;
[0218] This is the energy consumption amplification factor, calculated from the relative deviation of energy consumption from the benchmark.
[0219] The rationale and physical meaning of the selection of various complex functions in the composite reward function:
[0220] Gamma function : Used to amplify the positive reward after the defect repair success rate threshold (95%) is met or exceeded, so that the model still has room for further optimization after reaching the quality constraints, thereby avoiding the problem of "stopping optimization as soon as the target is met".
[0221] Complementary error function: This is used to smooth out the penalty for defect severity. When the defect severity is higher than the target baseline, the reward value decays rapidly, enhancing the model's sensitivity to high-risk defect states.
[0222] Riemann Zeta function It is used to characterize the long-term benefits of improved production efficiency, enabling the efficiency improvement to have a cumulative amplification effect in multi-cycle decision-making, and guiding the model to learn stable and efficient combinations of process parameters.
[0223] Complete Elliptic Integral Used to express the nonlinear relationship between efficiency changes and multiple process parameters, avoiding the underfitting of linear models under complex operating conditions.
[0224] Bessel function Used to describe the periodic or oscillating behavior of energy consumption that may occur during process adjustments, enabling the model to proactively avoid high energy consumption fluctuation ranges.
[0225] Exponential function It is used to amplify the penalty when energy consumption significantly exceeds the benchmark, thereby suppressing decisions with extremely high energy consumption.
[0226] Explanation of the calculation process for the current defect severity scalar:
[0227] In this invention, the current defect severity scalar is used to quantitatively characterize the overall defect state of the optical crystal at the current production stage. It is obtained by unified mapping and fusion calculation of the defect feature vector and serves as one of the key input parameters in the state space of the reinforcement learning decision module 2.
[0228] The defect feature vector is extracted by the defect monitoring module 1 based on multi-source detection data. It includes at least the degree of lattice distortion, microcrack density, non-uniformity of defect spatial distribution, and thermal stress fluctuation characteristics. Each feature reflects the internal structural defects of the crystal, the degree of local damage, and the defect evolution trend.
[0229] In the specific calculation process, the first step is to perform dimensionless processing on each feature component in the defect feature vector. Dimensionless processing employs an interval normalization method based on historical production statistics, ensuring that defect features with different physical dimensions are uniformly mapped to a preset numerical range, thus eliminating the influence of dimensional differences on subsequent calculations. The historical statistical range can be determined by long-term monitoring data from the stable production phase of similar crystals and is dynamically updated during system operation.
[0230] Secondly, the feature sensitivity of each defect feature component after normalization is adjusted. Different types of defects have varying degrees of impact on the optical properties and structural stability of the crystal. Therefore, when calculating the defect severity scalar, a nonlinear mapping mechanism related to the defect evolution risk is introduced, so that high-risk defects have a higher response intensity in the comprehensive calculation. This nonlinear mapping process is constructed based on the statistical distribution characteristics of defect features over time, and is used to amplify defect changes that have a significant impact on subsequent repair and process stability.
[0231] Subsequently, the defect feature components processed by nonlinear mapping are fused to obtain a single comprehensive defect characterization. This fusion process does not use a simple linear superposition method, but instead introduces a continuously differentiable smooth mapping function to aggregate multidimensional defect features. This allows the defect severity scalar to continuously reflect the overall risk level under the superposition of different defect types, thereby avoiding decision oscillations caused by abrupt changes in defect features.
[0232] Finally, the fusion result, after normalization constraints, is output as a scalar of the current defect severity. This defect severity scalar is a non-negative real number, and its value is positively correlated with the overall risk level of the current crystal defects: when the internal defects of the crystal are within a controllable range, the defect severity scalar is in a lower range; when multiple types of defects aggravate simultaneously or show a rapid evolution trend, the defect severity scalar increases significantly. As an important component of the input state of the reinforcement learning model, this scalar directly participates in the calculation of the composite reward function and decision-making strategy, guiding the system to prioritize in-situ repair or process optimization actions, thereby achieving dynamic suppression of crystal defects and stable control of the production process.
[0233] Explanation of the calculation process for the productivity improvement index:
[0234] In this invention, the production efficiency improvement index is used to characterize the degree of efficiency improvement of the current production state relative to the baseline production state. It is obtained by mapping the relative relationship between real-time production efficiency and target efficiency, and serves as one of the important input parameters of the reward function in the reinforcement learning decision module 2.
[0235] In the specific calculation process, the production system first collects current production efficiency data in real time. Production efficiency is the output of qualified optical crystals per unit time, or an equivalent efficiency index that reflects the overall level of production cycle and yield. Its calculation method is consistent within the same production system to ensure comparability between different decision-making cycles.
[0236] Subsequently, the real-time production efficiency is compared with a preset efficiency benchmark. The efficiency benchmark is determined by statistical data from historical stable production phases and is used to characterize the reference efficiency level under conditions of no significant defect interference and process stability. This ratio calculation eliminates absolute differences between different equipment sizes and production batches, allowing efficiency changes to be characterized as relative increases.
[0237] After obtaining the efficiency ratio, an exponential mapping mechanism is introduced to amplify or compress the efficiency change trend. When the real-time production efficiency is higher than the baseline efficiency, the exponential mapping can strengthen the positive contribution of efficiency improvement to the reward function; when the real-time production efficiency is lower than the baseline efficiency, the mapping result tends to converge to avoid excessive interference from short-term efficiency fluctuations on the reinforcement learning strategy. After the above mapping processing, the final output is the production efficiency improvement index.
[0238] The production efficiency improvement index is a dimensionless real number greater than zero, and its value increases monotonically with the relative improvement in production efficiency. This index is introduced into the composite reward function to guide the reinforcement learning model to gradually learn combinations of process parameters that are conducive to improving the overall production rhythm and output capacity, while ensuring the quality constraints of defect repair. This solves the problem in existing technologies where efficiency optimization is easily suppressed by quality constraints and is difficult to improve in a coordinated manner.
[0239] Explanation of the calculation process for the energy amplification factor:
[0240] In this invention, the energy consumption amplification factor is used to quantify the negative impact of energy consumption exceeding the target range during the production process. It is obtained by nonlinear mapping of the deviation between real-time energy consumption data and the target energy consumption benchmark, and is used to form a significant penalty mechanism for high energy consumption decisions in the composite reward function.
[0241] In the specific calculation process, the energy consumption monitoring unit first collects energy consumption data in real time during the current production cycle. The energy consumption data is the comprehensive energy consumption value per unit output or per unit time, which can reflect the overall energy consumption level of the crystal growth furnace, in-situ repair device and auxiliary systems.
[0242] Subsequently, the deviation between the real-time energy consumption data and the preset target energy consumption benchmark value is calculated. The target energy consumption benchmark value is obtained from the statistics of the historical best production conditions that meet quality and efficiency requirements, and is used to limit the reasonable energy consumption level of the system under the condition of balancing economy and stability. By calculating the deviation ratio of real-time energy consumption relative to target energy consumption, the degree of energy consumption risk of the current production state can be characterized.
[0243] After obtaining the energy consumption deviation ratio, an exponential amplification mapping mechanism is introduced to process the energy consumption deviation. When the real-time energy consumption is close to or lower than the target energy consumption benchmark, the energy consumption amplification factor takes a small value close to zero, so that the negative impact of energy consumption on the reward function remains within a controllable range; when the real-time energy consumption is significantly higher than the target energy consumption benchmark, the exponential mapping result increases rapidly, thereby forming a strong penalty term in the reward function, prompting the reinforcement learning model to actively avoid high-energy-consuming process parameter combinations.
[0244] The resulting energy amplification factor is a non-negative real number, and its magnitude is positively correlated with the degree of energy consumption exceeding the limit. This factor is used as the energy consumption penalty term in the composite reward function, and works together with the energy consumption fluctuation characteristics characterized by the Bessel function to enable the reinforcement learning model to gradually converge to a production strategy with stable energy consumption and better economic efficiency during multi-cycle decision-making. This solves the problem that energy consumption control in existing technologies is easily masked by short-term efficiency improvements.
[0245] Through the nonlinear design of the composite reward function, the reinforcement learning model can form a more stable trade-off between defect repair quality, production efficiency and energy consumption control. This effectively solves the problems of target bias, policy oscillation and long-term training instability caused by the overly simple reward function in the existing technology, and improves the reliability and interpretability of the decision-making strategy in complex production scenarios.
[0246] Composite reward function The theoretical range of is (-∞, +∞), where a positive value indicates that the current decision has a positive return under the constraints of comprehensive quality, efficiency, and energy consumption, and a negative value indicates that the decision scheme has a significant deficiency in at least one key indicator; the reinforcement learning model aims to maximize The strategy is updated to meet the target, thereby gradually converging to a control strategy that satisfies the quality threshold and has the best overall performance.
[0247] Overall technical effects of Example 2:
[0248] By introducing the optimized composite reward function, the training convergence speed and policy stability of reinforcement learning decision module 2 under multi-objective constraints are significantly improved. It can achieve synergistic optimization of production efficiency and energy consumption control while ensuring a defect repair success rate of no less than 95%, providing more engineering-practical intelligent decision support for the optical crystal production process.
[0249] VI. Training and Real-time Decision-Making Process of Reinforcement Learning Models;
[0250] 1. Model initialization phase: Model initialization subunit 241 loads the historical training sample set and pre-trains the reinforcement learning network to enable the model to have basic defect-action mapping capabilities.
[0251] 2. Real-time decision-making stage: The real-time decision-making subunit 242 receives the current state vector and outputs in-situ repair actions and process optimization actions based on the pre-trained policy network. While outputting decision instructions, it simultaneously introduces the correction factor sent by the accuracy correction module 6 and the global optimization parameters sent by the federated learning collaboration module 7 to dynamically correct the actions.
[0252] 3. Parameter update stage: The parameter update subunit 243 updates the network parameters online based on the repair effect data and process execution results returned by the data feedback module 5, combined with the correction factor, to achieve continuous iterative optimization of the model.
[0253] VII. Strengthen the core decision-making integration formula for learning;
[0254] To achieve unified decision-making reasoning based on defect characteristics, production status, reward feedback, and cross-domain collaborative information, this embodiment constructs the following integrated decision evaluation function:
[0255] ;
[0256] in, This is a comprehensive decision evaluation value, used to represent the goodness of the reinforcement learning model's decision in the current state, and its value is used to guide action selection; The defect feature complexity scalar is calculated from the defect feature vector; The production state coupling index is obtained by mapping growth rate, temperature distribution, and raw material ratio. The order parameter of the first-kind Bessel function is used to describe the oscillation characteristics of the defect space; This is the intensity coefficient of local repair action, reflecting the influence weight of in-situ repair actions; The current energy consumption value is derived from energy consumption monitoring data; The target energy consumption benchmark value; This is an energy consumption tolerance scale parameter used to adjust the range of influence of energy consumption deviation; This is a real-time production efficiency value, obtained from the efficiency monitoring system. To ensure efficient and stable offsets, and to avoid logarithmic singularity; This is a global collaborative regularization factor, derived from the aggregated parameters of federated learning; This is the prediction error statistic, calculated from historical prediction residuals; The normalized state entropy reflects the uncertainty of the current system state. This is the normalization function; The upper limit of the decision time window integration represents the state evolution interval within a decision cycle.
[0257] The above formula incorporates defect complexity, production status coupling, energy consumption deviation, efficiency changes, and cross-domain collaborative information into a unified integral-fractional structure, thereby achieving a comprehensive evaluation of local repair effects and global process optimization effects. This effectively solves the problems of existing reinforcement learning models that only focus on a single indicator, are prone to decision bias, and lack stability.
[0258] Comprehensive decision evaluation value The theoretical range of is (0, +∞), where A higher value indicates a better overall performance of the current action combination in terms of defect repair quality, production efficiency, and energy consumption control; when When the threshold is lower than the preset decision threshold, the reinforcement learning model will automatically adjust its strategy and prioritize actions with lower risk or better energy efficiency.
[0259] VIII. Summary of the technical effects of Example 2;
[0260] Through the specific implementation of reinforcement learning decision module 2 as described in Example 2, the system can achieve high-precision joint decision-making for defect repair and process optimization in a complex and perturbation-prone optical crystal production environment, effectively suppress prediction accuracy drift, improve the model's generalization ability in small sample and cross-process scenarios, and thus further enhance the stability and engineering applicability of the entire dynamic optimization control system.
[0261] Example 3 is the third embodiment of the present invention.
[0262] I. Explanation of the key technical aspects of Example 3;
[0263] Based on the overall system architecture of Embodiment 1 and the decision-making algorithm of Embodiment 2, this third embodiment further explains the collaborative operation mechanism of the defect monitoring module 1, the in-situ repair execution module 3, the process optimization module 4, and the data feedback module 5. It focuses on clarifying the closed-loop working principle of the execution layer between multimodal defect perception, in-situ repair execution, global process adjustment, and training sample feedback, so as to verify the feasibility and stable control effect of the present invention in the actual optical crystal production process.
[0264] II. Implementation method and working principle of defect monitoring module 1;
[0265] The defect monitoring module 1 is used to acquire defect information inside and on the surface of the crystal in real time during the optical crystal production process, providing high-precision input data for the reinforcement learning decision module 2.
[0266] In this embodiment, the defect monitoring module 1 includes a multimodal sensing unit 11, a noise filtering unit 12, and a feature extraction unit 13.
[0267] The multimodal sensing unit 11 simultaneously acquires spectral data, ultrasonic data, and infrared thermal imaging data throughout the crystal growth process to reflect defect characteristics such as lattice distortion, microcracks, and impurity inclusions. To suppress electromagnetic interference and equipment operating noise present in the industrial environment, the acquired multi-source monitoring data first enters the noise filtering unit 12.
[0268] The noise filtering unit 12 adopts a multi-scale decomposition and reconstruction mechanism based on wavelet packet transform to identify and suppress noise components in different frequency bands, thereby reducing the impact of noise interference on the accuracy of subsequent feature extraction while preserving effective defect information.
[0269] The noise-filtered multi-source monitoring data is input to the feature extraction unit 13. The feature extraction unit 13 extracts the core feature parameters of the defect based on a multimodal feature fusion algorithm and generates a defect feature vector. The defect feature vector includes at least the defect type, defect location, defect size, and defect growth rate, and is used to comprehensively describe the spatial attributes and evolution trend of the defect.
[0270] The defect feature vector is output to the reinforcement learning decision module 2 and the data feedback module 5 at a sampling frequency of not less than a first preset threshold, wherein the first preset threshold is 10Hz, so as to ensure that the decision module can respond to the dynamic changes of the defect in a timely manner.
[0271] III. Implementation method and working principle of in-situ repair execution module 3;
[0272] The in-situ repair execution module 3 is used to perform precise repair on the detected defects based on the in-situ repair actions output by the reinforcement learning decision module 2.
[0273] In this embodiment, the in-situ repair execution module 3 includes a positioning control unit 31, a laser execution unit 32, and a temperature field execution unit 33.
[0274] After the reinforcement learning decision module 2 outputs a decision command containing defect location information, the positioning control unit 31 controls the three-dimensional motion platform according to the defect location information to accurately position the focus of the repair device to the center of the defect. The positioning accuracy of the positioning control unit 31 is not lower than a first positioning accuracy threshold, which is ±5μm, to meet the accuracy requirements for microscale defect repair.
[0275] After positioning is completed, the laser execution unit 32 controls the high-precision pulsed laser to perform pulsed laser repair operation on the defect area according to the laser power and laser action time parameters given in the decision command; at the same time, the temperature field execution unit 33 controls the partition temperature control device to adjust the local temperature field of the repair area according to the local temperature field gradient parameters in the decision command, so as to reduce the risk of thermal stress concentration generated during the repair process.
[0276] During the repair operation, the laser execution unit 32 and the temperature field execution unit 33 simultaneously collect temperature and stress data of the repair area. The temperature and stress data are sent to the data feedback module 5 in real time as repair process data for subsequent repair effect evaluation and model training.
[0277] IV. Implementation method and working principle of process optimization module 4;
[0278] The process optimization module 4 is used to dynamically adjust the global process parameters of optical crystal production based on the process optimization actions output by the reinforcement learning decision module 2, in order to prevent the generation of defects.
[0279] In this embodiment, the process optimization module 4 includes a parameter receiving unit 41, a control signal generating unit 42, an execution adjustment unit 43, and a process status acquisition unit 44.
[0280] The parameter receiving unit 41 receives process optimization actions included in the decision instructions. These actions include at least adjustments to the crystal growth rate, temperature, and raw material ratio. The control signal generation unit 42, based on an adaptive PID compensation algorithm, converts these adjustments into corresponding control signals to adapt to the dynamic response requirements of different production stages.
[0281] The execution adjustment unit 43 drives the crystal growth furnace control system, raw material supply adjustment device, and energy consumption monitoring system to perform corresponding process parameter adjustments based on the generated control signals. After the adjustment is executed, the process status acquisition unit 44 collects growth rate, temperature distribution, raw material ratio, and energy consumption data in real time and sends them as process status data to the data feedback module 5.
[0282] The adjustment accuracy of the raw material supply adjustment device is not lower than the first accuracy threshold, which is ±0.01g; the temperature control accuracy of the crystal growth furnace control system is not lower than the second accuracy threshold, which is ±0.1℃, so as to ensure the stability and repeatability of the process adjustment.
[0283] V. Implementation method and working principle of data feedback module 5;
[0284] The data feedback module 5 is used to manage the repair process data and process status data in a unified manner, and to provide training samples for the continuous iteration of the reinforcement learning model.
[0285] In this embodiment, the data feedback module 5 includes a data receiving unit 51, a data cleaning unit 52, a data storage unit 53, and a sample packaging unit 54.
[0286] The data receiving unit 51 receives repair process data from the in-situ repair execution module 3 and process status data from the process optimization module 4, and performs time synchronization and integration of data from different sources. The data cleaning unit 52 cleans the integrated data based on outlier detection and filtering algorithms to remove invalid data caused by sensor anomalies or transient interference.
[0287] The cleaned data is stored in a distributed database. The sample encapsulation unit 54 extracts data from the distributed database according to a preset period and encapsulates it into a training sample set containing state parameters, action instructions, and reward values. The training sample set is fed back to the reinforcement learning decision module 2 for iterative updates of its model parameters, thereby forming a closed-loop learning mechanism between the execution layer and the decision layer.
[0288] VI. Overall working principle and technical effects of Example 3;
[0289] Through the coordinated operation of the defect monitoring, in-situ repair, process optimization and data feedback module 5 in Example 3, the present invention can realize real-time defect perception, precise in-situ repair and global process prevention and adjustment in the optical crystal production process, and promote the adaptive iteration of the reinforcement learning model through continuous data feedback.
[0290] This closed-loop execution layer mechanism effectively reduces the risk of defects spreading from local to global, improves the success rate of repair operations and the stability of process adjustments, and enables the system to further improve production efficiency and suppress energy consumption fluctuations while ensuring a defect repair success rate of no less than 95%, fully demonstrating the engineering application value of this invention in the field of precision optical crystal manufacturing.
[0291] Example 4 is the fourth embodiment of the present invention.
[0292] I. Explanation of the key technical aspects of Example 4;
[0293] Based on the system structure, execution layer closed loop, and reinforcement learning decision-making mechanism of Examples 1 to 3, this Example 4 further explains the collaborative optimization role of the accuracy correction module 6 and the federated learning collaboration module 7 in the long-term operation process. It focuses on clarifying the prediction bias calibration mechanism caused by implicit dynamic variables and the implementation method of cross-batch, cross-process, and cross-supply chain data participating in decision optimization, so as to improve the stability and generalization ability of the system under complex working conditions and large-scale production conditions.
[0294] II. Composition and working principle of precision correction module 6;
[0295] The accuracy correction module 6 serves as the central hub for real-time calibration of prediction accuracy. It is used to introduce the influence of implicit dynamic variables in the reinforcement learning decision-making process, correct model prediction bias, and avoid accuracy drift during long-term operation.
[0296] In this embodiment, the accuracy correction module 6 includes a variable acquisition unit 61 and a correction factor unit 62, as shown in the figure. Figure 3 The correction factor unit 62 further includes a data preprocessing subunit 621, a weight allocation subunit 622, an anomaly handling subunit 623, and a correction factor calculation subunit 624.
[0297] (a) Data preprocessing subunit 621;
[0298] The data preprocessing subunit 621 is used to preprocess the raw material purity data, equipment vibration data, environmental data, and growth dynamic data from the variable acquisition unit 61. The preprocessing process includes time alignment, detrending, and scale normalization to eliminate differences in sampling frequency, units, and noise levels among different data sources, providing a unified data basis for subsequent weight allocation.
[0299] (ii) Weight allocation subunit 622;
[0300] The weight allocation subunit 622, based on the adaptive mutual information entropy-random forest algorithm, performs correlation analysis and importance assessment on various types of preprocessed dynamic data. By calculating the mutual information entropy values between various dynamic variables and historical prediction errors, and combining the evaluation results of variable contribution by the random forest model, real-time weights are assigned to four types of variables: raw material purity drift, equipment performance drift, environmental perturbation accumulation, and dynamic characteristics of the growth stage.
[0301] The weights are not fixed values, but are dynamically updated according to the production stage and operating status, so that the reinforcement learning model can perceive the changes of key perturbation factors in different stages.
[0302] (iii) Exception handling subunit 623;
[0303] The anomaly handling subunit 623 identifies anomalies in dynamic data based on preset statistical criteria. In this embodiment, the statistical criterion adopted is the 3σ criterion; when a dynamic variable exceeds its historical statistical distribution range, it is determined to be abnormal data. For the identified abnormal data, the anomaly handling subunit 623 temporarily blocks its direct participation in the calculation of the correction factor and temporarily reduces the weight of the corresponding variable to 0.05 to weaken the impact of abnormal disturbances on the decision-making results.
[0304] (iv) Correction factor calculation subunit 624;
[0305] The correction factor calculation subunit 624, based on the weighted dynamic data and the historical training sample set provided by the data feedback module 5, uses an online gradient descent algorithm to calculate the correction factors for the repair parameters and the process parameters. The generated correction factors are used to dynamically correct the in-situ repair parameters and the process optimization parameters, respectively.
[0306] In this embodiment, the values of both the repair parameter correction factor and the process parameter correction factor are limited to [0.9, 1.1] to ensure that the correction magnitude is within a safe and controllable range and to avoid over-adjustment of the execution layer. The correction factor calculation subunit 624 sends the correction factor to the reinforcement learning decision module 2 with a response time not exceeding a preset delay threshold, wherein the preset delay threshold is 5ms to meet real-time control requirements.
[0307] III. Data Acquisition and Cross-Domain Integration Mechanism of Federated Learning Collaboration Module 7;
[0308] Federated Learning Collaboration Module 7 serves as the central hub for cross-domain collaborative optimization, breaking the limitation of data in a single production unit and introducing cross-batch, cross-process, and supply chain data to participate in global optimization.
[0309] Reference Figure 4The data acquisition unit 71 includes a process data acquisition subunit 711 and a supply chain data interface subunit 712.
[0310] The process data acquisition subunit 711 collects raw material particle size data from the raw material pretreatment process, surface roughness data from the processing process, and optical performance data from the testing process, respectively, to reflect the evolution characteristics of crystal quality in different processes.
[0311] The supply chain data interface subunit 712 obtains raw material batch quality data and equipment maintenance record data from the supply chain database and equipment maintenance management system through the API interface, which is used to characterize the long-term impact of raw material stability and equipment health status on the production process.
[0312] The collected data is integrated into cross-domain related data and sent to the collaborative optimization unit 72.
[0313] IV. Collaborative optimization mechanism of Federated Learning Collaborative Module 7;
[0314] Reference Figure 5 The collaborative optimization unit 72 includes a parameter aggregation subunit 721, a reward function adjustment subunit 722, and a cross-domain processing subunit 723.
[0315] (a) Parametric aggregation subunit 721;
[0316] The parameter aggregation subunit 721 processes cross-domain correlated data based on a hierarchical federated learning architecture. It aggregates model parameters from different production batches through horizontal federated learning, correlates feature parameters from different processes through vertical federated learning, and protects local data with a homomorphic encryption mechanism, thereby generating global optimization parameters without leaking the original data.
[0317] The generated global optimization parameters are sent to reinforcement learning decision module 2 to improve the model's generalization ability in different production scenarios.
[0318] (ii) Reward function adjustment subunit 722;
[0319] The reward function adjustment subunit 722 extracts supply chain-related features from cross-domain correlated data, generates a supply chain-related weight vector, and introduces this weight vector into the composite reward function of the reinforcement learning decision module 2 to dynamically adjust the reward function. This approach enables the reinforcement learning model to comprehensively consider the impact of raw material batch quality fluctuations and equipment maintenance status on production stability during the decision-making process.
[0320] ;
[0321] in, The output value of the dynamically adjusted composite reward function. The value of the supply chain association weight vector scalar mapping generated by the reward function adjustment sub-unit based on cross-domain association data is a comprehensive reflection of the raw material batch quality stability, equipment maintenance status and cross-process quality consistency. Its value is a normalized non-negative real number.
[0322] Q is a cross-domain quality confidence index, used to characterize the quality confidence level of the current production batch under the conditions of supply chain and process collaboration. This index is calculated by the federated learning collaboration module without exposing the original data.
[0323] δ is the cross-domain reward modulation coefficient, with a value of 0.1. It is used to limit the influence of cross-domain information on the original reward function structure and avoid destroying the stability of the main reward term.
[0324] Explanation of the dynamic adjustment mechanism:
[0325] In the dynamically adjusted composite reward function, the first three terms correspond to the defect repair quality, production efficiency and energy consumption control targets, respectively, maintaining the same reward function structure as in Example 2; the newly added fourth term serves as a cross-domain collaborative modulation term, used to introduce supply chain and multi-process collaborative information into the reinforcement learning decision-making process.
[0326] When the supply chain association weight vector and cross-domain quality confidence are high, the modulation term has a positive reinforcing effect on the overall reward, guiding the reinforcement learning model to prioritize decision strategies that match high-quality raw material batches and good equipment status. When cross-domain anomalies are identified or quality confidence decreases, the contribution of the modulation term weakens accordingly, thereby reducing the interference of abnormal cross-domain factors on the decision strategy.
[0327] The above methods enable dynamic and adaptive adjustment of the reward function without altering the original main structure of the composite reward function, thereby enhancing the stability and generalization ability of the reinforcement learning model under conditions of multiple batches, multiple processes, and supply chain fluctuations.
[0328] (iii) Cross-domain processing subunit 723;
[0329] The cross-domain processing subunit 723 uses the isolated forest algorithm to identify abnormal features in cross-domain associated data. When an abnormal pattern is detected, the root cause of the anomaly is located through the blockchain traceability node, and a preset alternative optimization scheme is triggered to avoid the abnormal data from having a negative impact on the global model.
[0330] V. Overall working principle and technical effects of Example 4;
[0331] In this embodiment, the optical crystal production dynamic optimization control system based on defect monitoring does not simply use a reinforcement learning model to perform single closed-loop control of defect repair and process parameters. Instead, it constructs a two-layer closed-loop, real-time interactive, and collaboratively enhanced decision optimization architecture around the reinforcement learning decision module 2. This architecture significantly improves the system's stability, adaptability, and global optimization capability in complex production environments through the coupling of the accuracy correction module 6 and the federated learning collaboration module 7 at different time scales and information levels.
[0332] (i) Inner closed loop: The "real-time perception-decision correction" mechanism based on the accuracy correction module 6;
[0333] In the inner control loop of reinforcement learning decision module 2, the system does not directly input dynamic variables such as raw material purity, equipment vibration, and environmental fluctuations as ordinary state variables into the reinforcement learning model. Instead, it independently models and analyzes the above-mentioned implicit dynamic variables in real time through the accuracy correction module 6.
[0334] Specifically, the accuracy correction module 6 first preprocesses the raw material purity data, equipment vibration data, environmental data, and growth dynamic data from the variable acquisition unit 61, and dynamically evaluates the impact of various variables on decision-making errors at the current production stage based on the adaptive mutual information entropy-random forest algorithm, thereby allocating real-time weights. Subsequently, through the anomaly processing subunit 623, abnormal fluctuation data is identified and masked based on preset statistical criteria to avoid the amplification effect of instantaneous anomalies on decision stability. On this basis, the correction factor calculation subunit 624 uses an online gradient descent algorithm, combined with historical training sample sets, to generate repair parameter correction factors and process parameter correction factors, and sends them to the reinforcement learning decision module 2 within a response time not exceeding a preset delay threshold.
[0335] When outputting decision instructions, reinforcement learning decision module 2 simultaneously receives correction factors and uses them as direct corrections to the model inference results. These factors are then applied to the parameter output processes of in-situ repair actions and process optimization actions, thereby compensating for prediction biases caused by raw material fluctuations, equipment state drift, and environmental disturbances within a millisecond timescale.
[0336] Through this inner closed-loop mechanism, the system is equivalent to introducing a "real-time decision calibration channel" for reinforcement learning decision module 2, which effectively suppresses the accuracy drift problem that occurs in the long-term operation of the model and significantly improves the real-time accuracy and robustness of defect repair and process control.
[0337] (ii) Outer closed loop: "Cross-domain collaboration - decision enhancement" mechanism based on Federated Learning Collaboration Module 7;
[0338] In the outer control loop of reinforcement learning decision module 2, the system further introduces federated learning collaboration module 7 to overcome the optimization bottleneck caused by the limited data scale of a single production batch or a single process.
[0339] The federated learning collaboration module 7 collects raw material particle size data, surface roughness data, and optical performance data from the raw material pretreatment, processing, and testing processes through the data acquisition unit 71, and obtains raw material batch quality data and equipment maintenance record data through the supply chain data interface subunit 712, integrating the above multi-source information into cross-domain correlated data. In the collaborative optimization unit 72, the parameter aggregation subunit 721, based on the hierarchical federated learning architecture, aggregates model parameters from different production batches and processes after homomorphic encryption processing of the cross-domain correlated data, generating global optimization parameters. At the same time, the reward function adjustment subunit 722 extracts supply chain features from the cross-domain correlated data, generates a correlation weight vector, and uses it to dynamically adjust the composite reward function in the reinforcement learning decision module 2.
[0340] Unlike traditional federated learning, which is only used for periodic model updates, in this embodiment, the global optimization parameters and the dynamically adjusted reward function are directly injected into the real-time decision-making process of the reinforcement learning decision module 2 as enhanced constraint information in the decision reasoning stage. This allows the decision results of the current batch to simultaneously absorb global optimization experience across batches, processes, and even the supply chain, thereby avoiding the long-term constraints of local optima on system performance.
[0341] (III) The synergistic integration and efficiency enhancement mechanism of the two-layer closed loop;
[0342] In this embodiment, the inner closed loop and the outer closed loop do not operate independently, but form a highly coupled synergistic structure at the reinforcement learning decision module 2.
[0343] On the one hand, the decision results and execution feedback data corrected in real time by the accuracy correction module 6 in the inner closed loop are continuously stored and encapsulated into a training sample set, providing high-quality and highly localized data support for the aggregation of federated learning parameters and adjustment of reward functions in the outer closed loop. On the other hand, the global optimization parameters and cross-domain reward modulation information generated by the outer closed loop also have a reverse effect on the decision reasoning process of the inner closed loop, enabling real-time repair and process optimization decisions to have a long-term global optimal orientation while considering the instantaneous production state.
[0344] Through the deep nesting of this double-layer closed loop in terms of time scale and information level, this system achieves the unity of "real-time optimization" and "global collaboration" as a whole, producing a significant collaborative control efficiency effect. It fundamentally solves the problems of disconnect between repair and optimization and imbalance between local adjustment and global target that are common in the existing optical crystal production process, and significantly improves the system's stable operation capability and large-scale adaptability under complex and variable production conditions.
[0345] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principle of the present invention should also be considered within the scope of protection of the present invention.
Claims
1. A dynamic optimization control system for optical crystal production based on defect monitoring, characterized in that, include: The defect monitoring module is used to collect production data, extract features, and output defect feature vectors. The reinforcement learning decision module, connected to the defect monitoring module, is used to receive the defect feature vector and production status parameters, and output decision instructions including in-situ repair actions and / or process optimization actions through the reinforcement learning model. The in-situ repair execution module, connected to the reinforcement learning decision module, is used to execute the in-situ repair action and collect repair data; The process optimization module, connected to the reinforcement learning decision module, is used to execute the process optimization actions and collect process data; The data feedback module is connected to the in-situ repair execution module and the process optimization module, respectively, and is used to store the repair data and process data and feed them back to the reinforcement learning decision module. An accuracy correction module, connected to both the defect monitoring module and the reinforcement learning decision module, includes: The variable acquisition unit is used to collect data on raw material purity, equipment vibration, environmental data, and growth dynamics to form the acquired data. The correction factor unit, connected to the variable acquisition unit, is used to generate correction factors based on the acquired data and historical data and send them to the reinforcement learning decision module. The federated learning collaboration module, which connects the reinforcement learning decision module and the accuracy correction module respectively, includes: The data acquisition unit is used to acquire cross-domain data from external systems and the supply chain. A collaborative optimization unit, connected to the data acquisition unit, is used to perform privacy aggregation on the cross-domain data through federated learning, generate global optimization parameters, and send them to the reinforcement learning decision module. The reinforcement learning decision module dynamically corrects the decision instructions based on the correction factor and the global optimization parameters.
2. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 1, characterized in that, The reinforcement learning decision module includes: A state space construction unit is used to construct a state space containing defect feature vectors and production system state parameters, wherein the production system state parameters include at least crystal growth rate, furnace temperature distribution, raw material ratio, energy consumption, and production efficiency. An action space construction unit is used to construct an action space that includes in-situ repair actions and process optimization actions. The in-situ repair actions include at least laser power, laser action time, and local temperature field gradient. The process optimization actions include at least growth rate adjustment, temperature adjustment, and raw material ratio adjustment. The reward function construction unit is used to construct a composite reward function, which calculates a comprehensive reward value by weighting the defect repair quality, production efficiency and energy consumption indicators based on preset quality weight coefficient, efficiency weight coefficient and energy consumption weight coefficient. The training and decision-making unit is connected to the state space construction unit, action space construction unit, and reward function construction unit, respectively, and is used to perform the training and real-time decision-making tasks of the reinforcement learning model.
3. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 2, characterized in that, The training and decision-making unit includes: The model initialization subunit is used to load the historical training sample set to pre-train the reinforcement learning model. The real-time decision-making subunit is used to receive real-time state parameters and output the optimal decision instruction based on the pre-trained model. The parameter update subunit is used to iteratively update the network parameters of the reinforcement learning model based on feedback data from the data feedback module and correction factors from the accuracy correction module. When outputting decision instructions, the real-time decision subunit synchronously receives and applies the correction factor sent by the accuracy correction module and the global optimization parameters sent by the federated learning collaboration module to dynamically correct the decision instructions.
4. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 1, characterized in that, The defect monitoring module includes: A multimodal sensing unit is used to collect multi-source monitoring data of the optical crystal production process, wherein the multi-source monitoring data includes at least spectral data, ultrasonic data and infrared thermal imaging data; A noise filtering unit, connected to the multimodal sensing unit, is used to process the multi-source monitoring data using a wavelet packet transform algorithm to filter out electromagnetic interference and equipment noise. The feature extraction unit, connected to the noise filtering unit, is used to extract the core feature parameters of the defect based on the filtered multi-source monitoring data, and generate a defect feature vector containing the defect type, defect location, defect size and defect growth rate. The defect feature vector is output to the reinforcement learning decision module at a sampling frequency not lower than a first preset threshold.
5. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 1, characterized in that, The in-situ repair execution module includes: The positioning control unit is used to control the three-dimensional motion platform to position the focus of the repair device to the center of the defect according to the defect location information in the decision instruction, with a positioning accuracy not lower than the first positioning accuracy threshold. The laser execution unit, connected to the positioning control unit, is used to control a high-precision pulsed laser to perform pulsed laser repair operations according to the laser power and laser action time parameters in the decision command; The temperature field execution unit, connected to the positioning control unit, is used to control the partition temperature control device to adjust the local temperature field of the repair area according to the local temperature field gradient parameters in the decision instruction. The repair device includes the high-precision pulsed laser and the zoned temperature control device. The laser execution unit and the temperature field execution unit synchronously collect temperature data and stress data of the repair area when performing the repair operation, and send the temperature data and stress data as repair process data to the data feedback module.
6. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 5, characterized in that, The process optimization module includes: The parameter receiving unit is used to receive the process optimization actions contained in the decision instruction, wherein the process optimization actions include at least the crystal growth rate adjustment amount, the temperature adjustment amount, and the raw material ratio adjustment amount; A control signal generation unit, connected to the parameter receiving unit, is used to convert the process optimization action into a corresponding control signal based on an adaptive PID compensation algorithm. An execution adjustment unit, connected to the control signal generation unit, is used to drive the crystal growth furnace control system, raw material supply adjustment device, and energy consumption monitoring instrument to perform process parameter adjustments according to the control signal. The process status acquisition unit is connected to the execution adjustment unit and is used to collect the adjusted growth rate, temperature distribution, raw material ratio and energy consumption data in real time, and send them to the data feedback module as process status data. The adjustment accuracy of the raw material supply adjustment device is not lower than the first accuracy threshold, and the temperature control accuracy of the crystal growth furnace control system is not lower than the second accuracy threshold.
7. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 6, characterized in that, The data feedback module includes: The data receiving unit is used to receive and integrate the repair process data and the process status data; A data cleaning unit, connected to the data receiving unit, is used to perform outlier detection and filtering on the integrated data using a data cleaning algorithm. A data storage unit, connected to the data cleaning unit, is used to store the cleaned data in a distributed database; The sample encapsulation unit, connected to the data storage unit, is used to extract data from the distributed database at a preset period and encapsulate it into a training sample set containing state parameters, action instructions and reward values. The training sample set is fed back to the reinforcement learning decision module for iterative updates of its model parameters.
8. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 1, characterized in that, The correction factor unit in the accuracy correction module includes: The data preprocessing subunit is used to preprocess the raw material purity data, equipment vibration data, environmental data, and growth dynamic data from the variable acquisition unit. The weight allocation subunit, connected to the data preprocessing subunit, is used to allocate real-time weights to various types of dynamic data based on the adaptive mutual information entropy-random forest algorithm. An anomaly handling subunit, connected to the weight allocation subunit, is used to identify and block abnormal data based on preset statistical criteria, and adjust the weight of the blocked data accordingly. The correction factor calculation subunit is connected to the weight allocation subunit, the anomaly handling subunit, and the data feedback module, respectively. It is used to calculate the repair parameter correction factor and the process parameter correction factor based on the weighted dynamic data and the historical training sample set provided by the data feedback module through an online gradient descent algorithm. The correction factor calculation subunit sends the generated correction factor to the reinforcement learning decision module with a response time not exceeding a preset delay threshold.
9. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 3, characterized in that, The data acquisition unit in the federated learning collaboration module includes: The process data acquisition subunit is used to acquire raw material particle size data from the raw material pretreatment process, surface roughness data from the processing process, and optical performance data from the testing process. The supply chain data interface subunit is used to obtain raw material batch quality data and equipment maintenance record data from the supply chain database and equipment maintenance management system via API interface; The data acquisition unit integrates the collected raw material particle size data, surface roughness data, optical performance data, raw material batch quality data, and equipment maintenance record data into cross-domain correlated data and sends it to the collaborative optimization unit.
10. The dynamic optimization control system for optical crystal production based on defect monitoring according to claim 9, characterized in that, The collaborative optimization unit in the federated learning collaborative module includes: The parameter aggregation subunit is used to perform homomorphic encryption processing on the cross-domain associated data based on the hierarchical federated learning architecture, and to aggregate model parameters from different production batches and processes to generate global optimization parameters. The reward function adjustment subunit is used to extract supply chain features from the cross-domain associated data and generate an association weight vector, and to dynamically adjust the composite reward function of the reinforcement learning decision module using the association weight vector; The cross-domain processing subunit is used to identify abnormal features in the cross-domain associated data through the isolated forest algorithm, locate the root cause of the anomaly based on the blockchain traceability node, and trigger alternative optimization schemes. Specifically, the parameter aggregation subunit sends the global optimization parameters to the reinforcement learning decision module, and the reward function adjustment subunit sends the dynamically adjusted composite reward function to the reinforcement learning decision module.
Citation Information
Patent Citations
Semiconductor defect detection and process optimization method based on deep learning
CN120107239A
Automatic alignment and integrated control system and method for optical crystal array
CN121541612A