Germanium single crystal growth intelligent optimization control system based on machine learning
Through machine learning, the intelligent optimization control system for germanium single crystal growth is solved, and the problems of multi-source data inconsistency and sensor drift are generated, high-confidence data sets are realized, precise control and dynamic optimization of germanium single crystal growth process is achieved, and production stability and fault warning capabilities are improved.
Patent Information
- Application Number
- CN202510437075.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-11
AI Technical Summary
The prior art is difficult to solve the problems of inconsistent standardization of multi-source data, sensor drift and data loss in germanium single crystal growth control, resulting in uneven data quality, affecting the practical application of intelligent optimization control systems.
Using a machine learning-based intelligent optimization control system for germanium single crystal growth, data dimension conversion and time stamp synchronization is achieved through the protocol conversion module, combined with anomaly detection algorithm and interpolation and repair technology, a high-reliability cleaning data set is generated, and a process-specific database is constructed to realize unified data collection, cleaning and real-time monitoring of data.
It significantly improves the control accuracy and product quality of the germanium single crystal growth process, ensures high consistency and accuracy of data, improves production stability and fault warning capabilities, and realizes accurate monitoring and dynamic optimization of the germanium single crystal straight pull process.
Smart Images

Figure CN120295244A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of crystal growth control, and particularly to an intelligent optimization control system for germanium single crystal growth based on machine learning. Background Art
[0002] With the rapid development of the semiconductor industry and the continuous improvement of the demand for high-performance materials, germanium single crystals have received extensive attention due to their unique optoelectronic properties in the fields of infrared detection, high-speed communication, and quantum devices. Currently, the Czochralski method is the mainstream process for germanium single crystal growth, which requires stable and continuous crystal growth under high temperature, high purity, and strict process parameter conditions. In the actual industrial production process, it is necessary to collect multiple key indicators such as temperature, melt weight, pulling speed, rotation speed, and crystal defects in real time; however, there are significant differences in communication protocols, sensor calibrations, sampling frequencies, etc. among different crystal furnace devices, resulting in the diversity and heterogeneity of data sources. In addition, in harsh environments such as high temperature and strong radiation, sensors are prone to drift, failure, and data loss, which in turn affects the accuracy of real-time monitoring and process control. To meet the requirements of large-scale industrial production of germanium single crystals for high quality and high consistency, it is necessary to build a unified data acquisition interface to achieve dimensional conversion, timestamp synchronization, and centralized indexing of multi-source data, and at the same time introduce a fault detection and warning mechanism to ensure high accuracy and continuity of data in the acquisition stage, providing a solid data foundation for subsequent intelligent optimization control and closed-loop regulation.
[0003] The existing technologies mainly adopt traditional PID regulation and single-sensor monitoring in the control of germanium single crystal growth, which are difficult to solve the problems of inconsistent standardization of multi-source data, sensor drift, and data loss, and thus lead to uneven data quality for machine learning model training, restricting the practical application of the intelligent optimization control system. Specifically, due to differences in communication protocols and sampling frequencies among different crystal furnace devices, it is difficult to uniformly process the raw data output by each sensor in the same coordinate system and unit system; at the same time, sensors are prone to drift in high-temperature environments and lack a unified calibration standard; in addition, data loss and outliers caused by equipment failures or external interferences are not repaired in time, resulting in the interruption of the data chain and affecting the model prediction and closed-loop regulation effects.
[0004] Therefore, the present invention provides an intelligent optimization control system for germanium single crystal growth based on machine learning. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides an intelligent optimization control system for germanium single crystal growth based on machine learning. By adopting a protocol conversion module, the dimension conversion and timestamp synchronization of the original data are realized, and the data integrity is ensured by using a fault detection mark. At the same time, combining the anomaly detection algorithm and the interpolation repair technology, the multi-source original data is converted into a cleaned data set with high credibility, providing reliable data support for machine learning modeling, online prediction and intelligent decision-making, thus significantly improving the control accuracy and product quality of the germanium single crystal growth process and solving the technical problems recorded in the background art.
[0006] To achieve the above object, the present invention is realized through the following technical solutions: An intelligent optimization control system for germanium single crystal growth based on machine learning, including, when the process monitoring is started, the original data M i (t) of the multi-source sensors is uniformly collected and indexed centrally and the fault is detected with the global timestamp T global and the furnace batch number FID k to generate a standardized data stream with a health status mark attached.
[0007] When the standardized data stream is generated, anomaly detection, drift correction and missing data interpolation are performed on it by using the vector state construction and weighted matrix method. By correcting the noise elimination process of the log records, a cleaned data set is obtained.
[0008] Through the fusion storage and multi-dimensional joint index, the structured process parameters and the unstructured CFD simulation files are uniformly imported into the relational and non-relational databases, and combined with the permission and version control, the dynamic update of the data is realized, and then a process-specific database is constructed.
[0009] Relying on real-time data extraction and custom convolution kernel smoothing processing, using the online prediction model to calculate and generate a risk index, and constructing an optimization objective function, automatically generating tuning parameter suggestions and forming a closed-loop feedback, realizing the intelligent adjustment of process parameters and the visual support of decision-making.
[0010] Preferably, a globally unique identifier ID is assigned to each sensor of each crystal furnace i ; after collecting and obtaining the original reading M i (t), each original reading M i (t) is converted into a standardized value S i (t) in a unified coordinate system through data preprocessing. The obtained standardized value S i (t), the corresponding ID i and t are combined to form a new data record, denoted as the intermediate data record.
[0011] Preferably, a unified time tag T is introduced for each intermediate data record global , and each record is bound to the corresponding heat number FID k . If an abnormal or overlapping timestamp span is detected, a mark is triggered and a prompt is sent for subsequent review; if sensor drift or abnormal deviation is detected, a health status field is added to the data record All records that are time-aligned, heat-identified, and marked with health status are aggregated into a standardized data stream
[0012] Preferably, a state vector x i (t) and an expected vector μ i (t) are constructed for each sensor i at time t; an adjustable weight matrix is introduced and the p i th power is used to comprehensively obtain an anomaly metric function A i (t):
[0013] When the anomaly metric function A i (t) exceeds the anomaly threshold Θ i set for this sensor, the reading at this moment is marked as a mutation anomaly and subsequent handling is performed, such as temporary removal or triggering an alarm; if within an approximately stable period, the measurement mean of the sensor deviates from the historical benchmark by more than the preset deviation threshold δ i , it is considered to have drift, and a list of anomaly records including temporarily corrected data is output
[0014] Preferably, if sensor i at a certain time t * is severely abnormal and excluded, an interpolation estimate value is obtained according to the interpolation formula The interpolation estimate value is substituted into the anomaly detection formula A i (t * ) for verification
[0015] If the secondary detection still exceeds the anomaly threshold Θ i , the interpolation point is marked as high-risk data. If it passes the verification, the interpolation estimate value is formally incorporated into the data record to complete the repair of the reading at this moment; all records after interpolation and anomaly correction are merged to form a cleaned data set At the same time, a repair log is output
[0016] Preferably, the cleaned data set Import the keyword fields related to process control into a relational database. Use the heat ID and sampling time as a composite primary key in this database to form a time-series access table structure for storing unstructured objects generated by CFD simulation; when creating an index for the composite key in the actual database, use a distributed search engine;
[0017] Define a similarity coefficient Υ(q,v) for measuring the similarity between a query vector and an attribute vector. Quickly eliminate the records that are closest or farthest from the query vector at the index level through the similarity coefficient γ(q,v);
[0018] Create a composite key index for the heat ID sampling time defect level. Establish spatial index metadata for the 3D field file in the document library or a dedicated retrieval engine and map it to the heat information in the relational database;
[0019] Preferably, when it is detected that a new production heat is completed, or after interpolation and repair of existing data, write the new or corrected content into the database incrementally. Define two modes of hot update and cold update, as well as user roles and access policies; set a short-term buffer for temporary records during the hot update period, and compare with the cleaned data set If a self-conflict is found for the same record, enable the cold update mode to perform a secondary merge;
[0020] Preferably, extract the key process parameters from the cleaned data set in real time, define a trend function and perform smoothing processing using a custom convolution kernel. Through a customized dashboard, synchronously present the curve of the trend function with the real-time raw data S k (t), and overlay the temperature field and flow field distribution maps from unstructured data on the 3D visualization interface;
[0021] Preferably, extract a multi-parameter vector sequence representing key process parameters such as temperature and drawing speed within a time window Γ from the cleaned data set Use weighted integration to construct an input feature vector F(t), input the feature vector F(t) into a machine learning prediction model f (·) that has been trained offline, and obtain the predicted output ML and construct a risk index
[0022] Preferably, after constructing an online optimization objective function and solving the optimization problem to obtain the optimal tuning advice, calculate and obtain the target state; directly push the generated target state P opt (t) to the device execution end, or implement it after prompting the operator to confirm through the interface. The control execution result will be written back to the relational database in the form of a new version record.
[0023] The present invention provides an intelligent optimization control system for germanium single crystal growth based on machine learning, which has the following beneficial effects:
[0024] Centralize the indexing of all data items and set a fault detection flag For real-time identification of abnormal data, which can ensure high consistency and accuracy of the data at the initial stage;
[0025] Utilize the anomaly detection function A i (t) to finely detect the data, and adopt a dynamic drift correction and missing value imputation mechanism to generate a cleaned data set Effectively reduce the data noise and measurement deviation caused by environmental interference; construct a multi-dimensional joint index through the logarithmic distance function Υ(q,v) to achieve efficient retrieval of furnace ID, sampling time, and defect level. The database also adopts version management and permission control to ensure data update and access security, providing real-time and accurate data support for subsequent machine learning and process optimization;
[0026] Through a custom convolution kernel Construct a trend extraction function To achieve smoothing of real-time data and presentation of dynamic trends. Use the feature vector F(t) constructed by the discretized weighted integral method as the input, and the offline trained prediction model f ML (·) to online predict the future process state and calculate the risk index After that, generate online optimization parameter adjustment suggestions, and obtain the optimal parameter adjustment vector ΔP * (t) through solving the optimization problem, so that the parameter adjustment suggestions can be pushed to the device execution end in real time.
[0027] Generally speaking, through unified data acquisition, fine data cleaning, construction of a fusion database, and real-time visualization and intelligent decision-making, the precise monitoring and dynamic optimization of the entire process of the Czochralski process for germanium single crystals are realized, thereby significantly improving production stability, fault warning ability, and overall process optimization level. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic structural diagram of the intelligent optimization control system for germanium single crystal growth based on machine learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0030] Please refer to Figure 1 , the present invention provides an intelligent optimization control system for germanium single crystal growth based on machine learning, including:
[0031] Step 1: When the process monitoring is started, the original data M i (t) of the multi-source sensors is collected uniformly and indexed centrally and fault detected with the global timestamp T global and the furnace batch number FID k to generate a standardized data stream with attached health status markers
[0032] The said Step 1 includes the following contents:
[0033] Step 101: Coordinate normalization and high-level dimension conversion of sensor readings
[0034] Assign a globally unique identifier ID i to each sensor of each crystal furnace, and record its rated range and the original unit i represents the sensor number (such as temperature sensor, melt weight sensor, etc.), and t represents the sampling moment;
[0035] After collecting the original readings M i (t), each original reading M i (t) is converted into a standardized value S i (t) in a unified coordinate system through data preprocessing, and the obtained standardized value S i (t) is formed into a new data record together with the corresponding ID i and t, denoted as the intermediate data record wherein: each original reading M i (t) is converted into a standardized value S i (t) in a unified coordinate system through the following formula:
[0036]
[0037] wherein, α i is the zero-point correction amount, with a value between -0.05 and 0.05, used to offset the inherent offset of the device; β i is the offset threshold, with a value that can be between 0.1 and 0.9, used to control the non-linear gain of the formula when the sensor reading deviates from a certain reference point; γ i is the power factor, with a value that can be between 1 and 3, determining the overall steepness of the conversion curve and can be mapped to a unified range; θ i is the exponential scale coefficient, with a value that can be between 0.01 and 0.2, such that when β iThe transition in the vicinity is smoother;
[0038] Step 102, Data Indexing and Time Synchronization
[0039] Obtain intermediate data records After that, further complete timestamp fusion, heat number marking, and sensor health status synchronization to generate the final standardized data stream Among them:
[0040] Introduce a unified time tag T for each intermediate data record global , bind each record to the corresponding heat number FID k , if an abnormal or overlapping timestamp span is detected, trigger marking and send a prompt for subsequent review;
[0041] If sensor drift or abnormal deviation is detected, add a health status field to the data record When the health status field shows a fault or maintenance required status, then pay key attention to avoid incorrect values from entering the system;
[0042] Summarize all records that are time-aligned, heat-identified, and marked with health status into a standardized data stream Containing the following core information:
[0043]
[0044] At the metadata level of the standardized data stream save the value ranges and unit descriptions of key parameters to provide a one-to-one mapping table for automatic reading and parsing in subsequent steps;
[0045] When in use, globally process the time axis, which can effectively prevent registration errors caused by different sensor sampling delays or unclear heat change points; simultaneously retain the health status, heat number, and standardized readings in a single structured record, avoid confusion caused by decentralized management of multiple files, and use health marking information to assist in screening suspicious records, improving the efficiency and accuracy of downstream processing.
[0046] Step Two, when the standardized data stream is generated, use the vector state construction and weighted matrix method to perform anomaly detection, drift correction, and missing data imputation on it, and obtain the cleaned dataset by modifying the log record noise elimination process
[0047] The said Step Two includes the following content:
[0048] Step 201, Anomaly Detection and Drift Correction
[0049] The output is a list of abnormal records that are temporarily excluded or marked, as well as the corrected intermediate data (hereinafter referred to as the temporary corrected data); let S i (t) represents the standardized measurement value (such as temperature or casting speed) of the i-th sensor at time t after the dimension conversion has been completed and the time is aligned. A state vector x i (t) is constructed for each sensor i at time t so that the measurement value, discrete derivative, and second-order difference can be considered simultaneously in the anomaly metric. In the most common single-channel time series case, it can be defined as:
[0050]
[0051] In the formula: S i (t) is the standardized measurement value (dimension converted and timestamp aligned); is the discrete first-order difference of this measurement value. If the forward difference is used, it can be written as:
[0052]
[0053] is the discrete second-order difference, such as:
[0054]
[0055] corresponds to x i (t), and an expected vector (or reference vector) μ i (t) is defined to characterize the reasonable state of this sensor under normal conditions. Its components can be obtained by moving average, exponential smoothing, or the similar furnace runs estimation described above:
[0056]
[0057] When calculating the degree of deviation, an adjustable weight matrix (different weights or coupling relationships can be assigned to different components) is introduced, and the p i -th power is used to enhance the sensitivity to large deviations, and the anomaly metric function A i (t) is comprehensively obtained:
[0058]
[0059] In the formula: ||·||2 represents the Euclidean norm (L2 norm);
[0060] p i > 1 is the power exponent; W i is a symmetric positive definite weighting matrix, which can be in diagonal form or contain non-diagonal elements. The selection of W i can vary according to the process concerns, <x i (t), μ i(t) represents the dot product of vectors, which is used to measure the matching degree of them in the same dimension; η i > 0 is the exponential correction coefficient;
[0061] When the anomaly metric function A i (t) exceeds the preset anomaly threshold Θ i at a certain moment, it is determined that the sensor readings (and their derivative information) at this moment are abnormal. The specific anomaly threshold Θi can be calibrated in the training set by combining historical operation experience or through the leave-one-out verification method; when the anomaly metric function A i (t) exceeds the anomaly threshold Θ set for this sensor i at a certain moment, the reading at this moment is marked as a mutation anomaly and subsequent disposal is performed, such as temporarily removing it or triggering an alarm;
[0062] For sensors in high-temperature environments, chronic offset phenomena often occur and drift parameter estimation is required. Part of the offset has been corrected in step one, but it may still accumulate further over time; if within an approximately stable period, the measurement mean of the sensor deviates from the historical benchmark by more than the preset deviation threshold δ i , it is considered that there is drift, and the corresponding zero-point correction amount or the power correction amount is updated and the data records during this period are retrospectively corrected; the newly generated correction amount or will continue to be used in the next period until a large offset is detected again. All update operations are recorded in the repair log with a timestamp; finally, an anomaly record list is output, which contains the time t and the detailed information of the anomaly metric function A i (t) exceeding the threshold; the temporarily corrected data after drift correction, excluding or marking mutation points;
[0063] When in use, it can effectively identify outliers caused by equipment failures and environmental interferences, prevent the accumulation of chronic drift, and significantly improve the stability of data: The partial corrections already performed at this stage can reduce the burden of subsequent interpolation and analysis, introduce a dynamic drift correction mechanism, and form a linkage with step one, which can continue to adjust the correction coefficient of the sensor when the actual process state changes significantly, maintaining the consistency of the data dimension.
[0064] Step 202, Imputation of Missing Measurement Points and Output of Repair Results
[0065] Based on the temporarily corrected data, further imputation processing is performed on the blanks still existing inside it (due to reasons such as suspension of production due to failure or being excluded due to serious anomalies), and the final cleaned data set is output together with the repair log, where:
[0066] If at a certain time t *Sensor i with serious anomalies is excluded, and the interpolation estimate needs to be calculated. The following interpolation formula is given to obtain the interpolation estimate.
[0067]
[0068] In the formula: Δ i represents the backtracking distance of the time series interpolation within this heat; represents the average value of temperature or drawing speed at the corresponding moment in another historical heat with a process curve highly similar to the current heat;
[0069] Φ i {·} is a function for prediction based on the nearest normal record on the time axis of this heat, which can be constructed by methods such as linear interpolation, exponential smoothing, or polynomial fitting, and can incorporate drift correction information. Ψ i {·} is the reference value extracted from the similar heat; λ i is the fusion weight, λ i ∈[0,1], and its value can be automatically determined by process experience or model evaluation;
[0070] To prevent the newly interpolated value from causing significant deviations in subsequent evolution, the interpolation estimate needs to be substituted into the anomaly detection formula A i (t * ) for verification;
[0071] If the secondary detection still exceeds the anomaly threshold Θ i , then mark this interpolation point as high-risk data. If it passes the verification, then incorporate the interpolation estimate into the data record formally to complete the repair of the reading at this moment;
[0072] Merge all the records after interpolation and anomaly correction to form a cleaned dataset Retain the marker field of the interpolation source (time interpolation or similar heat interpolation) in it, and at the same time output the repair log, including information such as the anomaly reason, interpolation method, and drift correction record for each moment {t *} processed;
[0073] In use, through the two-stage interpolation method that fuses time backtracking and similar heat data, the spatio-temporal correlation of existing data can be maximally exploited, the accuracy of filling missing points can be improved, and the quadratic anomaly metric verification can ensure the rationality of the results after interpolation, reducing the risk of outliers reappearing; introducing the concept of similar heats into missing interpolation enables repair based on historical heat data with the same or similar recipe parameters even when there is no continuous time data for reference, significantly improving the estimation accuracy. Anomaly metric verification is immediately performed after interpolation to reduce the impact on subsequent analysis and even database storage. The deep cleaning of the standardized data stream is achieved. This enables data records to maintain high integrity and credibility under the multiple influences of external environmental interference, equipment failures, or heat differences.
[0074] Step 3: Through fusion storage and multi-dimensional joint indexing, import structured process parameters and unstructured CFD simulation files into relational and non-relational databases, and combine permission and version control to realize dynamic data update and construct a process-specific database.
[0075] The above Step 3 includes the following content:
[0076] Step 301: Fusion storage and multi-dimensional data indexing
[0077] Import the cleaned data set Keywords related to process control (such as heat ID, sampling interval, drawing speed, temperature) into a relational database (R-DB). In this database, use the heat ID and sampling time as the combined primary key to form a time-series access table structure.
[0078] For unstructured objects such as three-dimensional temperature field files, flow field distribution files, images, and logs generated by CFD simulations, store them through a distributed file system or a document-oriented database (NoSQL); each unstructured file is associated and marked with the corresponding heat ID, and a reference link is established in the relational database for unified retrieval.
[0079] To quickly locate specific intervals or features (such as temperature distribution in a certain time period, samples of a certain defect level) among a large number of data records and file objects, a multi-dimensional joint indexing engine needs to be constructed. In use, not only perform simple combined indexing on heat ID, sampling time, defect level, etc., but also introduce advanced indexing for parameter similarity for quick screening in subsequent machine learning training or searching for similar heats; in an actual database, when establishing an index on the combined key (heat ID, sampling time, defect level), a distributed search engine (such as Elasticsearch) can be used to implement a custom query operator for logarithmic distance functions to quickly screen data records that meet the similarity requirements.
[0080] Assume that each record in the current database has a set of attribute vectors v = [v1, v2,..., v m (such as casting speed, temperature gradient, defect code, etc.), and q = [q1, q2,..., q m represents an external query vector (such as query conditions or similar recipe features). Define a similarity coefficient γ(q, v) for measuring similarity:
[0081]
[0082] where: m is the number of dimensions, corresponding to the comparable attributes in the database;
[0083] ω > 1 is an amplification factor used to amplify the logarithmic difference, usually taking 1.2 ≤ ω ≤ 3.0;
[0084] is the attenuation coefficient, controlling the attenuation rate of the end exponential term;
[0085] Π(q, v) can be any additional metric (such as similarity score of process segments or consistency evaluation of furnace environment. For example, cosine similarity, correlation coefficient, weighted integral or other statistical methods are used to calculate the matching degree between two vectors), which is maintained by the user in the database;
[0086] When the similarity coefficient γ(q, v) is smaller, it means that the query vector q is closer to the attribute vector v on the logarithmic scale. And if its end exponential term is large, a penalty or reward mechanism for strong matching of certain features can be achieved. Through this function, records that are closest or farthest from the query vector can be quickly excluded at the index level, meeting the differential requirements of machine learning modeling and statistical analysis for data distribution;
[0087] Establish a composite key index for heat ID, sampling time, and defect level for regular retrieval;
[0088] Establish spatial index metadata (such as file size, temperature field sampling resolution) for 3D field files in a document library or a dedicated retrieval engine, and map it to the heat information in a relational database (R-DB);
[0089] Pre-set the similarity coefficient γ(·) and the additional metric ∏(·) in the custom operator of the retrieval engine, allowing users or algorithms to enable similarity sorting when initiating a query;
[0090] When in use, with the help of the logarithmic distance function and the multi-database fusion structure, users can quickly locate the time-series data of the target heat and the corresponding simulation files. During anomaly repair, machine learning training, or cross-heat comparative analysis, reference data with similar attributes can be easily obtained using this index model. Since irrelevant or low-similarity data can be filtered out at the index level, the response time in the visualization and algorithm decision-making stages is significantly shortened.
[0091] Step 302, Dynamic Update and Versioning Management
[0092] When a new production furnace batch is detected as completed, or after the interpolation and repair of existing data, the newly added or corrected content needs to be written into the database in an incremental manner. For this purpose, two modes of hot update and cold update are defined:
[0093] Hot update: Collect and perform preliminary retrieval and registration in real time during the production process. Mainly incrementally write to structured time series tables and unstructured object indexes, without large-scale index reconstruction;
[0094] Cold update: After the furnace batch ends or the interpolation and patching are completed, perform batch reconstruction on the index engine (including binning of logarithmic distance functions or updating of spatial metadata), ensuring that the database has complete retrieval capabilities for new data, and marking the update batch ID for all newly added records;
[0095] After each interpolation (or large-scale modification) of the data for the same furnace batch, a new version number Ver is generated n And write a version change record into the database. For large CFD simulation files, if only some grid regions or post-processing intermediate results are updated, an incremental file can be newly created in the document library and have a derivative link with the old version to reduce duplicate storage;
[0096] Establish an access policy corresponding to user roles (operators, process engineers, management, external partners) and data tags (public, restricted, confidential). Both relational tables and unstructured files need to be authenticated based on a unified identity authentication module, and access logs or download records are recorded when necessary.
[0097] Set a short-term buffer for temporary records during the hot update period, and regularly compare with the finally cleaned dataset output in step two If a self-conflict of the same record is found (such as the new interpolation point being inconsistent with the output of the original interpolation algorithm), then enable the cold update mode to perform secondary merging. At the index level, if a key attribute (such as defect level) is updated, synchronous index reconstruction or a minimum-range local update must be triggered to ensure that subsequent retrieval results are not ambiguous;
[0098] When in use, the hot update mechanism ensures write without downtime during the production process, while the cold update mechanism guarantees the integrity of the index after quality repair, enhancing the system's elasticity. The permission grading and access log policy ensure that confidential formulas or patent data will not be illegally leaked; Combining hot update and cold update achieves a balance between low-latency writing during the production process and phased index reconstruction. Through incremental storage and derivative file relationships, the amount of duplicate data is greatly reduced and high traceability is maintained, completing the transition from the cleaned dataset The key leap segment to the dedicated germanium single crystal database significantly enhances the scalability and traceability in the face of large-scale, frequently changing, and multi-security-level data.
[0099] Step 4: Relying on real-time data extraction and custom convolution kernel smoothing processing, use the online prediction model to calculate and generate the risk index and construct the optimization objective function, automatically generate tuning suggestions and form a closed-loop feedback to achieve intelligent adjustment of process parameters and visual support for decision-making;
[0100] The content of Step 4 is as follows:
[0101] Step 401: Real-time data visualization and trend presentation
[0102] Real-time extract the key process parameters in the cleaned dataset and use the advanced trend extraction function to dynamically visualize the data; among them, for key parameters such as temperature and drawing speed, define the trend function k represents the parameter type, and use a custom convolution kernel for smoothing processing. Its expression is:
[0103]
[0104] In the formula: S k (τ) is the measured value of parameter k at time τ in the standardized value; is the attenuation convolution kernel dedicated to parameter k, λ k is the time constant, r k is the exponential adjustment factor, 0.5 ≤ λ k ≤ 5, 1.5 ≤ r k ≤ 3.0; T0 is the selected smoothing window width;
[0105] Through the customized dashboard, present the curve and the real-time raw data S k (t) synchronously, and superimpose the temperature field and flow field distribution maps from unstructured data on the three-dimensional visualization interface;
[0106] When in use, use the custom convolution kernel to smooth the data trend. Compared with the simple moving average method, its non-linear attenuation characteristic can more accurately capture the gradual change signal in the process; the data visualization presentation is intuitive and clear, providing a real-time and global perspective for process monitoring and at the same time providing benchmark information for the subsequent prediction model.
[0107] Step 402: Online prediction and abnormal risk warning
[0108] Use the machine learning prediction model f trained offline ML(·) Online prediction of historical data and current real-time data extracted from relational databases, identifying possible process anomalies in the future in advance, and calculating risk indicators to trigger early warnings, including:
[0109] From the cleaned dataset The multi-parameter vector sequence representing key process parameters such as temperature and casting speed is extracted within the time window Γ in:
[0110] S(t)=[S1(t),S2(t),…,S n (t)] T
[0111] Construct the input feature vector F(t) using weighted integration:
[0112]
[0113] in is the weight function, μ and ν are the time decay parameter and exponential adjustment factor, respectively, and their values can be between 0.5 and 5, and between 1.5 and 3;
[0114] Input the feature vector F(t) into the prediction model f ML (·), get the predicted output (such as predicting future crystal defect rates or process deviations) and defining risk indices
[0115]
[0116] Where: Y safe is the target vector of the safe process state; δ Y is the sensitivity adjustment parameter, the value is greater than 0, and it is recommended to be between 0.01 and 0.2. ||·||1 represents the L1 norm, which measures the absolute deviation between the prediction and the safe state;
[0117] When the risk index Exceeding the preset threshold When abnormal risk warning is triggered automatically;
[0118] To realize the online prediction function, the offline trained prediction model f ML (·) You can choose a recurrent neural network (RNN), a long short-term memory network (LSTM) or a hybrid model (such as an integrated random forest and deep neural network). The training data can be obtained from the cleaned data set output in step 2. From the cleaned dataset Extract the time series data containing key process parameters such as temperature and drawing speed, and construct the feature vector F(t) within the time window as the input of the prediction model. The label Y(t) can be defined as the average or peak value of the key parameters (such as crystal defect rate or process deviation) within a certain future time window (for example, in the future T pred minutes);
[0119] When in use, construct the input features through weighted integration, make full use of the historical time series data information, and improve the robustness of the prediction model. The risk index introduces logarithmic transformation and L1 norm, which not only avoids the limitations of traditional statistics but also can sensitively reflect the degree of abnormal deviation; the prediction output and risk indicators are fed back in real time to ensure that abnormal risks can be warned in advance and provide a basis for control decisions.
[0120] Step 403, Generation of control suggestions and execution of feedback closed-loop
[0121] Use the online optimization algorithm to generate specific process parameter adjustment suggestions, and achieve automatic execution and feedback recording through closed-loop control to ensure that the production process continuously approaches the optimal state, where:
[0122] Construct the online optimization objective function as follows:
[0123]
[0124] In the formula: ΔP(t) is the parameter adjustment vector (for example, the adjustment amount of drawing speed and rotation speed), s opt is the process optimal state vector, T opt is the prediction optimization time window; r opt is the adjustment index, and the value can be between 1.5 and 3, ||·|| F represents the Frobenius norm;
[0125] By solving the optimization problem:
[0126]
[0127] After obtaining the optimal parameter adjustment suggestion, calculate the target state P opt (t):
[0128] P opt (t) = S(t) + ΔP * (t)
[0129] Push the generated target state P opt (t) directly to the device execution end, or implement it after prompting the operator to confirm through the interface. The control execution results (including the states S(t) before and after actual parameter adjustment and the feedback actual process data) will be written back to the relational database constructed in step three in the form of a new version record to achieve closed-loop feedback.
[0130] Meanwhile, all key parameters (such as T opt , r opt , ΔP * (t)) and the risk index are logged for subsequent model updates and policy upgrades;
[0131] When in use, the online optimization objective function combines the Frobenius norm and the integral form, which can not only smoothly handle the changes in time-series data, but also seek the global optimum within the entire prediction window. Through the closed-loop feedback mechanism, it realizes the automatic execution and dynamic correction of the tuning suggestions, further shortens the response cycle, and improves the stability of the production process.
[0132] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.
Claims
1. An intelligent optimization control system for germanium single crystal growth based on machine learning, characterized in that: including When the process monitoring is started, the raw data of multi-source sensors are uniformly collected, and centralized indexing and fault detection are completed with the global timestamp and furnace number, generating a standardized data stream with health status tags attached. After the standardized data stream is generated, anomaly detection, drift correction, and missing data imputation are performed on it using the vector state construction and weighted matrix method, and the cleaned data set is obtained after recording the noise rejection process. Through fusion storage and multi-dimensional joint indexing, structured process parameters and unstructured CFD simulation files are uniformly imported into the database, and combined with permission and version control, a process-specific database is constructed after dynamic data update. Relying on real-time data extraction and custom convolution kernel smoothing processing, a risk index is calculated using an online prediction model, and an optimization objective function is constructed to automatically generate tuning suggestions and form a closed-loop feedback, realizing intelligent adjustment of process parameters and visual support for decision-making.
2. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 1, wherein An identifier is assigned to each sensor of each crystal furnace. After collecting the raw readings, each raw reading is converted into a standardized value in a unified coordinate system through data preprocessing, and the obtained standardized value, corresponding identifier, and timestamp are formed into an intermediate data record together.
3. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 2, wherein A time tag is introduced for each intermediate data record and bound to the corresponding identifier. If an abnormal or overlapping timestamp span is detected, a mark is triggered and a prompt is sent. If sensor drift or abnormal deviation is detected, a health status field is added to the data record, and all records that are time-aligned, added with furnace numbers, and have health status tags are summarized into a standardized data stream.
4. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 1, wherein Construct a state vector and an expected vector for each sensor i at time t, and introduce an adjustable weight matrix and use p i to the power of, and comprehensively obtain the anomaly metric function A i (t); When the abnormal measurement function A i (t) exceeds the abnormal threshold set by this sensor, mark the reading at this moment as a mutation anomaly and perform subsequent handling; if within an approximately stable period, the measurement mean of the sensor deviates from the historical benchmark by more than the preset deviation threshold, it is regarded as having drift, and output an anomaly record list including temporarily corrected data.
5. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 4, wherein If a sensor is severely abnormal and removed at a certain moment, an interpolation estimate value is obtained according to the interpolation formula, and the interpolation estimate value is substituted into the anomaly detection formula again for verification. If the secondary detection still exceeds the anomaly threshold, the interpolation point is marked as high-risk data. If it passes the verification, the interpolation estimate value is formally incorporated into the data record to complete the repair of the reading at that moment. All records after interpolation and anomaly correction are merged to form a cleaned data set and a repair log is output.
6. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 5, wherein The cleaned data set and keyword fields related to process control are imported into a relational database to form a time-series access table structure, and unstructured objects generated by CFD simulation are stored. A composite key index is established for the furnace ID, sampling time, and defect level. Spatial index metadata is established for the three-dimensional field file in the document library or a dedicated retrieval engine and mapped to the furnace information in the relational database.
7. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 6, characterized in that: When it is detected that a new round of production furnace runs is completed, or after the interpolation repair of the existing data, the newly added or corrected content is written into the database in an incremental manner, and two modes of hot update and cold update, as well as user roles and access policies are defined; A short-term buffer is set for the temporary records during the hot update period, and it is compared with the cleaned dataset regularly. If the same self-conflict is found, the cold update mode is enabled to perform secondary merging.
8. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 7, characterized in that: The key process parameters in the cleaned dataset are extracted in real time, a trend function is defined and smoothed using a custom convolution kernel; through a customized dashboard, the curve of the trend function and the real-time raw data are synchronously displayed on the dashboard, and the temperature field and flow field distribution maps from unstructured data are superimposed on the three-dimensional visualization interface.
9. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 8, characterized in that: A time series vector sequence containing multiple process parameters within a time window is extracted from the cleaned dataset, and an input feature vector is constructed using weighted integration; The feature vector is input into the machine learning prediction model trained offline, and the prediction output is obtained and a risk index is constructed. When the risk index exceeds the preset threshold, an abnormal risk warning is automatically triggered.
10. The intelligent optimization control system for germanium single crystal growth based on machine learning according to claim 9, characterized in that: After constructing the online optimization objective function and solving the optimization problem, the optimal tuning parameter suggestions are obtained, and then the target state is calculated and obtained. The generated optimal tuning parameter suggestions are directly pushed to the device execution end, or implemented after being confirmed by the operator through an interface prompt. The control execution result will be written back to the relational database in the form of a new version record.
Citation Information
Cited By
Optical crystal production dynamic optimization control system based on defect monitoring
CN121806513A