Germanium single crystal growth process regulation and control system and method

By establishing a cross-device calibration system and real-time abnormality detection, the problem of inconsistent data quality in germanium single crystal growth experiments is solved, efficient data cleaning and adaptive model optimization are achieved, and the precise regulation of process parameters is improved.

CN120273030APending Publication Date: 2025-07-08KUNMING YUNZHE HIGH TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510446069.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In high-throughput germanium single crystal growth experiments, sensor drift, calibration deviation and data record loss lead to uneven data quality, affecting the accuracy of subsequent analysis and model optimization.

Method used

Establish a consistent calibration system across devices, monitor abnormalities in real time and perform multi-channel redundancy verification, combine hierarchical cleaning and consistency analysis to form a closed-loop adaptive model training and feedback process to ensure data quality.

Benefits of technology

It improves data consistency and accuracy, reduces noise interference, enhances model stability and process optimization accuracy, and improves the management level of high-throughput experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120273030A_ABST
    Figure CN120273030A_ABST
Patent Text Reader

Abstract

The invention discloses a germanium single crystal growth process regulation and control system and method, relates to the technical field of process optimization, and aims to solve the problems of non-uniform measurement of equipment and difficulty in timely recognition of abnormity in multi-furnace germanium single crystal pulling. And a high-throughput data quality control method with cross-device dynamic calibration, real-time anomaly detection, multi-channel redundancy check, graded cleaning and feedback modeling is provided. Calibration and a data template are unified through the first step; 2, monitoring abnormity based on local kernel density and redundancy similarity; step 3, carrying out hierarchical cleaning by adopting outlier integration and batch difference; 4, carrying out model training by utilizing a quality weight and monitoring an error on line, and if high-quality data also deviates, backtracking a calibration or updating strategy, so that the comparability and the reliability of the data are both considered, and the optimization efficiency of a crystal pulling process is remarkably improved; according to the method, multi-equipment measurement and threshold dynamic adjustment requirements are considered, and improvement of single crystal growth quality and data management efficiency is practically assisted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of process optimization, and specifically to a germanium single crystal growth process control system and method. Background Art

[0002] In the field of semiconductor material preparation, especially in the large-scale germanium single crystal growth process, it is often necessary to conduct multiple rounds of crystal pulling experiments for different furnace batches and equipment conditions to explore the effects of parameters such as temperature gradient, pulling speed, and atmosphere on crystal quality. To improve the R & D efficiency, laboratories usually adopt automated crystal pulling equipment and various sensors, and use high-throughput experimental methods to complete batch parameter combination tests in a short time, and record massive information such as temperature, stress, defect count, and crystal growth speed in real time. On the one hand, these data can be used to quickly screen potential process parameter combinations, and on the other hand, they also provide materials for subsequent machine learning or statistical modeling. However, due to the fact that the crystal pulling process involves multiple devices, different furnace batch production lines, and various environmental interferences, the accuracy and consistency of data have become the core factors restricting the R & D process and the effect of process optimization. If the high-throughput experiment is not equipped with a perfect data standardization process and quality control mechanism, it is very easy to cause problems such as the inability to compare the same index between different furnace batches, no redundant verification when the sensor fails, and abnormal data flowing into the database and being regarded as valid samples, which will expand the subsequent modeling error and waste a large amount of human and material costs.

[0003] In the above high-throughput experimental environment, the core technical problems are as follows:

[0004] How to address the problem of uneven data quality caused by factors such as sensor drift, calibration deviation, and data recording loss during the multi-device data acquisition process. Specifically, due to the inherent measurement errors of each sensor on different furnace batches and devices, and occasional sensor failures or signal losses during the automatic acquisition process, if cross-device calibration and format standardization are not achieved through a dynamic calibration reference library at the data source, the collected data will have serious systematic biases and noises. This problem directly affects the accuracy of subsequent data-driven anomaly detection, hierarchical cleaning, and consistency analysis, and causes biases in the model using quality weights for weighted training, and may ultimately lead to distorted conclusions in process parameter optimization. For this reason, a closed-loop feedback optimization method including multi-level technical means such as dynamic calibration, real-time anomaly detection, multi-channel redundant verification, outlier integration, and batch consistency evaluation is proposed to ensure that each piece of data undergoes strict quality control before entering the modeling, so as to provide a reliable basis for the precise control of the germanium single crystal growth process.

[0005] For this reason, the present invention provides a germanium single crystal growth process control system and method. Summary of the Invention

[0006] (1) Technical Problems to be Solved

[0007] In view of the deficiencies of the prior art, the present invention provides a germanium single crystal growth process control system and method, particularly focusing on the pain point that it is difficult to detect sensor drift or signal loss in a timely manner during the automated crystal pulling process and it is easy to introduce system deviations. Through multi-level technical means such as establishing a cross-device consistent calibration system at the data source, real-time monitoring of anomalies and multi-channel redundant verification, and post-event batch hierarchical cleaning and consistency analysis, a closed-loop adaptive model training and feedback process is finally formed, so as to be able to balance experimental efficiency and data quality in high-throughput large-scale production scenarios and provide solid data support for the accurate optimization of germanium single crystal growth parameters.

[0008] (2) Technical solutions

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0010] A method for controlling the germanium single crystal growth process, including: when multiple batches of equipment enter the pre-crystal pulling stage, reading the original values of each sensor, calling a dynamic calibration reference library and a unified acquisition template to perform automatic calibration and format standardization on it, and generating calibrated data embedded with a calibration coefficient θ i to improve the source comparability and data consistency, and generating real-time tracking logs for each device; When the calibrated data

[0011] continuously flows into the real-time monitoring unit, online anomaly recognition is implemented according to the local kernel density Θ (t) and the redundant kernel similarity Π i (t), automatically marking significantly deviated values, and synchronously writing them into the metadata field to quickly isolate the abnormal channel and prevent distorted data from being injected into subsequent links; i When the archived quantity meets the batch analysis threshold, perform hierarchical cleaning and inter-batch consistency inspection on the data with anomaly marks, call the outlier γ

[0012] (t) and the batch difference degree i to eliminate seriously abnormal records, down-weight local suspected values or include them in the to-be-reviewed, and label batches with overall offsets, hierarchically measure data reliability and write back quality labels; When multi-level cleaning is completed and the available data volume is sufficient, extract normal data and the suspected anomalies retained after down-weighting for weighted loss training, strengthen the contribution of high-quality samples through the quality weight ω

[0013] (t), and introduce the deviation metric Δ i (t) to monitor the model stability. Once the deviation of high-quality data exceeds the limit, trigger backtracking for re-calibration or update the cleaning threshold to form a closed-loop iteration. i When the high-quality data deviation exceeds the limit, trigger backtracking for re-calibration or update the cleaning threshold to form a closed-loop iteration.

[0014] Preferably, define the original reading of each device during the calibration phase as S i (t), and the reference curve corresponding to the reference library as R i (t). Quantify the cumulative overall deviation through the reference consistency Ω i (T), and calculate the calibration coefficient θ for each device according to the result of the reference consistency Ω i (T); i ;

[0015] Preferably, write the calibration coefficient θ i into a unified data acquisition template, and perform online correction or compensation on the original sensor readings S i (t), and output the calibrated data after compensation Reserve a consistency verification mechanism in the acquisition template, including naming rule verification, range verification, and timestamp completeness:

[0016] Preferably, introduce the local kernel density Θ i (t) to quantify the aggregation degree of the original reading S i (t) at a certain moment relative to the recent historical observations to determine whether it is an outlier;

[0017] To achieve real-time detection, calculate the local kernel density Θ immediately after obtaining the calibrated data at each new sampling point t i (t), and compare it with the preset density threshold Γ2. If Θ i (t) < Γ2, mark it as abnormal and trigger an alarm;

[0018] Preferably, for key performance indicators, configure multiple sensors of the same type or different types to measure simultaneously

[0019] When real-time detection of a certain channel triggers an abnormality, calculate the redundant kernel similarity Π i (t) among the redundant channels to determine whether it is a single-point failure or an overall abnormality

[0020] When the redundant kernel similarity Π i (t) is lower than the similarity threshold Λ2, write a multi-channel abnormality mark in the metadata;

[0021] When the redundant kernel similarity ∏ i (t) is not lower than the similarity threshold Λ2, it can be preliminarily judged that only a single channel has drifted or failed, and the whole batch is not alarmed for the time being, and the automatic compensation logic can be executed or the data of the backup channel can be called.

[0022] Preferably, define the outlier degree γ i (t) to measure a certain calibrated data The overall deviation from the recent history after archiving, according to the outlier degree γ i (t) and the anomaly flag, the data is divided into three levels:

[0023] Severe anomaly: When the calibrated data is marked as anomalous data, or the outlier degree γ i (t) is lower than the high outlier threshold , it is regarded as a severe anomaly;

[0024] Suspected anomaly: When the calibrated data is not marked as a severe anomaly, but the outlier degree γ i (t) is lower than the secondary outlier threshold , it is classified as a suspected anomaly and listed in the pending review list;

[0025] Normal data: The remaining cases are classified as normal data;

[0026] Finally, for each calibrated data write the cleaned label in the database

[0027] Preferably, taking the heat number or experimental batch as the basic unit, aggregate the data determined to be normal data to the batch level for distribution comparison. Suppose there are a total of K batches, denoted as B k (k = 1,…,K), within each batch B k , the representative value M can be calculated according to the calibrated data k ; Introduce the batch difference degree to measure the overall distribution gap between batch and ,

[0028] When the batch difference degree is less than the threshold Φ3, it is determined that there is a significant difference between batch and batch ;

[0029] If it is found that a certain batch B j has a large difference from most other batches {B m ∣m≠j}, then it is determined that batch B j has an overall deviation problem; If a certain batch is confirmed to have an overall deviation, mark it as an abnormal batch or a suspected abnormal batch;

[0030] Preferably, the training data all come from the normal data marked after cleaning and consistency analysis, or the suspected abnormal records that have been confirmed to be available after manual review. The data includes:

[0031] The original readings S that have completed cross-device calibration *(t); Global or batch-level quality labels, metadata, and feature variables;

[0032] To highlight the role of data quality in model training, a quality weight ω is introduced during the training process i (t);

[0033] During specific model training, the quality weight ω i (t) can be incorporated into the loss function Ψ(Θ). During the process of optimizing the loss function Ψ(Θ), each iteration will update each training sample weighted based on the quality weight ω i (t). The algorithm can adopt gradient descent or other numerical optimization methods. After training is completed, the obtained optimal parameters and the corresponding model structure are saved.

[0034] Preferably, after the model training is completed and put into operation, real-time or batch prediction is performed on the newly collected original readings S i (t) to obtain the prediction result If predicting or simulating a key target, it is compared with the actual observed value y i (t) to form an error metric Δ i (t); If the error metric Δ i (t) exceeds the preset threshold ε4, a warning will be automatically triggered;

[0035] As the online prediction data accumulates, the scheme introduces a joint evaluation index Π4 to measure the influence degree of data with different quality levels on the model performance,

[0036] When it is recognized that the model error metric Δ i (t) exceeds the preset threshold ε4, or it shows that there are large deviations in high-quality data under the evaluation of the joint evaluation index Π4, it will automatically backtrack. When the model runs stably and the prediction accuracy meets the process requirements, its prediction or suggestion can be applied to the actual crystal pulling process scheduling.

[0037] The germanium single crystal growth process control system includes,

[0038] A data calibration module. When multiple furnace equipment enters the pre-crystal pulling stage, it reads the original values of each sensor, calls the dynamic calibration reference library and the unified acquisition template to perform automatic calibration and format standardization on them, and generates calibrated data embedded with the calibration coefficient θ i To achieve source comparability and improve data consistency, a real-time tracking log is generated for each device;

[0039] An anomaly detection module. When the calibrated data continuously flows into the real-time monitoring unit, according to the local kernel density Θ i (t) and the redundant kernel similarity Π i(t) Implement online anomaly recognition, automatically mark significantly deviated values, and synchronously write them into metadata fields to quickly isolate abnormal channels and prevent distorted data from being injected into subsequent processes;

[0040] The cleaning and analysis module, when the archiving volume meets the batch analysis threshold, performs hierarchical cleaning and inter-batch consistency verification on the data with anomaly marks, and calls the outlier γ i (t) And the batch difference degree Eliminate severely abnormal records, downgrade local suspected values or include them in the pending review, and label batches with overall offsets, hierarchically measure data reliability and write back quality labels;

[0041] The model construction module, when multi-level cleaning is completed and the available data volume is sufficient, extracts normal data and retained suspected anomalies with downgraded weights for weighted loss training, through the quality weight ω i (t) Strengthen the contribution of high-quality samples, and introduce the deviation measure Δ i (t) Monitor the model stability. Once the deviation of high-quality data exceeds the limit, trigger backtracking for recalibration or update the cleaning threshold to form a closed-loop iteration.

[0042] (III) Beneficial effects

[0043] The present invention provides a germanium single crystal growth process control system and method, having the following beneficial effects:

[0044] Establish a unified cross-device calibration and data acquisition standard, using the calibration coefficient θ i And the unified data template to ensure that the calibrated data output by each sensor Maintains consistency under the same metrological framework, reducing measurement deviations caused by equipment or furnace differences.

[0045] Adopt real-time anomaly detection and multi-channel redundant verification, through the local kernel density Θ i (t) And the redundant kernel similarity Π i (t) And other technologies to timely isolate abnormal data and single-channel fault data, preventing significantly distorted records from entering subsequent analysis.

[0046] Perform hierarchical cleaning and consistency analysis on the archived data, using the outlier γ i (t) And the batch difference degree And other robust statistical methods to synchronously check from the single data and batch levels, write back labels such as severely abnormal, suspected abnormal, and abnormal batches to the database, enabling high-throughput experiments to accurately filter local anomalies and overall deviations.

[0047] After data cleaning, use the high-quality data set and the quality weight ω i(t)'s suspected abnormal data are jointly incorporated into the training process of machine learning or statistical models to form an adaptive optimization model for germanium single crystal technology. Through online error monitoring and evaluation, dynamic tracking of prediction deviations is achieved. Once the model has a large mismatch with high-quality data, the previous steps can be quickly traced back to correct the calibration coefficient θ i Or re-determine the cleaning threshold to maintain the continuous accuracy of the model.

[0048] In summary, the overall solution has four aspects of collaborative benefits:

[0049] First, the high-standard standardization of source data makes the measurement results across devices comparable, laying a solid data foundation:

[0050] Second, the dual measures of real-time detection and redundant verification significantly reduce noise interference:

[0051] Third, multi-level cleaning and batch-to-batch consistency analysis simultaneously enhance reliability and batch comparability:

[0052] Fourth, the closed-loop model training can adaptively respond to process changes and efficiently assist in parameter tuning.

[0053] Through these interlocking technical features, the solution can provide high-precision decision support for the process optimization of germanium single crystal pulling, further improving the automation and refinement management level in the high-throughput experimental environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic flow diagram of the method for regulating the germanium single crystal growth process of the present invention;

[0055] Figure 2 It is a schematic structural diagram of the system for regulating the germanium single crystal growth process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0057] Please refer to Figure 1 , the present invention provides a method for regulating the germanium single crystal growth process, including,

[0058] Step 1. When multiple furnace equipment enters the pre-pulling crystal stage, read the original values of each sensor, and call the dynamic calibration reference library and unified acquisition template to perform automatic calibration and format standardization on them, generating calibrated data embedded with the calibration coefficient θ i of Achieve source comparability and improve data consistency, and generate real-time tracking logs for each device;

[0059] The first step includes the following contents:

[0060] Step 101, Cross-device dynamic calibration

[0061] Build a dynamic calibration reference library in the experimental center in advance, which includes the ideal response curves, environmental compensation parameters and historical reference data of various types of sensors. The reference library is indexed by the device identifier i, corresponding to the ideal output ranges and compensation models of different measuring devices such as thermometers and drawing speed sensors;

[0062] To achieve dynamic calibration of the device in the actual operating environment, it is necessary to monitor the deviation trend between the sensor and the reference library. Define the original readings of each device during the calibration phase (time window [0, T]) as S i (t), and the reference curve corresponding to the reference library is R i (t). Quantify the overall deviation accumulation through the reference consistency Ω i (T):

[0063]

[0064] Among them, Ω i (T) represents the reference consistency of device i within the time period [0, T]; α > 0 is the attenuation coefficient, which controls the sensitivity to large deviations; β > 1 is the deviation exponent, which non-linearly amplifies or reduces the deviation magnitude; t is the time independent variable, which is used for continuous accumulation throughout the calibration period.

[0065] The higher the value of the reference consistency Ω i (T), the more consistent the sensor output is with the reference library; according to the result of the reference consistency Ω i (T), calculate the calibration coefficient θ i for each device, which is used to uniformly correct the sensor readings of this device during subsequent formal experiments:

[0066] θ i = ln(1 + Ω i (T))

[0067] When the reference consistency Ω i (T) is large (indicating that it is more consistent with the reference), the calibration coefficient θ i also increases accordingly;

[0068] When the deviation between the sensor and the reference is obvious, the value of the reference consistency Ω i (T) becomes smaller, and the calibration coefficient θ iAs it decreases, this logarithmic mapping can avoid being overly sensitive to extremely large or small values and take into account a certain degree of smoothness;

[0069] Record the calibration coefficient θ i and so on into the metadata field bound to the device identifier, which is used to correct the original readings of the same device in real time during the subsequent experimental operation phase. For example, the calibrated data of the sensor output during the formal experiment can be:

[0070]

[0071] Compensation is carried out, and the specific function form can be set according to the actual crystal pulling process requirements. For example, during the formal experiment, the calibrated data output by the sensor can be adjusted and corrected through the following compensation function form:

[0072]

[0073] In the formula: S i (t) is the original signal output by the sensor at time t;

[0074] θ i is the calibration offset of device i, representing the systematic deviation of the sensor;

[0075] α i is the proportionality coefficient of device i, which is used to scale the signal and is usually related to the sensitivity of the device;

[0076] β i is the reference value of device i, which is used to adjust the zero point of the signal to align it with the standard response curve R i (t).

[0077] When in use, by integrating and exponentiating the deviation through the reference consistency Ω i (T), it can more accurately reflect the consistency between the overall range of the device and the reference library, reduce the calibration error between different types or different states of devices, eliminate the dimensional differences between devices, and convert the reference consistency Ω i (T) into the calibration coefficient θ i through ln(1 + Ω i (T)), which provides a smoothing mechanism for extreme values and ensures more robustness during subsequent acquisition.

[0078] Step 102, Unify the data acquisition template

[0079] Write the calibration coefficient θ i into the unified data acquisition template. Each data field includes:

[0080] Device identifier, furnace batch number, calibration coefficient θi , timestamp t, and original reading S i (t) and calibrated data Auxiliary metadata such as unit and range definitions, measurement environment descriptions, etc.; at the same time, clarify the sampling frequency f and its floating range to ensure that data for each furnace can be aligned and compared under high-throughput conditions;

[0081] For the original sensor readings S i (t), perform necessary online corrections or compensations to ensure that the calibrated data recorded in the database is already the original reading comparable across devices. The correction processing logic can call the functions described above:

[0082]

[0083] And write it into a unified template in real time to prevent compatibility issues caused by the lack of original calibration information in the later stage;

[0084] Reserve a consistency check mechanism in the acquisition template, including,

[0085] Naming rule check: When a new data field is written, immediately compare whether the field name meets the naming constraints;

[0086] Range check: Compare the calibrated data with the reference interval in the reference library. If unexplainable values appear, mark them as pending manual review status;

[0087] Timestamp completeness: If there is no valid data at a certain moment, automatically trace back and write the reason for the missing data in the metadata, which is convenient for abnormal detection to call relevant marking information in subsequent steps;

[0088] When in use, apply the calibration coefficient θ i directly to the processing of the original reading S i (t), which ensures that the data stored at the database level is the data aligned across devices, reduces the error accumulation caused by repeated conversions. Through the unified template, all downstream analysis links can directly read the timestamp, device identifier, compensated value, and anomaly mark, reducing information silos and enhancing data traceability; automatically write the calibration coefficient θ in the acquisition template i and perform original reading correction, which can reduce the cumbersome processes of independent processing of different sensors at the manual or script level and ensure data consistency. Establish a dynamic deviation measure and calibration coefficient for multiple devices and the reference library; directly apply the calibration coefficient to the real-time compensation of the original reading, which ensures the comparability and consistency of cross-device data, provides clear alignment identifiers and metadata, and can minimize misjudgments or noise accumulation caused by device differences to the greatest extent.

[0089] Step 2. When the calibrated data When continuously flowing into the real-time monitoring unit, according to the local kernel density Θ i (t) and the redundancy kernel similarity Π i (t), perform online anomaly recognition, automatically mark significant deviation values, and synchronously write them into the metadata field to quickly isolate the abnormal channel and prevent distorted data from being injected into subsequent links;

[0090] The second step includes the following content:

[0091] Step 201, Real-time anomaly detection based on local kernel density

[0092] The input data monitored in this step is directly inherited from the calibrated data defined in the first step Among them:

[0093] Introduce the local kernel density Θ i (t) to quantify the aggregation degree of the original reading S i (t) at a certain moment relative to recent historical observations to determine whether it is an outlier, and define:

[0094]

[0095] In the formula: ω2>0 represents a time window with a fixed length, usually set according to the response speed of the crystal pulling process;

[0096] κ2>0 is the sensitivity coefficient in the kernel function, used to control the suppression intensity of deviation values; γ2>1 is the deviation power, used to amplify the difference degree of the original reading;

[0097] When the calibrated data is close to most data within the past time window ω2, the local kernel density Θ i (t) will also be larger; conversely, if the calibrated data deviates significantly from recent data, the calibration coefficient Θ i (t) will be significantly reduced, indicating that the original reading at this moment has an outlier trend;

[0098] To achieve real-time detection, after obtaining the calibrated data at each new sampling point t, immediately calculate the local kernel density Θ i (t), and compare it with the preset density threshold Γ2. If Θ i (t)<Γ2, mark it as abnormal and trigger an alarm;

[0099] When an alarm is triggered, in addition to making a log record, the flag information of the data, that is, the abnormal data flag, will also be written into the metadata field and the original value will be retained; if continuous monitoring of different batches of experiments is required, independent time windows and thresholds can be set for each batch during initialization and distinguished by device identification; in this way, even if the behavior patterns or data distributions of different furnace runs are different, abnormal discrimination can be independently performed.

[0100] During use, through the calculation of the local kernel density Θ i (t), large deviations can be identified instantaneously during data acquisition and an alarm can be issued, quickly preventing a large amount of bad data from flowing into subsequent links, which is more suitable for high-throughput experimental data situations with multiple modalities or unknown distributions. Further, κ2 and γ2 can be set or automatically learned according to the specific crystal pulling speed and temperature fluctuation range to improve the adaptability of this method.

[0101] Step 202, multi-channel redundancy check

[0102] For key performance indicators (such as crystal defect count or key temperature points), multiple sensors of the same type or different types can be configured to measure simultaneously, denoted as c is the channel number, C is the number of multi-channels, and C≥2;

[0103] When real-time detection of a certain channel triggers an anomaly, that is, Θ i (t)<Γ2, calculate the redundancy kernel similarity Π i (t) between redundant channels to determine whether it is a single-point failure or an overall anomaly.

[0104] ω2>0 represents the length of the real-time detection sliding window, measuring the historical interval for multi-channel cross-comparison; η2>0 is the redundancy sensitivity coefficient; μ2>1 is the power coefficient; ζ2>0 is the time decay coefficient; is the calibrated original reading of the c-th channel of the i-th device (or the i-th furnace run) at time t; in the time interval [t - ω2, t], first define the channel difference matrix function D i (τ), whose elements are:

[0105]

[0106] By applying an exponential kernel function to these differences, integrating over time and attaching an exponential decay weight w(t - τ) =

[0107] exp(-ζ2[t - τ]) to obtain the redundancy kernel similarity Π i (t):

[0108]

[0109] In the formula: Used to measure the measurement differences of different channels at the same time τ; is the kernel similarity between channel pairs. The smaller the difference, the larger the value. The larger the difference, the closer the value is to 0. exp(-ζ2[t-τ]) is the time decay weight. The data closer to the current moment has a higher contribution, and the data farther away has a weaker influence. Z(ω2) is the normalization factor, which is used to ensure that the integral result does not have an order of magnitude imbalance due to the difference in ω2 or ζ2. It can be taken as Or other methods that can make redundant kernel similarity ∏ i (t) a constant maintained within a reasonable range of values;

[0110] When the difference between channels is large or the data close to the current moment deviates significantly from each other, the redundant kernel similarity ∏ i (t) If the temperature drops rapidly, it indicates an abnormality or failure; if the temperature drops rapidly, it indicates an overall abnormality and the experiment needs to be suspended or manual intervention is required;

[0111] When the redundant kernel similarity Π i (t) When the similarity is lower than the threshold Λ2, a multi-channel abnormality mark is written in the metadata;

[0112] When the redundant kernel similarity ∏ i When (t) is not lower than the similarity threshold Λ2 (indicating that most channels match each other), it can be preliminarily determined that only a single channel has drifted or failed, and no alarm will be issued for the entire batch. The automatic compensation logic can be executed or the backup channel data can be called.

[0113] Write the residual comparison results between each channel into the metadata field to form a residual verification report;

[0114] When in use, in a high-throughput environment, it is not uncommon for a single channel to experience temporary drift or noise. Multi-channel redundancy can effectively reduce the impact of such errors on the overall experimental process, avoid single-point measurement failures, and quickly distinguish between single-point channel failures and global anomalies, thereby improving decision-making accuracy.

[0115] Step 3: When the archive volume meets the batch analysis threshold, perform graded cleaning and batch consistency check on the data with abnormal labels, and call the outlier degree γ i (t) Difference between batches Severely abnormal records are eliminated, local suspected values ​​are downgraded or included in the list for review, and batches with overall deviations are marked. Data reliability is measured in layers and quality labels are written back.

[0116] The step three includes the following contents:

[0117] Step 301: Multi-stage cleaning

[0118] Defining outlier degree γ i(t), which is used to measure a certain calibrated data The overall deviation from the recent historical records after archiving, and the specific formula is as follows:

[0119]

[0120] Where: ω3>0, which is the length of the sliding time window and determines how long the historical data in the backtracking interval is; Med i (t, ω3) represents a certain robust central value within [t - ω3, t], such as the multi-point median or the compensated median; κ3>0, which is the sensitivity coefficient, and γ3>1, which is the power amplification coefficient exponent of the deviation;

[0121] If the calibrated data is always relatively close to the robust central value Med i in the near future, the outlier degree γ i (t) is larger; on the contrary, if the recent data shows obvious deviations multiple times within a local range, the outlier degree γ i (t) is significantly reduced, reflecting the enhanced cumulative outlier property; according to the outlier degree γ i (t) and the anomaly mark, the data is divided into three levels:

[0122] Severe anomaly: When the calibrated data is marked as abnormal data, or the outlier degree γ i (t) is lower than the high outlier threshold , it is immediately regarded as a severe anomaly and excluded or specially marked in the subsequent analysis;

[0123] Suspected anomaly: When the calibrated data is not marked as a severe anomaly, but the outlier degree γ i (t) is lower than the secondary outlier threshold , it is classified into the suspected anomaly category and listed in the pending review list, and needs to be reconfirmed according to the experimental records or manual remeasurement;

[0124] Normal data: In other cases, it is classified as normal data and directly incorporated into the subsequent analysis and modeling. The high outlier threshold and the secondary outlier threshold here can be set by small-scale pre-experiments or domain expert experience; it can also be dynamically adjusted during long-term operation to adapt to the changes in different furnace batches or equipment characteristics.

[0125] Finally, for each calibrated data Write the cleaned labels in the database, such as: severely abnormal, suspected abnormal, or normal data, so that subsequent processes (including step 302 or the fourth step of model training) can perform differential processing on data of different quality levels. If a certain piece of data has both a single-channel fault and a suspected abnormality mark, it shall be stated side by side in the metadata for convenient subsequent traceability;

[0126] When in use, multi-level cleaning can distinguish between inevitable abnormalities and possible abnormalities, balance automation efficiency and accuracy in high-throughput experiments, will not roughly delete a large number of suspicious data at one time, use the abnormality marks given in the second step to strengthen the determination of severe abnormalities, and can also take into account the redundant verification results to accurately isolate the data of the faulty channel; and through abnormality grading, the cleaning strategy can be dynamically adjusted according to the needs of the experimental stage and is not limited to a single rule, realizing flexible multi-level thresholds.

[0127] Step 302, Inter-batch consistency analysis

[0128] Taking the heat number or experimental batch as the basic unit, aggregate the data determined to be normal data to the batch level for distribution comparison. Suppose there are a total of K batches, denoted as B k (k = 1, …, K), within each batch B k internally, one or more robust statistics (such as the coordinate median vector method, the spatial median method) or a representative value M that can characterize the overall characteristics of this batch can be calculated according to the calibrated data ; Introduce the batch difference degree k to measure the overall distribution gap between the batch and as follows:

[0129]

[0130] In the formula: is a certain mixed reference center, and the mean value or weighted fusion of the representative values (such as the median vector) of the two batches can be taken; ρ3 > 0 is the batch difference sensitivity coefficient, ν3 > 1 is the power coefficient; ∥·∥ represents the Euclidean distance or other norms suitable for multi-dimensional quantities; the integration interval means to perform statistics within the range of the combined data of the two batches, and can also be refined into multiple dimensions (such as temperature, drawing speed, etc. combined for measurement).

[0131] When the data distributions of the two batches are similar, ∥S * (τ)-Mix(…)∥ is relatively small as a whole, the exponential term is relatively large, and the batch difference degree rises. On the contrary, if the two batches are significantly different, the batch difference degree falls. Thus, the consistency between different batches can be quantitatively measured;

[0132] When the batch difference degree is less than the threshold Φ3, it is determined that the batch has a significant difference from the batch ; there is a significant difference;

[0133] If it is found that a certain batch B j has a large difference from most other batches {B m |m≠j}, then it is determined that batch B j has an overall deviation problem. Possible reasons include: the experimental equipment for this batch was not updated in time θ i resulting in thermometer drift, or changes in raw material purity, etc.;

[0134] If a certain batch is confirmed to have an overall deviation, it is marked as an abnormal batch or a suspected abnormal batch and written into the batch-level label field of the database. In the subsequent model training session, a downweighting or exclusion strategy can be adopted for the data of this batch. If it is confirmed through tracing that it is only an environmental accidental event and has been repaired, it can be noted in the label that it has been rechecked and the data of this batch is retained for subsequent analysis;

[0135] When in use, through the design of the batch difference degree , the distribution proximity between multiple batches can be measured simultaneously, identifying abnormal batches or batches that are very different from other batches. The flexibility at the batch level, mixed reference center and other operations can select different mixing centers according to specific scenarios (for example, performing weighted averaging on multi-dimensional vectors such as temperature, drawing speed, and defect count) and have a highly scalable traceability ability. Based on the calibration and anomaly information recorded in the first and second steps, it is possible to quickly locate sensor, furnace, or raw material problems related to batch deviation, providing a direction for subsequent process improvement. The data is divided into three levels: severely abnormal, suspected abnormal, and normal data, ensuring rapid identification and isolation of extreme anomalies in a large amount of data, while leaving some room for rechecking suspicious data: By means of indicators such as batch difference degree quantify the batch distribution difference and detect overall deviation, providing a traceability basis for possible equipment calibration deviation or process fluctuation.

[0136] Step 4: When multi-level cleaning is completed and the amount of available data is sufficient, extract normal data and downweighted and retained suspected anomalies for weighted loss training. Through the quality weight ω i (t) strengthen the contribution of high-quality samples, and introduce the deviation metric Δ i (t) to monitor the model stability. Once the deviation of high-quality data exceeds the limit, trigger backtracking for recalibration or update the cleaning threshold to form a closed-loop iteration;

[0137] The said step 4 includes the following contents:

[0138] Step 401: Model training driven by high-quality data

[0139] The training data are all from the normal data marked after cleaning and consistency analysis, or the suspected abnormal records confirmed to be available after manual review. The data includes:

[0140] The original readings S * (t) that have completed cross-device calibration; global or batch-level quality labels (such as suspected abnormal batches or abnormal batches); metadata and feature variables (such as temperature gradient, drawing speed, crystal defect count, etc.); for data still in a severely abnormal state, they are excluded and not involved in model training; if there are suspected abnormalities but no substantial problems are judged manually, they can also be included in the training with reduced weights;

[0141] To highlight the role of data quality in model training, a quality weight ω i (t) is introduced during the training process. This weight is based on the outlier degree γ i (t) and batch difference degree and is comprehensively calculated based on data marking information, such as:

[0142]

[0143] In the formula: Q i (t) is the comprehensive quality value, and the outlier degree γ i (t) and batch consistency results and other information are normalized and then obtained through weighted average or multiplication; ξ4>0 is the sensitivity coefficient, δ4>1 is the power amplification coefficient; Γ4 is the basic weight constant, used to ensure that ω i (t) is maintained within the expected numerical range;

[0144] When the comprehensive quality value Q i (t) is small (good data quality), the quality weight ω i (t) remains at a high value; if the comprehensive quality value Q i (t) is large, the quality weight decreases rapidly due to exponential decay;

[0145] During the training of a specific model (whether it is a machine learning model or a statistical model), the quality weight ω i (t) can be incorporated into the loss function Ψ(Θ) (Θ represents the model parameter vector) to form:

[0146]

[0147] In the formula: is the prediction of the target value by the model under the parameters Θ; y i(t) is the true (or measured) target value; e(·,·) is a function used to measure the prediction error (non-variance-based metrics such as absolute deviation, logarithmic difference, etc. can be selected).

[0148] Through this weighted form, data with higher quality has a greater impact on training, while data with poor quality is automatically downweighted, thereby improving the robustness of the model.

[0149] During the process of optimizing the loss function ψ(Θ), each iteration will update each training sample based on the quality weight ω i (t). The algorithm can adopt gradient descent or other numerical optimization methods. After training is completed, the obtained optimal parameters and the corresponding model structure are saved for subsequent online inference or process parameter adjustment.

[0150] When in use, through the loss function with quality weights, the model can prevent the higher noise of suspected abnormal data from affecting the overall fitting while not completely discarding the suspected abnormal data; if the quality of subsequent batch data improves, the quality weight ω i (t) also increases, enabling the model to learn more accurate feature relationships faster; conversely, if there are a large number of suspected abnormal data, the model's dependence on them will automatically weaken during training, excluding labels such as severe anomalies, ensuring that the model is not interfered by extreme errors and differentiating weights for normal data and suspected anomalies, which is more flexible and in line with the actual situation.

[0151] Step 402, Loop Feedback Optimization and Model Application

[0152] After the model is trained and put into use, real-time or batch prediction is performed on the newly collected original readings S i (t) to obtain the prediction result If predictions or simulations are made on key targets (such as the amount of crystal defects, yield), then it can be compared with the actual observed value y i (t) to form an error metric Δ i (t) (non-variance-based metrics such as absolute deviation, logarithmic difference, etc. can still be used);

[0153] If the error metric Δ i (t) exceeds the preset threshold ε4, a warning is automatically triggered, and the model mismatch flag is recorded in the metadata, indicating that there may be data distribution drift or model mismatch;

[0154] As the online prediction data accumulates, the scheme introduces a joint evaluation index Π4 to measure the impact degree of data with different quality levels on the model performance. Exemplary formula:

[0155]

[0156] In the formula: Δ i(t) is the current model prediction deviation, i.e., the error metric;

[0157] w grade (Q i (t)) is a piecewise function divided by data quality level (e.g., assign 1.0 for normal data and 0.5 for suspected anomalies, etc.), to distinguish and analyze the contribution of data with different quality stratifications to the overall performance of the model;

[0158] τ4 > 0 is the sensitivity coefficient, and χ4 > 1 is the deviation power amplification coefficient;

[0159] If the error metric Δ i (t) of high-quality data is still higher than expected, it indicates that the model itself may need to be updated; if only low-quality data produces large deviations, it means that the data collection or anomaly detection process needs to be improved again;

[0160] When it is recognized that the model error metric Δ i (t) exceeds the preset threshold ε4, or shows large deviations in high-quality data under the evaluation of the combined evaluation index Π4, it will automatically backtrack, where:

[0161] Re-examine the sensor calibration coefficient θ i : Check whether there is the latest drift in the cross-device calibration link of the first step;

[0162] Evaluate the anomaly detection and cleaning thresholds of the second and third steps: such as whether ω2 and ω3 need to be updated;

[0163] Collect the data during this period (including model mismatch marks), re-perform the cleaning and consistency analysis of the third step, and then enter step 401 to retrain the model, forming a closed loop;

[0164] If it is confirmed that there are only individual channel or furnace failures, the faulty channels can be isolated separately in the multi-channel redundancy check of the second step without large-scale retraining;

[0165] When the model runs stably and the prediction accuracy meets the process requirements, its predictions or suggestions (such as drawing speed, heating power curve, etc.) can be applied to the actual crystal pulling process scheduling. At this time, the process personnel can be coordinated to adjust the key setting parameters (such as temperature gradient), forming a closed loop of prediction-execution-verification, further shortening the iteration cycle of process optimization and reducing resource consumption.

[0166] When in use, combined with online monitoring and historical cleaning labels, the model can dynamically evaluate its own performance. Once a serious mismatch is found, it will trigger traceability and back-training to ensure that better accuracy is always maintained in high-throughput scenarios; if the feedback indicates that there is a problem with the sensor or data cleaning, you can go back to the first, second or third step and make targeted improvements to the corresponding modules to achieve in-depth collaborative rapid iteration process optimization. The model output can directly act on the process parameter adjustment link, and through continuous prediction-execution-measurement-retraining, the rapid iteration optimization of single crystal growth parameters can be achieved. Integrate the various data quality tags, anomaly detection results, and consistency analysis labels produced by the first three steps into the model training process, and continuously monitor the model prediction error during actual operation. Once a warning situation occurs, it can automatically go back to the previous step to find potential defects in calibration anomaly detection or cleaning links, and then train and update the model again to form an adaptive closed-loop process.

[0167] See also Figure 2 The present invention provides a germanium single crystal growth process control system, comprising:

[0168] Data calibration module, when the multi-furnace equipment enters the pre-crystal pulling stage, reads the original value of each sensor, calls the dynamic calibration reference library and the unified acquisition template to perform automatic calibration and format specification, and generates the embedded calibration coefficient θ i The calibrated data Improve source comparability and data consistency, and generate real-time tracking logs for each device;

[0169] Anomaly detection module, when the calibrated data When it continues to flow into the real-time monitoring unit, according to the local kernel density θ i (t) Similarity with redundant core ∏ i (t) Implement online anomaly identification, automatically mark significant deviation values, and synchronously write them into metadata fields to quickly isolate abnormal channels and prevent distorted data from being injected into subsequent links;

[0170] The cleaning analysis module performs graded cleaning and batch consistency checks on the data with abnormal labels when the archive volume meets the batch analysis threshold, calling the outlier degree γ i (t) Difference between batches Severely abnormal records are eliminated, local suspected values ​​are downgraded or included in the list for review, and batches with overall deviations are marked. Data reliability is measured in layers and quality labels are written back.

[0171] Model building module, when multi-level cleaning is completed and the amount of available data is sufficient, normal data and suspected anomalies retained after downgrading are extracted for weighted loss training, and the quality weight ω i (t) Strengthen the contribution of high-quality samples and introduce the deviation metric Δ i(t) Monitor the stability of the model. Once the deviation of high-quality data exceeds the limit, trigger backtracking for recalibration or update the cleaning threshold to form a closed-loop iteration.

[0172] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for regulating the germanium single crystal growth process, comprising, characterized in that: When multiple batches of equipment enter the pre-pulling stage, read the original sensor values, call the dynamic calibration reference library and perform automatic calibration and template matching to generate calibrated data embedded with calibration coefficients, achieving source comparability; After the calibrated data flows into the real-time monitoring process, detect transient anomalies based on local kernel density and trigger multi-channel redundant comparison, set marks for significantly distorted data and store them in the database, implementing anomaly isolation; When the archived quantity reaches the batch analysis threshold, summarize the data mark information and perform hierarchical cleaning and batch consistency inspection using the outlier degree and batch difference degree. After performing data cleaning, write the data mark information back to the metadata; If the cleaned dataset meets the modeling conditions, include both normal data and suspected anomalies retained with reduced weights in the weighted loss training to generate a process optimization model and backtrack calibration or retrain when the online prediction deviation exceeds the expected value.

2. The method for regulating the germanium single crystal growth process according to claim 1, characterized in that: After obtaining the original readings of each device during the calibration stage and the reference curves corresponding to the reference library, quantify the cumulative overall deviation through the reference consistency, and calculate the calibration coefficient for each device according to the result of the reference consistency; Write the calibration coefficient into a unified data acquisition template, perform online correction or compensation on the original sensor readings collected, and output the calibrated data after compensation; Reserve a consistency verification mechanism in the acquisition template, including naming rule verification, range verification, and timestamp completeness.

3. The method for regulating the germanium single crystal growth process according to claim 2, characterized in that: Introduce local kernel density to quantify the aggregation degree of the original readings at a certain moment relative to the recent historical observations to determine whether it is an outlier; when obtaining the calibrated data at each new sampling point, immediately calculate the local kernel density and compare it with the preset density threshold. If the local kernel density is less than the density threshold, mark it as an anomaly and trigger an alarm.

4. The method for regulating the germanium single crystal growth process according to claim 3, characterized in that: Configure multiple sensors of the same type or different types to measure simultaneously for key performance indicators. If an anomaly is triggered in the real-time detection of a certain channel, construct a redundant kernel similarity for judging whether it is a single-point failure or an overall anomaly among the redundant channels; When the redundant kernel similarity is lower than the similarity threshold, write a multi-channel anomaly mark in the metadata; When the redundant kernel similarity is not lower than the similarity threshold, do not alarm the entire batch and execute the automatic compensation logic or call the data of the backup channel.

5. The method for regulating the germanium single crystal growth process according to claim 1, characterized in that: Define an outlier degree for measuring the overall deviation of a piece of calibrated data relative to the recent historical records after archiving. According to the outlier degree and anomaly marks, divide the data into: Severe anomaly: When the calibrated data is marked as abnormal data or the outlier degree is lower than the high outlier threshold, it is regarded as a severe anomaly; Suspected anomaly: When the calibrated data is not marked as a severe anomaly but the outlier degree is lower than the secondary outlier threshold, classify it as a suspected anomaly and list it in the pending review list; Normal data: The rest of the cases are classified as normal data.

6. The method for regulating germanium single crystal growth process according to claim 5, characterized in that: Taking the furnace number or experimental batch as the basic unit, aggregating the data determined to be normal data to the batch level for distribution comparison. After calculating the representative value according to the calibrated data within each batch, introducing the batch difference degree to measure the overall distribution gap between batches. When the batch difference degree is less than the difference threshold, it is determined that there is a significant difference between the two batches. If it is found that a certain batch shows large differences with most other batches, it is determined that the batch has an overall deviation, and it is marked as an abnormal batch or a suspected abnormal batch.

7. The method for regulating germanium single crystal growth process according to claim 1, characterized in that: The training data all come from the normal data marked after cleaning and consistency analysis, or the suspected abnormal records confirmed to be available after manual review. The data includes: the original readings that have completed cross-device calibration. The quality labels, metadata and characteristic variables at the global or batch level. Introducing quality weights during the training process, and integrating the quality weights into the loss function during model training.

8. The method for regulating germanium single crystal growth process according to claim 1, characterized in that: During the process of optimizing the loss function, each iteration will perform weighted update on each training sample based on the quality weights, and obtain the optimal parameters and the corresponding model structure after training. After completing model training and going online, perform real-time or batch prediction on the newly collected original readings to obtain the prediction results. If the prediction or simulation of the key target is compared with the actual observed value, an error metric is formed.

9. The method for regulating germanium single crystal growth process according to claim 8, characterized in that: If the error metric exceeds the error threshold, a warning will be automatically triggered. Introducing a joint evaluation index to measure the influence degree of data with different quality levels on the model performance. When it is identified that the model error metric exceeds the preset threshold, or it shows that there are also large deviations in high-quality data under the evaluation of the joint evaluation index, automatic backtracking will be performed.

10. A germanium single crystal growth process control system, characterized in that: Including A data calibration module. When multiple furnace equipment enters the pre-pulling stage of crystal growth, read the original sensor values, call the dynamic calibration reference library and perform automatic calibration and template matching to generate calibrated data embedded with calibration coefficients. An anomaly detection module. When the calibrated data flows into the real-time monitoring link, detect transient anomalies based on local kernel density and trigger multi-channel redundant comparison, set marks for significantly distorted data and store them in the database, and implement anomaly isolation. A cleaning and analysis module. When the archived quantity reaches the batch analysis threshold, summarize the data marking information and perform hierarchical cleaning and batch consistency inspection using the outlier degree and batch difference degree. After performing data cleaning, write the data marking information back to the metadata. A model construction module. When the cleaned dataset meets the modeling conditions, incorporate both the normal data and the suspected anomalies retained with reduced weights into the weighted loss training, generate a process optimization model, and backtrack for calibration or retraining when the online prediction deviation exceeds the expectation.

Citation Information

Cited By

  • Dynamic detection method and system for crystal growth

    CN120635093A