Electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis system based on decision tree model
By introducing variable-sensitive interval hit determination and perturbation handling mechanisms into the decision tree model, the prediction jump problem caused by input perturbation is identified and mitigated, thereby improving the stability and accuracy of the model in the electronic-grade anhydrous hydrogen fluoride arsenic removal process. It is suitable for real-time high-precision prediction and auxiliary decision-making in industrial scenarios.
Patent Information
- Application Number
- CN202511432300.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-09
AI Technical Summary
In the task of removing arsenic from anhydrous hydrogen fluoride in electronic grade, the decision tree model, which is based on variable threshold splitting, has the problem of discontinuous prediction results near the variable boundary. This causes slight input disturbances to lead to output jumps, affecting production control judgment.
A variable sensitivity interval hit determination module is introduced. By constructing a core variable sensitivity interval matching table, critical regions prone to prediction jumps are identified. When a sensitive interval is hit, a perturbation input sample sequence is generated, and the perturbation prediction output sequence is analyzed to determine the fluctuation judgment and processing branch path and perform targeted processing to stabilize the model prediction.
It improves the stability and accuracy of the model within the critical range of key variables, enhances the model's adaptability and robustness, prevents large prediction shifts caused by slight input perturbations, and ensures the stability and reliability of the arsenic removal process prediction results.
Smart Images

Figure CN120913679A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of decision tree, and particularly discloses an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model. BACKGROUND
[0002] In the production process of electronic-grade anhydrous hydrogen fluoride, arsenic as a key impurity needs to be strictly controlled, and its residual amount often needs to be as low as ppb or even lower. Since the removal effect of arsenic is jointly affected by multiple process parameters such as raw material quality, oxidant type and molar ratio, reaction temperature, reaction residence time, crystallization condition, separation mode, etc., the variables present a complex relationship of nonlinearity, multi-coupling and high dimension, and it is difficult for traditional linear regression or single factor model to effectively depict the interaction mechanism. Therefore, in order to realize accurate modeling and prediction control of residual arsenic concentration, a decision tree model (such as CART, GBDT or XGBoost) with nonlinear modeling capability and interpretable structure is used as a core data analysis tool.
[0003] The advantage of the decision tree model in industrial process data modeling is that its structure is clear, the output is a tree path of condition-result, which is convenient for direct mapping to the process logic; its training process does not require data normalization, and can adapt to process inputs of multiple scales and heterogeneous types; it supports feature importance analysis, PDP and SHAP explanation, which helps to assist in discovering key factors affecting residual arsenic concentration and their action direction. In addition, the decision tree model performs stably in a small sample environment, and is suitable for the reality that the sample size is limited in the production process of electronic-grade chemicals.
[0004] For example, the Chinese invention patent application with the publication number CN114169537B discloses a federated learning method and system of longitudinal xgboost decision tree, which provides a joint training process and a joint inference process of longitudinal xgboost decision tree. In the joint training process, the split points are calculated, and in the joint inference process, each node is discriminated. The disclosed information in the joint training process is the maximum split value of each participant, without directly leaking the feature information of each participant. The security of the joint inference process relies on a homomorphic encryption scheme.
[0005] The above-mentioned technology at least has the following technical problems:
[0006] In the specific application process of the decision tree model in the electronic grade HF arsenic removal task, due to the essence of the tree model is based on variable threshold splitting, there is a problem of discontinuity of prediction results near the variable boundary. Taking the NaF mole ratio as an example, if the variable is fine-tuned from 5.0 mol to 5.2 mol, the path may be switched to another leaf node, causing the predicted As concentration to jump from 18 ppb to 27 ppb, although the input change is very slight. The problem of output jump caused by input disturbance will directly interfere with the production control judgment, especially when the predicted value is close to the qualified critical value, it is easy to cause misjudgment or omission. SUMMARY
[0007] In order to solve the above technical problems existing in the prior art, the embodiment of the present application provides an electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model. The technical scheme is as follows:
[0008] The variable sensitive interval hit determination module is used for collecting an electronic grade anhydrous hydrogen fluoride arsenic removal data variable set, forming an arsenic removal data input sample, performing variable sensitive interval hit determination, and outputting an electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result through the currently executed decision tree model when the sensitive interval is not hit.
[0009] The fluctuation judgment processing branch path determination module is used for, when the sensitive interval is hit, counting the number of hit variables, generating a disturbance input sample sequence, inputting the sequence into the currently executed decision tree model, obtaining a disturbance prediction output sequence, analyzing a prediction fluctuation score, determining a fluctuation judgment processing branch path, and executing targeted processing of the disturbance fluctuation.
[0010] The disturbance processing effect analysis module is used for, after the targeted processing of the disturbance fluctuation is executed, performing disturbance processing effect analysis, and outputting an electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result after effective disturbance processing.
[0011] The technical scheme provided by the embodiment of the present application has at least the following beneficial effects:
[0012] 1. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model provided by the present application can identify and alleviate the prediction jump problem caused by input disturbance in the arsenic removal process through the integration of sensitive region identification, disturbance sample construction, prediction fluctuation score, model adaptive adjustment and prediction smoothing mechanism, and improve the stability and precision of the model in the critical interval of the key variable. The system has the ability to realize online adjustment without retraining the main model, has strong adaptability and high calculation efficiency, is suitable for deploying in industrial scenes to perform real-time prediction and auxiliary decision of the arsenic removal effect with high precision and high robustness, and enhances the data support and intelligent level of the arsenic removal process control.
[0013] 2、The present application can identify whether the input sample is in the critical region prone to prediction jump in the model structure before model prediction by performing variable sensitive interval hit determination. Once the variable hits the sensitive interval, the system will trigger the disturbance analysis and stability detection process, avoiding the output discontinuity caused by the decision tree model splitting structure, improving the reliability and robustness of the model near the core variable fluctuation, preventing large prediction deviation caused by slight input disturbance, enhancing the fault tolerance and engineering usability of the decision tree model in the actual process control scene, thereby ensuring the stability and reliability of the arsenic removal process prediction result.
[0014] 3、The present application can further identify the fluctuation source type in the unstable region of model prediction by performing disturbance uncertainty type analysis, that is, determine whether the current prediction uncertainty belongs to input level noise disturbance or model structure level knowledge blind area, and take different response strategies according to different types. For the uncertainty caused by input noise, the prediction robustness is enhanced by smoothing fusion; for the uncertainty caused by model blind area, structure fine tuning is triggered preferentially. Avoiding the one-size-fits-all processing mode, improving the adaptability, sensitivity and misjudgment tolerance of the decision tree model, and providing more intelligent and stable prediction guarantee for electronic grade anhydrous hydrogen fluoride arsenic removal process. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 is the structure diagram of the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on the decision tree model provided by the embodiment of the present application.
[0017] Figure 2 is the disturbance processing effect analysis flowchart involved in the embodiment.
[0018] Figure 3 is the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis whole process schematic diagram based on the decision tree model involved in the embodiment of the present application.
[0019] Figure 4 is the model architecture diagram of the decision tree model involved in the embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the present application will be described below with reference to the drawings.
[0021] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.
[0022] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.
[0023] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.
[0024] To make the technical problem to be solved, the technical solution and advantages of the present invention clearer, a detailed description will be given below with reference to the accompanying drawings for the purpose of data analysis to predict arsenic content.
[0025] like Figure 1 The diagram shown is a data analysis system structure for arsenic removal using an electronic-grade anhydrous hydrogen fluoride based on a decision tree model. It includes: a variable sensitivity interval hit determination module, a fluctuation judgment and processing branch path determination module, and a disturbance processing effect analysis module.
[0026] like Figure 3As shown, the embodiment of the present application relates to a full-process schematic diagram of electronic grade anhydrous hydrogen fluoride arsenic removal data analysis based on a decision tree model. First, an arsenic removal data input sample is formed and a variable sensitive interval hit determination is performed. If the determination result is that the sensitive interval is not hit, the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result is directly output. If the determination result is that the sensitive interval is hit, a disturbance input sample sequence is generated and further analysis and prediction of a fluctuation score are performed. When the prediction fluctuation score is less than a prediction fluctuation score threshold value, the arsenic removal data input sample is directly taken as an input to output the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result. When the prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold value, the system determines the uncertainty type according to whether the prediction result range is less than a prediction result range threshold value. If the prediction result range is less than the prediction result range threshold value, the input noise type uncertainty is recorded and output smoothing is performed. If the prediction result range is greater than or equal to the prediction result range threshold value, the model structure type uncertainty is recorded and model structure lightweight adjustment is performed. Subsequently, a disturbance processing effect analysis link is entered. If the result is effective disturbance processing in the disturbance processing effect analysis, the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result is output. Otherwise, a warning prompt is output.
[0027] The variable sensitive interval hit determination module is used for collecting an electronic grade anhydrous hydrogen fluoride arsenic removal data variable set, forming an arsenic removal data input sample, performing variable sensitive interval hit determination, and outputting an electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result through a current decision tree model when the sensitive interval is not hit.
[0028] Further, the variable sensitive interval hit determination is performed. The specific determination process is as follows:
[0029] The core variable data of the electronic grade anhydrous hydrogen fluoride arsenic removal process is collected at a preset collection period to form an arsenic removal data input sample.
[0030] In the electronic grade anhydrous hydrogen fluoride arsenic removal process, accurate control of residual arsenic concentration is the key to guarantee product purity and meet the needs of downstream high-end processes. In order to realize intelligent modeling and stable operation monitoring of the arsenic removal process, the system needs to collect and analyze multiple core variables with engineering quantifiable characteristics in real time. These variables directly affect the oxidation reaction efficiency, arsenic species migration behavior and final residual concentration. In the embodiment, they specifically include: NaF molar ratio, oxidant addition rate, reaction temperature, average residence time, HF raw liquid purity and initial arsenic concentration.
[0031] Among them, sodium fluoride is used as an auxiliary precipitant in the arsenic removal process, and the ratio of its dosage to the molar amount of HF is defined as the NaF molar ratio. The NaF molar ratio will significantly affect the stability of the arsenic fluoride complex generated in the system and the precipitation efficiency. The molar ratio control range is generally between 4.5 and 6.5 mol / mol, and a slight change can cause process behavior mutation, which is one of the high-sensitive variables in modeling.
[0032] The oxidation state conversion efficiency of arsenic depends largely on the dosage rate of the oxidizing agent (such as or ), the oxidizing agent dosage rate (mL / min) controls the reaction rate and the heat release rate, which plays a decisive role in the generation of arsenic and subsequent separation process. Its dosage rate is generally controlled in the range of 0.5 to 2.0 mL / min, and too fast or too slow may cause nonlinear reaction behavior.
[0033] Reaction temperature (℃) is one of the core parameters for controlling the activity of oxidation reaction and the stability of arsenic fluoride complex. Too high temperature may cause side reactions or increase the solubility of arsenic species, affecting the arsenic removal efficiency; while too low temperature reduces the reaction rate. In industrial operation, this variable is usually adjusted in the range of 40 to 80℃, and is regulated by the dynamic feedback of heating / cooling system.
[0034] The average residence time (s) of the reaction solution in the main reaction section directly determines the completion of arsenic oxidation and precipitation reaction, and is an important variable for characterizing the kinetics of the reactor. The average residence time is affected by flow rate, liquid level and equipment structure, and is usually controlled in the range of 60 to 300 seconds. Together with temperature and oxidizing agent, it constitutes the three factors of reaction dynamics.
[0035] The starting purity of anhydrous HF has a substantial impact on the migration path of arsenic and the behavior of the subsequent generated precipitate phase. If the water or impurity content in HF is not stable, it will affect the arsenic speciation distribution and reaction path. The purity of HF stock solution (%) is obtained by raw material detection, which is not usually collected in real time, but as a static characteristic variable input into the modeling system.
[0036] The initial arsenic concentration of the stock solution before entering the reaction section determines the load level of the system, which is an important benchmark input for model prediction. The initial arsenic concentration (ppb) may be caused by batch differences, equipment residues and other factors, and is usually measured by online ICP-MS or laboratory samples, with high accuracy and a range of generally 20 to 150 ppb.
[0037] It should be noted that the above variables are only the core variables set in this embodiment, and in specific embodiments, specific settings can be made according to specific circumstances, and this embodiment does not limit.
[0038] The extraction system has a built-in matching table of sensitive intervals of each core variable.
[0039] In the process of electronic grade anhydrous hydrogen fluoride arsenic removal data analysis based on decision tree model, due to the decision tree model essentially relies on variable split threshold for structure division, small disturbance of part of input variables is easy to cause the jump of prediction path, thus leading to the mutation or instability of model output results. In order to identify such high risk input situation in advance in the process of system running, guarantee the prediction stability and result reliability, the system designs the core variable sensitive interval matching table as the pre trigger mechanism of disturbance analysis and judgment process.
[0040] The core variable sensitive interval matching table is a kind of pre defined structured data table, which is used to record the variable split threshold interval which is highly related to model structure and sensitive to prediction behavior. The construction of the table usually comes from the following aspects:
[0041] Model structure analysis, by analyzing the structure of the trained decision tree model, the key split points and their statistical frequencies of all split variables are extracted, and the areas which are easy to trigger path change are identified.
[0042] Prediction fluctuation backtracking analysis, in the history sample prediction record, identify which input value interval exists significant prediction fluctuation or model switching behavior.
[0043] Business risk critical value superposition, combined with the actual control critical point of key variables in the process (such as NaF mole ratio close to the process alarm value), further set the safety buffer zone.
[0044] Finally, the core variable sensitive interval matching table is organized in the form of key value pair or structured dictionary, each record includes variable name, sensitive interval range (such as [4.8, 5.3] mol) and other fields.
[0045] Based on the core variable sensitive interval matching table, the arsenic removal data input sample is judged. If no core variable falls into its corresponding core variable sensitive interval in the arsenic removal data input sample, the variable sensitive interval hit judgment result is recorded as a miss.
[0046] If no core variable falls into its corresponding core variable sensitive interval in the arsenic removal data input sample, it means that the key operating parameters of the current electronic grade anhydrous hydrogen fluoride arsenic removal process (such as NaF mole ratio, oxidant addition rate, reaction temperature, etc.) are in the relatively stable and safe normal operating range. At this time, the variable sensitive interval hit judgment result is a miss, which means that the process running state does not appear abnormal variable disturbance which may cause the prediction result to fluctuate sharply or the model to be distorted. The system can directly process the arsenic removal data input sample through the currently executed decision tree model, output stable and reliable arsenic removal data analysis results, without starting the additional disturbance processing process.
[0047] In specific embodiments, when the sensitive interval is not hit, the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output by the currently executed decision tree model, specifically: taking the arsenic removal data input sample as input, the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output by the currently executed decision tree model, the above-mentioned currently executed decision tree model refers to the decision tree model version actually used for prediction operation on the input data sample at the current time and under the current system configuration, which has the following characteristics: having been trained, deployed online and possibly having been lightly fine-tuned.
[0048] As Figure 4 shown is a model architecture diagram of the decision tree model involved in the embodiments of the application. It includes a data input stage, a sub-tree structure layer, a sub-tree fusion module and a final output node. Among them, the data input stage takes the arsenic removal data input sample as the input of the decision tree model, that is, the arsenic removal data input sample at the top of the figure represents the electronic-grade HF arsenic removal process core variable data input collected by the system. These variables include but are not limited to NaF molar ratio, oxidant addition rate, reaction temperature, etc., all of which are structured numerical data with clear dimensions, forming a complete input sample vector. The sub-tree structure layer includes main model sub-tree structure 1, main model sub-tree structure 2 and main model sub-tree structure 3. The three sub-tree structures in the middle part respectively represent three typical sub-trees in the currently executed decision tree integrated model (such as XGBoost). They constitute the structural basis of the main model, and each tree is modeled based on different training samples or feature subsets, with independent splitting paths and prediction behaviors. This multi-sub-tree integrated structure helps to enhance the generalization ability and local stability of the model. The sub-tree fusion module is used to fuse the output information of the sub-tree structure layer and output a unified prediction value. The fusion methods include but are not limited to weighted average, model confidence weighting, disturbance sample weighting, etc. In the presence of input disturbance or hitting the variable boundary region, the system can dynamically adjust the weights of each sub-tree to improve the stability and accuracy of the prediction. The final output node outputs the residual arsenic concentration prediction result, which directly serves the downstream quality evaluation, parameter optimization suggestion or process stability control in specific embodiments.
[0049] It should be noted that the three sub-tree structures in the figure are only an example to illustrate the core mechanism of the application. In actual application, the number of sub-trees can be flexibly adjusted according to specific engineering requirements, data complexity or computing resource configuration, and the application does not limit this.
[0050] In this embodiment, to achieve high-precision prediction of residual arsenic content in the process of removing arsenic from electronic-grade anhydrous hydrogen fluoride, a supervised learning model is constructed based on historical process operation data, and an interpretable boosting decision tree structure is used as the main model (such as the regression tree ensemble implemented by XGBoost or LightGBM). The input data used for model training includes NaF molar ratio, oxidant addition rate, reaction temperature, average residence time, HF stock solution purity, and initial arsenic concentration, a total of six key process variables. These variables are numerical, quantifiable continuous features. The system first standardizes the original batch data, including missing value completion, outlier removal, and normalization conversion, to ensure the stability and availability of the training data. After the feature preparation is completed, the training set sample pair (X, y) is constructed, where X is a six-dimensional input vector and y is the corresponding residual arsenic concentration value. During the training process, the mean square error is used as the loss function, and the gradient boosting strategy is used to generate weak regression trees iteratively, and the overall prediction performance is improved through ensemble.
[0051] To avoid model overfitting and enhance generalization ability, a cross-validation mechanism is introduced during the training process, and multiple key hyperparameters (such as maximum tree depth, minimum leaf node sample size, minimum split gain, regularization weight, etc.) are optimized through grid search. After training is completed, the system solidifies the model and deploys it in the online prediction process as the current execution decision tree model, supporting subsequent real-time disturbance analysis and prediction stability evaluation processes.
[0052] If a core variable in the arsenic removal data falls within its corresponding core variable sensitive interval, the variable sensitive interval hit determination result is recorded as a hit sensitive interval, and the core variable is recorded as a hit variable.
[0053] The fluctuation judgment processing branch path determination module is used to count the number of hit variables when the sensitive interval is hit, generate a disturbance input sample sequence, input it into the current execution decision tree model, obtain a disturbance prediction output sequence, analyze the prediction fluctuation score, determine the fluctuation judgment processing branch path, and perform targeted processing of disturbance fluctuation.
[0054] If a core variable in the arsenic removal data falls into its corresponding core variable sensitive interval, it indicates that at least one key operating parameter (such as NaF molar ratio, oxidant dosage rate, reaction temperature, etc.) in the current electronic grade anhydrous hydrogen fluoride arsenic removal process is in a critical range that may cause model prediction uncertainty. At this time, the variable sensitive interval hit judgment result is a hit sensitive interval, and the core variable that falls into the sensitive interval is marked as a hit variable. This situation means that the current parameter state may cause the decision tree model to have problems such as output jump and increased dispersion in predicting residual arsenic content, so the system needs to further count the number of hit variables, generate a perturbation input sample sequence, determine the subsequent processing path by analyzing the prediction fluctuation score, and avoid the interference of model prediction distortion on process judgment.
[0055] Further, the number of hit variables is counted, and a perturbation input sample sequence is generated therefrom, and the specific process is as follows:
[0056] The number of hit variables is counted, and if the number of hit variables is 1, a perturbation value is constructed based on the hit variable and hit variable data to generate a perturbation input sample sequence of the hit variable.
[0057] In specific embodiments, when only one input variable hits the sensitive interval (for example, the NaF molar ratio falls into [4.8, 5.3] mol), the target of perturbation analysis at this time is to investigate the influence of small changes of the variable in its neighborhood on the model prediction result. Since the variable is single, there is no need to construct a multi-dimensional space, and there is no interaction between variables, so the central perturbation method (such as ±0.1, ±0.2, etc.) can be used to efficiently construct the perturbation input sequence. The advantage of this method is that the operation is simple, the calculation cost is low, and at the same time, it can accurately depict the prediction curve fluctuation trend in the local area of the current variable, directly reflecting the path jump behavior of the decision tree model near the split point.
[0058] If the number of hit variables is greater than 1, the Latin hypercube sampling method is used. For each hit variable, based on its corresponding sensitive interval, it is divided into K equal-width subintervals, and then a random selection is made in each subinterval to form a perturbation sample set of the hit variable. The perturbation sample sets of the hit variables are combined, shuffled, and matched to generate a linked perturbation sample sequence, which is recorded as the perturbation input sample sequence.
[0059] It should be noted that K represents the number of subintervals into which the sensitive interval is divided. In specific embodiments, the specific value of K should be dynamically set according to the actual engineering scene. When the sensitive interval is large, a larger K value can be selected to refine the perturbation; if the sensitive interval itself is narrow, the K value should not be too large, so as not to make the subintervals too narrow and cause no significant difference between the perturbation samples.
[0060] When multiple input variables (such as oxidant dosage rate, reaction temperature, average residence time, etc.) hit the sensitive interval at the same time, there is a synergistic or coupling relationship between the variables, which cannot be simply perturbed one by one. At this time, if the Cartesian product enumeration is still used, the number of samples will increase exponentially, which is difficult to meet the timeliness of online system response; if random sampling is used, the sample uniformity is insufficient and the coverage range is unstable. Therefore, the system selects Latin hypercube sampling as the perturbation sample generation method.
[0061] The Latin hypercube sampling method ensures that the perturbation interval of each variable is sampled representatively by equally dividing the interval of each sensitive variable, and introduces a shuffling and rearrangement mechanism when matching combinations, generating a set of perturbed input samples that are uniformly distributed and rich in information. This method shows good coverage and stability in high-dimensional perturbation space, and can effectively reveal the changes in the model response surface shape under the common perturbation of multiple variables.
[0062] In one specific embodiment, assuming a two-dimensional perturbation situation, and the hit variables are NaF molar ratio and reaction temperature, assuming that the perturbation range of NaF molar ratio is set to [4.8, 5.3] and the perturbation range of reaction temperature is set to [80.95] in the core variable sensitive interval matching table, and K is 5. Then the specific process of obtaining the perturbed input sample sequence is as follows:
[0063] A1, divide the NaF molar ratio into five segments: [4.8, 4.9), [4.9, 5.0), [5.0, 5.1), [5.1, 5.2), [5.2, 5.3].
[0064] Divide the reaction temperature into five segments: [80, 83), [83, 86), [86, 89), [89, 92), [92, 95].
[0065] A2, randomly sample a value in each sub-interval, for example:
[0066] NaF molar ratio sampling: {4.83, 4.97, 5.04, 5.21, 5.12}.
[0067] Reaction temperature sampling: {80.5, 85.2, 88.7, 92.1, 83.3}.
[0068] A3, shuffle the sampling set and combine it to form 5 two-dimensional perturbations:
[0069] X1' = (4.83, 88.7), X2' = (4.97, 80.5), X3' = (5.04, 83.3), X4' = (5.21, 85.2), X5' = (5.12, 92.1).
[0070] This gets a sample set {X1', X2', X3', X4', X5'}, i.e. a perturbed input sample sequence, each X' is a d-dimensional vector (d = 2 in this example).
[0071] Further, a perturbed prediction output sequence is obtained, and the specific process is as follows:
[0072] The system loads the decision tree model currently used for the electronic grade anhydrous hydrogen fluoride arsenic removal process data, calls the model inference interface, and inputs the perturbed input sample sequence as a batch of data to the currently executed decision tree model.
[0073] The prediction result corresponding to each perturbed input and output of the currently executed decision tree model represents the predicted value of the residual arsenic content under the perturbed combination, and all output results are combined to form a prediction sequence, denoted as a perturbed prediction output sequence.
[0074] Further, the prediction fluctuation score is analyzed, and the specific process is as follows:
[0075] Based on the perturbed prediction output sequence, the maximum and minimum values of the output prediction result are extracted, and the maximum value of the prediction result is subtracted from the minimum value of the prediction result, i.e. the maximum value of the prediction result is subtracted from the minimum value of the prediction result, to obtain the prediction result range, which is used to measure the upper and lower bound fluctuation range of the prediction result.
[0076] The standard deviation of the perturbed prediction output sequence is calculated, denoted as the prediction result standard deviation, which is used to measure the overall dispersion degree of the prediction result.
[0077] The perturbed prediction output sequence is fitted to a one-dimensional linear trend line, the root mean square error between the perturbed prediction output sequence and the fitted value is calculated, denoted as the prediction result trend residual score, which is used to capture nonlinear jumps or trend mutation behaviors.
[0078] In the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis process, when the input variable hits the model sensitive interval, the system will generate a perturbed input sample set, and these perturbed samples will be sequentially input into the currently executed decision tree model, thereby obtaining the corresponding perturbed prediction output sequence. The sequence records the response change of the model output to the input fine-tuning within the input neighborhood range, and is a numerical basis for analyzing the local stability and response trend of the model.
[0079] To further assess the trend stability of model output, the system takes the perturbation prediction output sequence as a set of discrete numerical points and performs a one-dimensional linear regression fitting to generate an optimal linear trend line. The trend line reflects the overall change direction of the predicted values with the perturbation input, i.e., the linear response tendency of the system to the variable perturbation. If the model's output is relatively stable within the perturbation interval, its predicted values should highly coincide with the trend line, indicating that the model has a stable response to small changes in the variable. Subsequently, the system calculates the root mean square error between the perturbation prediction values and their corresponding linear fitting values, and defines it as the prediction result trend residual score. The prediction result trend residual score quantifies the degree of fluctuation in the model's output and the deviation from the linear trend, and is the core indicator for evaluating whether the model's prediction trend is smooth within the current input neighborhood. If the prediction result trend residual score is small, it means that the perturbation prediction is approximately linear and changes smoothly, and the model has good stability; if the prediction result trend residual score is large, it indicates that the model is sensitive to perturbations and may have structural instability or overfitting problems, and the model structure needs to be further optimized or an output smoothing mechanism needs to be introduced.
[0080] The prediction result range, the prediction result standard deviation, and the prediction result trend residual score are scaled by the preset scaling coefficient, and the maximum value after scaling is taken as the prediction fluctuation score.
[0081] In the data analysis process of electronic-grade anhydrous hydrogen fluoride arsenic removal, the core purpose of the preset scaling coefficient is to normalize the three indicators with different physical meanings and numerical dimensions (prediction result range, prediction result standard deviation, and prediction result trend residual score) so that they can be compared on the same evaluation scale, ensuring that the maximum value replacing the weighted average composite fluctuation score mechanism is reasonable and discriminative. The following is the general setting logic and basis of the preset scaling coefficient: The scaling coefficients set in the system are a set of empirical parameters corresponding to the three types of core prediction fluctuation indicators. The setting method can be summarized as the following three core principles:
[0082] B1, dimension scaling driven by historical data: By statistics of the common value range of each type of fluctuation indicator in the historical perturbation prediction data, the target normal interval is set, for example: the prediction result range is usually distributed in 2 to 10 ppb; the prediction result standard deviation is mostly in 1 to 5 ppb; the trend residual score is mostly concentrated in 0.5 to 3 ppb. For these value ranges, the system can select a unified target interval (e.g., 0 to 1) as the output interval after standardization, and then define a linear scaling coefficient for each type of indicator so that it falls into the target range after normalization. For example, if the target standardized value range is [0, 1], and the typical upper limit of the standard deviation in the historical samples is 5 ppb, then the standard deviation scaling coefficient can be set to 1 / 5.
[0083] B2, Sensitivity adjustment guided by engineering tolerance: Different industrial scenarios have different tolerances to fluctuation types. For example, the prediction result extreme difference represents the span of the upper and lower limits of the prediction, and its sudden change is more likely to cause control misjudgment, so it should be given higher sensitivity; while the trend residual deviation may represent the existence of nonlinear change or boundary anomaly in the system, which also needs to be strengthened in some high requirement processes. Therefore, the scaling coefficient is not only the unit scaling of the physical quantity, but also can be used to reflect the engineering focus. For indicators with low tolerance, set a larger scaling ratio to make it easier to trigger high-risk judgment. Such coefficients can be adjusted by engineers or learned adaptively through Bayesian optimization and other methods.
[0084] B3, Continuous calibration of cross-batch stability: The scaling coefficient is not static configuration, but should have a certain self-calibration mechanism. The system can regularly evaluate the correlation between the prediction fluctuation score and the true arsenic removal error, and fine-tune the scaling coefficient to make the final score result most relevant to the actual error. This mechanism can be based on the evaluation accuracy within the sliding window for dynamic feedback, thereby enhancing the robustness of the model during long-term operation.
[0085] In summary, the core value of the preset scaling coefficient is to unify the fluctuation indicators of different dimensions into a comparable evaluation framework while retaining the ability to identify key abnormal patterns. Its setting is based on a three-in-one of historical value range analysis, engineering sensitivity priority, and dynamic adaptive updating mechanism, which is one of the key foundations to ensure the effectiveness, stability, and landing of the entire prediction fluctuation identification mechanism.
[0086] In specific embodiments, the prediction fluctuation score is specifically represented as follows:
[0087] ,
[0088] Wherein A is the prediction fluctuation score, a is the prediction result extreme difference, b is the prediction result standard deviation, c is the prediction result trend residual score, a1 is the prediction result extreme difference scaling coefficient, b1 is the prediction result standard deviation scaling coefficient, and c1 is the prediction result trend residual score scaling coefficient.
[0089] Further, the determination of the fluctuation judgment processing branch path is specifically analyzed as follows:
[0090] Extract the preset prediction fluctuation score threshold in the database.
[0091] In this embodiment, the prediction fluctuation score threshold is used to determine whether the current model response fluctuation is within an acceptable range. In this embodiment, the preset process of the prediction fluctuation score threshold is as follows: the system extracts a large number of input samples located in the non-sensitive region from the historical sample library, calculates the prediction fluctuation score under the perturbed sample set, and forms a fluctuation score distribution curve in a stable state. Then, the 90% quantile value or the verified optimal stability lower bound is selected from the distribution as the prediction fluctuation score threshold.
[0092] It should be noted that the preset process of the prediction fluctuation score threshold described above is only an example. In specific embodiments, the setting of the prediction fluctuation score threshold can be combined with actual working conditions and requirements for specific setting, and this embodiment is not limited.
[0093] If the prediction fluctuation score is less than the prediction fluctuation score threshold, the fluctuation judgment processing branch path is recorded as directly using the arsenic removal data input sample as input, based on the current decision tree model, to output the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result.
[0094] If the prediction fluctuation score is less than the prediction fluctuation score threshold, it indicates that the prediction result (such as the predicted value of residual arsenic content) obtained by perturbing the input sample sequence has a small overall fluctuation range and low dispersion degree, and has no obvious nonlinear jump or trend mutation behavior. This indicates that the current decision tree model is minimally affected by variable perturbation when processing the arsenic removal data input sample, and the model prediction stability and reliability are high, and there is no need to adjust the prediction result or the model structure. Therefore, the fluctuation judgment processing branch path is determined to directly use the original arsenic removal data input sample as input, calculate and output the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result by the current decision tree model, to ensure analysis efficiency while ensuring result accuracy.
[0095] If the prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, perturbation uncertainty type analysis is performed, and if the prediction result range is less than the preset prediction result range threshold, the perturbation uncertainty type is recorded as input noise type uncertainty, and the fluctuation judgment processing branch path is recorded as performing output smoothing.
[0096] If the prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, it indicates that the prediction result fluctuation caused by the perturbed input sample sequence exceeds the stable range allowed by the system, and there may be a prediction distortion risk, which requires further analysis of the specific type of perturbation uncertainty to develop a targeted processing scheme. This situation usually means that the variable perturbation in the current arsenic removal data input sample has affected the prediction output of the decision tree model, and if the original model output result is directly used, it may lead to production control errors (such as misjudging whether the residual arsenic content meets the standard), so the subsequent uncertainty type analysis and corresponding processing process must be started, rather than directly outputting the result.
[0097] If the prediction result extreme difference is less than the prediction result extreme difference threshold value at this time, it indicates that the overall fluctuation of the prediction result is mainly caused by the slight noise interference in the input data (such as accidental errors when the sensor collects data, slight random fluctuations of process parameters), rather than the sensitive response of the model structure to the variable change, that is, the disturbance uncertainty type is input noise type uncertainty. At this time, the upper and lower limit range of the prediction result is still in the controllable interval, and there is no need to adjust the model structure. Only by executing the output smoothing process, the influence of noise on the final result can be reduced, and the stability of the output result can be ensured.
[0098] Further, the output smoothing is performed, and the specific analysis process is as follows:
[0099] The prediction fluctuation score is subjected to difference processing with the prediction fluctuation score threshold value to obtain a prediction fluctuation deviation score, that is, the value obtained by subtracting the prediction fluctuation score threshold value from the prediction fluctuation score is taken as the prediction fluctuation deviation score.
[0100] The prediction fluctuation deviation score is taken as a query key value to query the smoothing fusion coefficient.
[0101] In this embodiment, the system pre-constructs a configuration table containing the corresponding relationship between the prediction fluctuation deviation score and the smoothing fusion coefficient. This table needs to be generated based on a large amount of measured data or simulation results, and clearly matches the smoothing fusion coefficient in different prediction fluctuation deviation score intervals (low deviation score corresponds to a smaller smoothing fusion coefficient to reduce excessive smoothing, and high deviation score corresponds to a larger coefficient to strengthen the data correction effect); then, the specific prediction fluctuation deviation score is obtained by analysis; finally, the calculated prediction fluctuation deviation score is taken as a query key value for interval matching or accurate searching in the preset configuration table, so that the smoothing fusion coefficient suitable for the current prediction fluctuation deviation state can be quickly obtained, and the smoothing fusion coefficient can accurately respond to the data prediction fluctuation deviation, thereby improving the accuracy and adaptability of data processing.
[0102] The smoothing fusion coefficient is a proportional factor for weighted fusion between the original prediction result and the disturbance prediction mean value, and the value range is 0 to 1. In specific embodiments, when the smoothing fusion coefficient tends to 1, the system tends to retain the original model prediction result; when the smoothing fusion coefficient tends to 0, the disturbance sample prediction mean value is relied on more, thereby relieving the prediction mutation risk.
[0103] The disturbance prediction output sequence is subjected to mean value processing to obtain a prediction result mean value.
[0104] Taking the arsenic removal data input sample as input, the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output by the current decision tree model, and is denoted as an original prediction result.
[0105] Based on the original prediction result, the prediction result average value and the smoothing fusion coefficient, the original prediction result is smoothed to obtain a smoothed prediction result, and a smoothing completion signal is generated synchronously.
[0106] In specific embodiments, the smoothed prediction result is specifically represented as: Wherein y is the smoothed prediction result, y1 is the original prediction result, y2 is the prediction result average value, and d is the smoothing fusion coefficient.
[0107] If the prediction result range is greater than or equal to the preset prediction result range threshold, the disturbance uncertainty type is recorded as a model structure type uncertainty, and thus the fluctuation judgment processing branch path is recorded as executing model structure lightweight adjustment.
[0108] If the prediction result range is greater than or equal to the prediction result range threshold at this time, it indicates that the prediction result has a large amplitude of upper and lower bound jump. Such fluctuation is not caused by input noise, but is caused by the excessive sensitivity of the decision tree model structure (such as variable threshold splitting rule) to the variable in the current hit sensitive interval. A small change in variable input triggers model path switching, resulting in a dramatic fluctuation of the output result, i.e. the disturbance uncertainty type is a model structure type uncertainty. At this time, only output smoothing cannot fundamentally solve the fluctuation problem, and the execution of model structure lightweight adjustment process is required to optimize the response logic of the model to the sensitive interval variable, to reduce the jump risk of the prediction result from the root and to ensure the reliability of the subsequent analysis result.
[0109] Further, the model structure lightweight adjustment is executed, and the specific execution process is as follows:
[0110] Based on the current execution decision tree model, a temporary copy model is generated.
[0111] In this embodiment, when the system detects that the input sample falls into the sensitive region of the model structure, and the disturbance prediction result fluctuation score exceeds the threshold, it indicates that the current model has a prediction instability problem in this local region. At this time, if the main model is directly adjusted, it may cause the degradation of the overall prediction ability, so the main model is kept unchanged to ensure the stable prediction of most samples; based on the current model structure, the model state is copied to generate a temporary copy model for temporary adjustment and experiment, avoiding the high time delay and resource occupation problem caused by complete retraining.
[0112] The core purpose of generating a temporary copy model is to adjust and optimize the structure of the local input sensitive area without affecting the stability of the main model, so as to improve the stability of the prediction result in the boundary area or jump area of the model. The temporary copy model is a copy of the structure and trained parameters of the current main model, which is an independent execution body of the main model. The temporary copy model is consistent with the main model in structure, but the hyperparameters of the control complexity and fitting strength can be adjusted locally.
[0113] The prediction fluctuation deviation score is used as a query key value to query the model structure fine-tuning parameter set.
[0114] The model structure fine-tuning parameter set includes a minimum sub-node sample number improvement value, a pruning threshold increase value, and a split step length shortening value.
[0115] In this embodiment, based on the prediction fluctuation deviation score, a pre-constructed model structure fine-tuning parameter mapping library is called. The mapping library is generated based on experimental data in different fluctuation scenarios, and internally stores the mapping relationship between different prediction fluctuation deviation score intervals and corresponding structure fine-tuning parameters. The mapping rule follows the principle that the prediction fluctuation deviation score is positively correlated with the structure fine-tuning strength: the higher the score (indicating more severe prediction fluctuation), the larger the minimum sub-node sample number improvement value (increasing the node split threshold to enhance model stability), the larger the pruning threshold increase value (strengthening the redundancy structure elimination intensity), and the larger the split step length shortening value (slowing down the tree growth speed to avoid overfitting); on the contrary, the lower the score, the smaller the parameter adjustment value.
[0116] Finally, the calculated prediction fluctuation deviation score is used as a query key value to match the interval in the above mapping library, and the corresponding minimum sub-node sample number improvement value, pruning threshold increase value, and split step length shortening value are extracted after locating the interval, that is, the query of the model structure fine-tuning parameter set is completed.
[0117] The minimum sub-node sample number refers to the minimum number of samples that a sub-node should contain after each split. If a split results in a sub-node with fewer samples than this value, the split will be canceled. This parameter controls whether the model allows the generation of a split node with a small number of samples, which is an important means to prevent overfitting.
[0118] The pruning threshold represents the minimum loss reduction value required for node splitting. Only when a split can bring a loss reduction greater than the threshold will it be executed. This parameter controls the "profit threshold" of the split, thereby avoiding unnecessary splitting in the case of small profit.
[0119] Splitting step or maximum step limit the maximum "moving range" of each tree in the splitting process, especially suitable for offset problems in classification or regression. It helps to slow down the model weight update speed and improve the stability of splitting.
[0120] Increasing the minimum number of child node samples requires each split child node to have more samples, which directly inhibits the splitting behavior optimized only for a small number of sample points. In the boundary region, due to the sparse distribution of samples, this constraint can prevent the model from making highly sensitive fitting in these regions, thereby reducing the possibility of prediction jump.
[0121] Increasing the pruning threshold will make the model need to evaluate a larger benefit before splitting, thus avoiding unnecessary small-scale splitting in sensitive areas of the model boundary. This strategy improves the stability of the model within the neighborhood, making it less sensitive to perturbed inputs.
[0122] By limiting the splitting step, the amplitude of node weight or splitting position in each update process of the model can be constrained, so that the generation process of the entire tree structure is more stable, avoiding some drastic response splitting rules near the variable threshold, thereby reducing the output mutation caused by input perturbation.
[0123] Based on the model structure fine-tuning parameter set, the temporary copy model structure is lightly adjusted, and an adjustment completion signal is generated after the adjustment is completed, and the temporary copy model after the adjustment is completed is assigned a to-be-verified model state identifier.
[0124] After generating the copy model, the system will fine-tune its structural hyperparameters, increase the minimum number of child node samples, increase the pruning threshold, and limit the maximum splitting depth. The adjustment of these parameters helps to suppress the model's overfitting to perturbed inputs in the boundary region and reduce the dramatic fluctuations in predictions. Since the copy model is a light adjustment based on the trained model structure, it does not need to retrain the full model, so it has the advantages of speed, low cost, and online operation.
[0125] The disturbance processing effect analysis module is used to analyze the effect of the disturbance processing after the targeted processing of the disturbance fluctuation is completed, and output the electronic grade anhydrous hydrogen fluoride arsenic removal data analysis result after effective disturbance processing.
[0126] As Figure 2The shown is a disturbance processing effect analysis flowchart related to the embodiment, including two parallel signal receiving paths: when receiving a smoothing completion signal, the system analyzes the original prediction residual standard deviation based on the disturbance prediction output sequence and the original prediction result, then analyzes the smoothing residual standard deviation based on the disturbance prediction output sequence and the smoothed prediction result, and determines whether the original prediction residual standard deviation is greater than the smoothing residual standard deviation, if yes, the disturbance processing effect analysis result is recorded as effective disturbance processing, if not, the disturbance processing effect analysis result is recorded as invalid disturbance processing; when receiving an adjustment end signal, the system inputs the disturbance input sample sequence into the model to be verified and outputs a new disturbance prediction output sequence, and recalculates the prediction fluctuation score based on the new disturbance prediction output sequence, if the recalculated prediction fluctuation score is less than the prediction fluctuation score threshold, the disturbance processing effect analysis result is recorded as effective disturbance processing, if the recalculated prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, the disturbance processing effect analysis result is recorded as invalid disturbance processing.
[0127] Further, the disturbance processing effect analysis is performed, and the specific execution process is as follows:
[0128] If the system receives a smoothing completion signal, the original prediction residual standard deviation is analyzed based on the disturbance prediction output sequence and the original prediction result.
[0129] The smoothing residual standard deviation is analyzed based on the disturbance prediction output sequence and the smoothed prediction result.
[0130] If the original prediction residual standard deviation is greater than the smoothing residual standard deviation, the analysis result of the disturbance processing effect is recorded as effective disturbance processing, otherwise, the analysis result of the disturbance processing effect is recorded as invalid disturbance processing.
[0131] If the original prediction residual standard deviation is greater than the smoothing residual standard deviation, it indicates that the disturbance prediction value is distributed more dispersedly (unstable) under the current model, and the prediction result is more concentrated after smoothing. The original prediction of the current model has instability in the disturbance range; the smoothing processing plays a role in filtering jumps and aggregating fluctuations; at this time, the smoothed result is more reliable than the original result; therefore, the system considers that the disturbance processing is effective, and the smoothed result can be retained for downstream decision or output.
[0132] On the contrary, it indicates that the disturbance prediction value itself is already relatively concentrated, and the smoothing processing after smoothing leads to an un-converged or even expanded fluctuation range, the current smoothing method is not suitable for the distribution structure, and does not play a role; if the smoothed result is forcibly used, it may introduce bias or information loss; therefore, the system judges that the disturbance processing is invalid, and generates a warning prompt information.
[0133] If the system receives an adjustment end signal, the disturbance input sample sequence is input into the model to be verified, and a new disturbance prediction output sequence is output.
[0134] Based on the new disturbance prediction output sequence, the prediction fluctuation score is recalculated, and if the recalculated prediction fluctuation score is less than the prediction fluctuation score threshold, the analysis result of the disturbance treatment effect is recorded as an effective disturbance treatment.
[0135] After the system completes the model structure adjustment (for example, adjusting the minimum sub-node sample number, pruning parameters, etc.), the previously constructed disturbance input sample sequence is re-input into the model to be verified (i.e., the adjusted model), and a new disturbance prediction output sequence is obtained. Then, the system recalculates the prediction fluctuation score of the sequence, which reflects the prediction stability of the adjusted model within the input disturbance interval. If the recalculated prediction fluctuation score is less than the prediction fluctuation score threshold, it indicates that the model adjustment significantly improves the prediction jump problem, and the unstable prediction caused by the original input disturbance has been suppressed or alleviated. The current model has higher robustness to input in the sensitive area. Therefore, the system can determine that this adjustment is an effective disturbance treatment, and the result can be directly adopted, entering the output or archiving stage, without further smoothing or alarm.
[0136] Further, after the effective disturbance treatment, the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output, and the specific execution process is as follows:
[0137] After the effective disturbance treatment, the system performs model switching, replacing the current execution decision tree model with the model to be verified, and assigning the model to be verified after the replacement with the state identifier of the current execution decision tree model. The switching event, adjustment parameters, score improvement amplitude, and operation timestamp are recorded synchronously for subsequent model version management.
[0138] The arsenic removal data input sample is input into the current execution decision tree model, and the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output.
[0139] If the recalculated prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, the analysis result of the disturbance treatment effect is recorded as an invalid disturbance treatment, and a warning prompt is generated.
[0140] If the recalculated prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, it indicates that the current model adjustment strategy fails to improve the disturbance area prediction fluctuation, and may even exacerbate the instability. The model still produces large output differences to small input disturbances, and the jump risk in structure has not been eliminated. The adjustment strategy may not be suitable for such disturbance patterns, and there is a risk of blind adjustment or overfitting repair. Therefore, the system determines that this disturbance treatment is invalid, and triggers the alarm prompt mechanism to prompt the operation and maintenance personnel or system administrator to perform manual intervention or call higher-level strategies (such as model replacement, transfer learning, etc.). This step can avoid repeated adjustment of the system in the wrong direction, ensuring the controllability and engineering safety of the overall prediction output.
[0141] The above-described embodiments can be implemented in whole or in part by software, hardware (e.g., circuitry), firmware, or any combination thereof. When implemented in software, the above-described embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions according to the embodiments of the present application are entirely or partially generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another computer-readable storage medium, for example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center through a wired (e.g., infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. containing one or more available medium collections. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. The semiconductor medium can be a solid-state disk.
[0142] It should be understood that the term "and / or" herein merely describes an association relationship of associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it, but it can also represent an "and / or" relationship, which can be understood according to the context before and after it.
[0143] In the present application, "at least one" means one or more, and "multiple" means two or more. "At least one of the following" or the like means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple.
[0144] It should be understood that in various embodiments of the present application, the size of the sequence number of the above-described processes does not mean the order of execution, and the execution order of the processes should be determined according to their functions and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0145] Those skilled in the art can clearly understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0146] Those skilled in the art can clearly understand that, for the convenience and brevity of the description, the specific working processes of the devices, apparatuses and units described above can refer to the corresponding processes in the foregoing method embodiments, which will not be repeated here.
[0147] In addition, each functional unit in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit.
[0148] If the functions are realized in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various program code storage media.
[0149] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. An electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model, characterized in that, The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device.
2. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 1, characterized in that, The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device.
3. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 1, characterized in that, The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device.
4. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 3, characterized in that, The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device.
5. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on decision tree model according to claim 1, characterized in that, The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application relates to an electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis method and device. The application a standard deviation of the perturbation prediction output sequence, denoted as a prediction result standard deviation, is calculated, and the prediction result standard deviation is used to measure the overall dispersion degree of the prediction result; a root mean square error between the perturbation prediction output sequence and a fitting value of a one-dimensional linear trend line is calculated, and the root mean square error is denoted as a prediction result trend residual error score, which is used to capture nonlinear jump or trend mutation behaviors; the prediction result range, the prediction result standard deviation and the prediction result trend residual error score are scaled by a preset scaling coefficient, and a maximum value after the scaling is taken as a prediction fluctuation score.
6. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on decision tree model according to claim 1, characterized in that, The specific analysis process of the determination of the fluctuation judgment processing branch path is as follows: a prediction fluctuation score threshold is extracted; if the prediction fluctuation score is less than the prediction fluctuation score threshold, the fluctuation judgment processing branch path is recorded as directly taking the arsenic removal data input sample as input, outputting the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result based on the currently executed decision tree model; if the prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, the perturbation uncertainty type analysis is performed, if the prediction result range is less than a preset prediction result range threshold, the perturbation uncertainty type is recorded as an input noise type uncertainty, and thus the fluctuation judgment processing branch path is recorded as executing output smoothing; if the prediction result range is greater than or equal to the preset prediction result range threshold, the perturbation uncertainty type is recorded as a model structure type uncertainty, and thus the fluctuation judgment processing branch path is recorded as executing model structure lightweight adjustment.
7. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 6, characterized in that, The specific analysis process of the execution of the output smoothing is as follows: a difference value between the prediction fluctuation score and the prediction fluctuation score threshold is processed to obtain a prediction fluctuation deviation score; the prediction fluctuation deviation score is taken as a query key value to query a smoothing fusion coefficient; a mean value of the perturbation prediction output sequence is obtained through mean processing; the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output by taking the arsenic removal data input sample as input through the currently executed decision tree model, and the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is recorded as an original prediction result; the original prediction result is smoothed based on the original prediction result, the prediction result mean value and the smoothing fusion coefficient to obtain a smoothed prediction result, and a smoothing completion signal is generated synchronously.
8. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 6, characterized in that, The specific execution process of the execution of the model structure lightweight adjustment is as follows: a temporary copy model is generated based on the currently executed decision tree model; the prediction fluctuation deviation score is taken as a query key value to query a model structure fine-tuning parameter set; the model structure fine-tuning parameter set includes a minimum sub-node sample number improvement value, a pruning threshold increase value and a split step length shortening value; the model structure fine-tuning parameter set is used to perform lightweight adjustment on the temporary copy model structure, and an adjustment completion signal is generated after the adjustment is completed, and the temporary copy model after the adjustment is given a to-be-inspected model state identifier synchronously.
9. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on decision tree model according to claim 1, characterized in that, The specific execution process of the perturbation processing effect analysis is as follows: if the system receives the smoothing completion signal, the original prediction residual standard deviation is analyzed based on the perturbation prediction output sequence and the original prediction result; the smoothing residual standard deviation is analyzed based on the perturbation prediction output sequence and the smoothed prediction result; If the original prediction residual standard deviation is greater than the smoothed residual standard deviation, the analysis result of the disturbance processing effect is recorded as effective disturbance processing, otherwise, the analysis result of the disturbance processing effect is recorded as ineffective disturbance processing; If the system receives the adjustment end signal, the disturbance input sample sequence is input into the model to be verified, and a new disturbance prediction output sequence is output; Based on the new disturbance prediction output sequence, the prediction fluctuation score is recalculated, and if the recalculated prediction fluctuation score is less than the prediction fluctuation score threshold, the analysis result of the disturbance processing effect is recorded as effective disturbance processing; If the recalculated prediction fluctuation score is greater than or equal to the prediction fluctuation score threshold, the analysis result of the disturbance processing effect is recorded as ineffective disturbance processing, and a warning prompt is generated.
10. The electronic grade anhydrous hydrogen fluoride arsenic removal data analysis system based on a decision tree model according to claim 9, characterized in that, After effective disturbance processing, the system outputs the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result, and the specific execution process is as follows: After effective disturbance processing, the system executes model switching, replaces the current execution decision tree model with the model to be verified, and assigns the replaced model to be verified with the current execution decision tree model state identifier, synchronously records the switching event, adjustment parameters, score improvement amplitude and operation timestamp for subsequent model version management; With the arsenic removal data input sample, the current execution decision tree model is input, and the electronic-grade anhydrous hydrogen fluoride arsenic removal data analysis result is output.
Citation Information
Patent Citations
A federated learning method and system for vertical XGBoost decision trees
CN114169537B
Principal component sensitivity analysis method and device for sewage treatment and medium
CN115034434A
Big data bamboo tableware manufacturing system based on production process monitoring
CN120388244A
Prediction method for memory effect of natural gas hydrate
CN120544734A
Multivariable time series data-oriented interpretability prediction analysis system
CN120596826A