Fluororubber formula regulation range sensing double-stage modeling method, electronic equipment and storage medium

Through the range-aware two-stage modeling method, the known functional group characteristics and unknown structural fingerprints of fluororubber are integrated, which solves the complex problem of monomer ratio and macroscopic performance mapping in fluororubber formula optimization, and realizes the scientific guidance of the analysis of unknown structural fingerprints and formula regulation.

CN120805683APending Publication Date: 2025-10-17SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510906805.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing technologies lack a quantitative conduction path model for optimizing fluororubber formulations, are unable to effectively utilize the structural information contained in NMR spectra, and traditional machine learning methods find it difficult to reveal the mapping relationship between monomer ratios and macroscopic properties.

Method used

A range-aware two-stage modeling method is adopted to integrate known functional group features with unknown structural fingerprints by constructing a model network structure, including input layer, data separation layer, enhanced feature engineering layer, range-aware feature selection layer, first-stage encoder, intermediate representation layer, enhanced PLS dimensionality reduction branch, second-stage decoder and output layer, and use an integrated regression model for feature mapping and performance prediction.

Benefits of technology

It achieves a clear mapping between monomer ratios and macroscopic properties, solves the shortcomings of unknown structural fingerprint analysis in existing technologies, provides scientific guidance for fluororubber formula regulation, and improves the efficiency and transparency of formula optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120805683A_ABST
    Figure CN120805683A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of chemical materials, and provides a fluororubber formula regulation range sensing double-stage modeling method, electronic equipment and a storage medium. The method comprises the steps of model network structure construction, data receiving, packet processing, enhanced feature processing, feature screening, first-stage coding, PLS dimension reduction, second-stage decoding, index prediction distribution output and model iteration training. According to the method, a complex monomer ratio and macroscopic performance mapping problem is decomposed into two relatively simple sub-problems with clear physical significance, and meanwhile, a range sensing technology is introduced, so that a nonlinear relationship is effectively captured, and the feature mining effect of low-variability data is improved; a more complete microstructure characterization system is constructed by integrating known functional group characteristics and unknown structure fingerprints.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of chemical materials, and particularly relates to a fluorine rubber formula regulated range perception two-stage modeling method, an electronic device and a storage medium. BACKGROUND

[0002] Vinylidene fluoride is abbreviated as VDF; tetrafluoroethylene is abbreviated as TFE; perfluoromethyl vinyl ether is abbreviated as PMVE; nuclear magnetic resonance spectroscopy is abbreviated as NMR. Fluororubber, as a kind of special elastomer material with good comprehensive performance, has become an important material in the fields of aerospace, automobile industry and chemical industry, etc. due to its excellent high temperature resistance, corrosion resistance, oil resistance and oxidation resistance. With the rapid development of industrial technology and the more stringent application environment, higher requirements are put forward for the performance of fluororubber, especially in maintaining the balance of multiple properties while achieving precise control of specific properties. Fluororubber formulation design is a complex system engineering involving multiple dimensions, including monomer ratio design, vulcanization system optimization, filler system configuration and other key links. Among these dimensions, monomer ratio design occupies an important position, to a large extent determines the chemical structure of the main chain, the distribution of functional groups and the characteristics of intermolecular interaction, and is a key factor for the realization of performance directional control and formulation optimization of fluororubber. Modern fluororubber usually adopts a multi-component copolymer design strategy, which realizes the directional control of performance by effectively controlling the ratio of different functional monomers. Taking the widely used ternary fluororubber system as an example, VDF provides the basic elastic skeleton, TFE enhances the chemical resistance and thermal stability, and PMVE adjusts the vulcanization performance and processing performance. The ratio change of the three monomers will cause systematic changes in the molecular chain sequence distribution, crystallinity, glass transition temperature and other microstructure parameters, and then have a complex synergistic effect on the macroscopic properties such as Mooney viscosity, tensile strength, elongation and hardness. However, the lack of in-depth understanding of this complex transmission mechanism limits the rational optimization design and mechanism understanding of fluororubber formulation. The traditional method of optimizing fluororubber formulation mainly relies on the trial-and-error method, which is widely used but inefficient and difficult to reveal the internal mechanism of formulation control, and cannot provide scientific guidance and transparent analysis tools for rational optimization. In order to establish a more accurate structure-property relationship and guide the formulation optimization, researchers have turned their attention to microstructure characterization techniques. NMR provides a powerful tool for accurate characterization of fluororubber molecular structure. Researchers have used multi-dimensional NMR techniques to analyze the fine sequence structure of VDF-HFP-TFE ternary fluororubber, and researchers have studied the influence mechanism of different end group structures on the performance of fluororubber through end group analysis. Although these studies have made important progress in understanding the local structure-property relationship, there are still important deficiencies in the application of formulation optimization: most existing researches are still at the qualitative analysis level, and lack of quantitative transmission path models that can guide the formulation control; the rich structure information contained in the NMR spectrum has not been fully utilized, and a large number of unassigned spectral peak regions may contain key "structure fingerprint" information; there is a lack of complete modeling framework to systematically correlate multi-dimensional microstructure characteristics with macroscopic properties, which cannot provide clear control strategies for formulation optimization. In recent years, the success of machine learning in the field of materials science has provided a new way for the optimization of fluororubber formulation. However, the unique complexity of fluororubber has posed new challenges for existing machine learning methods.On the one hand, the synthesis process of fluoroelastomer needs to be carried out under harsh conditions such as high temperature and high pressure, which not only makes the sample preparation cost high, but also requires strict control accuracy of the experiment; on the other hand, the performance data of fluoroelastomer usually presents low variability, and the slight change of performance is not easy to be captured by traditional regression models and machine learning methods. Therefore, the particularity of fluoroelastomer makes the traditional method face not small challenges in this field.

[0003] Although some traditional black box models such as linear regression, decision tree and the like have interpretability, existing researches are mostly limited to the statistical correlation level, lack of interpretability of the fluoroelastomer formula optimization process, and are difficult to reveal how the monomer ratio affects the final performance through the microstructure change. Such limitations make it difficult for these methods to provide transparent analysis tools for the formula optimization of fluoroelastomer from the perspective of molecular mechanism. SUMMARY

[0004] In order to overcome the deficiencies of the prior art, the purpose of the present application is to provide a range-aware two-stage modeling method for fluoroelastomer formula regulation, an electronic device and a storage medium, which solves the problems of complex mapping of monomer ratio and macroscopic performance and lack of unknown structure fingerprint analysis in the prior art.

[0005] To achieve the above-mentioned purpose, the present application provides the following scheme:

[0006] A range-aware two-stage modeling method for fluoroelastomer formula regulation, comprising:

[0007] constructing a model network structure; the model network structure comprises: an input layer, a data separation layer, an enhanced feature engineering layer, a range-aware feature selection layer, a first stage encoder, an intermediate representation layer, an enhanced PLS dimension reduction branch, a second stage decoder and an output layer connected in turn; the second stage decoder comprises a strategy selector;

[0008] receiving original multivariate chemical data through the input layer;

[0009] grouping the original multivariate chemical data by using the data separation layer to obtain grouped data;

[0010] performing basic ratio maintaining, interaction ratio calculation, functional group aggregation and correlation screening on the grouped data by using the enhanced feature engineering layer to obtain enhanced feature data;

[0011] performing feature screening on the enhanced feature data by using the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data; the expression of the comprehensive evaluation formula is: CS(X,y)=0.7×RS(X,y)+0.3×|Corr(X,y)|; wherein, CS(X,y) is the comprehensive score; RS(X,y) is the range score; Corr(X,y) is the absolute value of the correlation; σ global (y) is the global standard deviation of the target variable; Represents σ global (y) performs square operation; σ group (y) is the standard deviation of the target variable y within the i-th quartile group of the feature matrix X; k is the number of quartile groups;

[0012] In the first-stage encoder, an integrated regression model is used to construct a mapping from monomer ratios to the key feature data and a leave-one-out cross-validation is performed to obtain predicted feature data, and the predicted feature data is transferred to the second-stage decoder using the intermediate representation layer;

[0013] Performing dimensionality reduction and range indication generation on the predicted feature data using the enhanced PLS dimensionality reduction branch to obtain dimensionality reduction indication data, and transmitting the dimensionality reduction indication data to the second-stage decoder;

[0014] According to the dimensionality reduction indication data, the predicted feature data is processed in parallel using the second-stage decoder, the target variable is transformed, and the extreme value samples are resampled to obtain performance indicator data, and the performance indicator data is subjected to leave-one-out cross-validation maximization strategy matching using the strategy selector to obtain the indicator prediction distribution; the parallel processing paths include: full NMR, monomer and NMR fusion, PLS dimensionality reduction, and customized features;

[0015] Outputting the indicator prediction distribution using the output layer;

[0016] The constructed model network structure is trained according to the indicator prediction distribution to obtain a fluororubber formula control model.

[0017] Preferably, receiving the raw multivariate chemical data through the input layer comprises:

[0018] Performing Savitzky-Golay filtering to reduce noise on the original multivariate chemical data;

[0019] performing baseline correction on the raw multivariate chemical data after noise reduction;

[0020] The corrected raw multivariate chemical data were intensity normalized.

[0021] Preferably, the data separation layer is used to perform grouping processing on the original multi-dimensional chemical data to obtain grouped data, including:

[0022] extracting functional group feature data in the original multi-chemical data according to a preset chemical structure; the functional group feature data includes: OCF3, CH2-CF2, CF2-CF2;

[0023] performing interval feature recognition on the original multi-chemical data by using DBSCAN clustering to obtain unknown structure fingerprints;

[0024] fusing the functional group feature data and the unknown structure fingerprints to obtain the grouping data.

[0025] Preferably, the grouping data is subjected to basic ratio retention, interactive ratio calculation, functional group aggregation and correlation screening by using the enhanced feature engineering layer to obtain enhanced feature data, including:

[0026] extracting regional peak area, known point characteristic signal intensity, unknown regional peak area and range awareness feature of the grouping data after correlation screening;

[0027] constructing the feature matrix by using the known regional peak area, the known point characteristic signal intensity, the unknown regional peak area and the range awareness feature; the expression of the feature matrix is:

[0028]

[0029] wherein, r is a monomer ratio set; r1 is the first monomer ratio in the monomer ratio set; F is a set of the known regional peak area and the known point characteristic signal intensity; U is the unknown regional peak area; R is the range awareness feature; F1 is the first parameter in the known regional peak area; F 1* is the first parameter in the known point characteristic signal intensity; U1 is the first parameter in the unknown regional peak area; R1 is the first parameter in the range awareness feature; n is the total number of samples;

[0030] standardizing the feature matrix by using a StandardScaler standardization formula to obtain the enhanced feature data.

[0031] Preferably, in the first stage encoder, an integrated regression model is used to construct a mapping from monomer ratio to the key feature data and perform leave-one-out cross-validation to obtain predicted feature data, and the intermediate representation layer is used to pass the predicted feature data to the second stage decoder, including:

[0032] presetting the integrated regression model; the integrated regression model includes: a Ridge regression model, an ElasticNet regression model, a Huber regression model, a support vector regression model and a gradient boosting regression model;

[0033] setting a parameter optimization range of the integrated regression model; the parameter optimization range of the regularization parameter of the Ridge regression model is [0.1, 1, 10, 100]; the parameter optimization range of the alpha parameter of the ElasticNet regression model is [0.1, 1, 10]; the parameter optimization range of the L1 ratio parameter of the ElasticNet regression model is [0.1, 0.5, 0.7, 0.9]; the parameter optimization range of the complexity parameter of the support vector regression model is [0.1, 1, 10]; the learning rate of the gradient boosting regression model is 0.1, and the maximum depth is 3.

[0034] Preferably, the enhanced feature data is screened by the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data, including:

[0035] The enhanced feature data is calculated by the comprehensive evaluation formula to obtain the comprehensive score;

[0036] The comprehensive score is sorted in descending order, and a set number of features with the highest scores are selected to obtain the key feature data;

[0037] When the number of the same group of the comprehensive score is less than the set number, the top three features with the highest scores are selected to obtain the key feature data.

[0038] Preferably, the functional group feature data in the original multi-element chemical data is extracted according to a preset chemical structure, including:

[0039] Nine key chemical shift intervals and 33 characteristic peak positions in the original multi-element chemical data are identified to obtain original extraction data; the key chemical shift intervals include: -52.9ppm to -55.3ppm, -125.7ppm to -128.4ppm;

[0040] The features with a peak area ratio less than 0.5% of the full spectrum or a signal intensity lower than a preset signal threshold in the original extraction data are removed to obtain the original extraction data.

[0041] Preferably, an electronic device includes at least one processor and a memory connected in communication with the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to perform the aforementioned range-aware two-stage modeling method for regulating the formula of fluorine rubber.

[0042] Preferably, a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the aforementioned range-aware two-stage modeling method for regulating the formula of fluorine rubber.

[0043] The present application discloses the following technical effects:

[0044] The present application provides a fluorine rubber formula regulation range perception two-stage modeling method, electronic equipment and storage medium, through the two-stage modeling method of the first stage encoder and the second stage decoder, the problem that the single body ratio and the macro performance mapping are relatively complex is solved, the complex single body ratio and the macro performance mapping problem is decomposed into two relatively simple and clear physical meaning sub-problems, through the integration of known functional group characteristics and unknown structure fingerprint, the problem that the prior art lacks unknown structure fingerprint analysis is solved, and the microstructure characterization system is perfected. BRIEF DESCRIPTION OF DRAWINGS

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0046] Figure 1 The fluorine rubber formula regulation range perception two-stage modeling process schematic diagram provided for the embodiments of the present application;

[0047] Figure 2 The two-stage modeling flowchart provided for the embodiments of the present application;

[0048] Figure 3 The range perception two-stage modeling system architecture diagram provided for the embodiments of the present application;

[0049] Figure 4 The fluorine rubber performance index narrow distribution characteristic and distribution ratio schematic diagram provided for the embodiments of the present application;

[0050] Figure 5 The prediction accuracy schematic diagram of the three methods in different performance quantile interval provided for the embodiments of the present application. DETAILED DESCRIPTION

[0051] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only constitute some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0052] The application aims to provide a fluororubber formula regulation range-aware two-stage modeling method, an electronic device and a storage medium, solve the problem that mapping of monomer ratio and macroscopic performance is relatively complex, and the problem that the prior art lacks unknown structure fingerprint analysis.

[0053] In order to make the above-mentioned purposes, features and advantages of the application more obvious and easy to understand, the application will be further described in detail below in combination with the drawings and specific embodiments.

[0054] Figure 1 The fluororubber formula regulation range-aware two-stage modeling process schematic diagram provided by the embodiments of the application is shown in FIG. 1. Figure 1 As shown in FIG. 1, the application provides a fluororubber formula regulation range-aware two-stage modeling method, which comprises the following steps.

[0055] Step 100: constructing a model network structure; the model network structure comprises an input layer, a data separation layer, an enhanced feature engineering layer, a range-aware feature selection layer, a first-stage encoder, an intermediate representation layer, an enhanced PLS dimension reduction branch, a second-stage decoder and an output layer connected in sequence; the second-stage decoder comprises a strategy selector;

[0056] Step 200: receiving original multivariate chemical data through the input layer;

[0057] Step 300: grouping the original multivariate chemical data by using the data separation layer to obtain grouped data;

[0058] Step 400: performing basic ratio maintenance, interaction ratio calculation, functional group aggregation and correlation screening on the grouped data by using the enhanced feature engineering layer to obtain enhanced feature data;

[0059] Step 500: performing feature screening on the enhanced feature data by using the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data; the expression of the comprehensive evaluation formula is: CS(X,y)=0.7*RS(X,y)+0.3*|Corr(X,y)|; wherein, CS(X,y) is a comprehensive score; RS(X,y) is a range score; Corr(X,y) is an absolute value of correlation; σ global (y) is a global standard deviation of a target variable; represents squaring σ global (y); σ group (y) is a standard deviation of a target variable y in the i-th quartile group of a feature matrix X; k is the number of quartiles.

[0060] Step 600: constructing a mapping of monomer ratio to the key feature data using an ensemble regression model in the first stage encoder and performing leave-one-out cross-validation to obtain predicted feature data, and passing the predicted feature data to the second stage decoder using the intermediate representation layer;

[0061] Step 700: performing dimensionality reduction and range indication generation on the predicted feature data using the enhanced PLS dimensionality reduction branch to obtain dimensionality reduction indication data, and passing the dimensionality reduction indication data to the second stage decoder;

[0062] Step 800: performing parallel processing, target variable transformation, and extreme sample resampling on the predicted feature data using the second stage decoder according to the dimensionality reduction indication data to obtain performance indicator data, and performing leave-one-out cross-validation maximum strategy matching on the performance indicator data using the strategy selector to obtain an indicator prediction distribution; the parallel processing paths include: full NMR, monomer and NMR fusion, PLS dimensionality reduction, and custom features;

[0063] Step 900: outputting the indicator prediction distribution using the output layer;

[0064] Step 1000: training the constructed model network structure according to the indicator prediction distribution to obtain a fluoroelastomer formulation regulation model.

[0065] Specifically, the original multivariate chemical data is received through the input layer, including:

[0066] The original multivariate chemical data is subjected to Savitzky-Golay filtering and denoising;

[0067] The denoised original multivariate chemical data is subjected to baseline correction;

[0068] The corrected original multivariate chemical data is subjected to intensity normalization.

[0069] Further, the original multivariate chemical data is subjected to grouping processing using the data separation layer to obtain grouped data, including:

[0070] Functional group feature data is extracted from the original multivariate chemical data according to a preset chemical structure; the functional group feature data includes: OCF3, CH2-CF2, CF2-CF2;

[0071] Interval feature recognition is performed on the original multivariate chemical data using DBSCAN clustering to obtain unknown structure fingerprints;

[0072] The functional group feature data and the unknown structure fingerprints are fused to obtain the grouped data.

[0073] Specifically, the grouping data is subjected to basic proportioning maintenance, interaction ratio calculation, functional group aggregation and correlation screening by using the enhanced feature engineering layer to obtain enhanced feature data, including:

[0074] extracting regional peak area, known point characteristic signal intensity, unknown regional peak area and range awareness feature of the grouping data after correlation screening;

[0075] constructing the feature matrix by using the known regional peak area, the known point characteristic signal intensity, the unknown regional peak area and the range awareness feature; the expression of the feature matrix is:

[0076]

[0077] wherein, r is a monomer proportion set; r1 is the first monomer proportion in the monomer proportion set; F is a set of the known regional peak area and the known point characteristic signal intensity; U is the unknown regional peak area; R is the range awareness feature; F1 is the first parameter in the known regional peak area; F 1* is the first parameter in the known point characteristic signal intensity; U1 is the first parameter in the unknown regional peak area; R1 is the first parameter in the range awareness feature; n is the total number of samples;

[0078] standardizing the feature matrix by using a StandardScaler standardization formula to obtain the enhanced feature data.

[0079] Further, in the first stage encoder, an integrated regression model is used to construct a mapping of monomer proportion to the key feature data and perform leave-one-out cross-validation to obtain predicted feature data, and the intermediate representation layer is used to pass the predicted feature data to the second stage decoder, including:

[0080] presetting the integrated regression model; the integrated regression model includes a Ridge regression model, an ElasticNet regression model, a Huber regression model, a support vector regression model and a gradient boosting regression model;

[0081] setting a parameter optimization range of the integrated regression model; the parameter optimization range of the regularization parameter of the Ridge regression model is [0.1, 1, 10, 100]; the parameter optimization range of the alpha parameter of the ElasticNet regression model is [0.1, 1, 10]; the parameter optimization range of the L1 ratio parameter of the ElasticNet regression model is [0.1, 0.5, 0.7, 0.9]; the parameter optimization range of the complexity parameter of the support vector regression model is [0.1, 1, 10]; the learning rate of the gradient boosting regression model is 0.1, and the maximum depth is 3.

[0082] Specifically, the enhanced feature data is filtered by the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data, including:

[0083] The enhanced feature data is calculated by the comprehensive evaluation formula to obtain the comprehensive score;

[0084] The comprehensive score is sorted in descending order, and the top set number of features with the highest scores are selected to obtain the key feature data;

[0085] When the number of the same group of the comprehensive score is less than the set number, the top three features with the highest scores are selected to obtain the key feature data.

[0086] Further, functional group feature data in the original multi-element chemical data is extracted according to a preset chemical structure, including:

[0087] Nine key chemical shift intervals and 33 feature peak positions in the original multi-element chemical data are identified to obtain original extraction data; the key chemical shift intervals include: -52.9ppm to -55.3ppm, -125.7ppm to -128.4ppm;

[0088] Features with a peak area ratio less than 0.5% of the full spectrum or a signal intensity lower than a preset signal threshold in the original extraction data are removed to obtain the original extraction data.

[0089] Specifically, Figure 2 The complete framework of the fluororubber structure-property range-aware two-stage modeling method proposed in this embodiment is shown. The method includes four core levels: Figure 2 The top part decomposes the complex mapping into two sub-problems; the middle part constructs a systematic feature engineering workflow; and the bottom part establishes a multi-dimensional explainability analysis path.

[0090] Optionally, based on 52 groups of industrial-grade ternary fluororubber (PMVE / VDF / TFE) production data, the complete formula, microstructure, and performance information are covered. The data set includes monomer ratio parameters (satisfying the mass balance constraint, the sum of monomer ratios is 100%), four key performance indicators (Mooney viscosity ML(1+10)121℃, tensile strength, compression set, and elongation), and high-resolution 19F-NMR spectrum data (400MHz, chemical shift range-218.56ppm to 18.56ppm, resolution 0.01ppm). All samples are prepared and tested according to the industrial standardization process.

[0091] Further, all samples were prepared according to the industry standardization process, and NMR data was quality controlled using standard pretreatment process, including Savitzky-Golay filter denoising, baseline correction, intensity normalization, etc.

[0092] Specifically, range-aware feature engineering. Range score indicator (RS), the present embodiment proposes the concept of range score (RangeScore), which quantifies the feature's ability to distinguish the range of the target variable by calculating the ratio of the standard deviation of the target variable within the four quartile groups of the feature to the global standard deviation. Its formula is:

[0093]

[0094] wherein σ global (y) is the global standard deviation of the target variable, σ group (y) is the standard deviation of the target variable y within the i-th quartile group of the feature X, and k is the number of quartile groups. The larger the ratio, the better the feature can distinguish the target variable.

[0095] Using a comprehensive scoring mechanism, combining range score (70% weight) and absolute value of correlation (30% weight), a comprehensive score is created to balance the statistical correlation and range distinguishing ability of the feature:

[0096] CS(X, y) = 0.7 x RS(X, y) + 0.3 x |Corr(X, y)|

[0097] Highlighting the range distinguishing ability while maintaining the statistical correlation. Therefore, the present embodiment designs multiple scoring group experiments to verify that the weights 0.7 and 0.3 are finally obtained as the scoring weight distribution.

[0098] Preferably, the RS score formula application process: the application process of the range discrimination ability score formula, the range discrimination ability score formula Range_score plays a key role in the range perception feature selection. When the system receives the feature value and target value data, first, data preprocessing is performed to ensure that the input is in one-dimensional array format, and if the input is in one-dimensional array format, the first column is automatically extracted, and at the same time, it is verified that the feature value has a sufficient number of samples and is not all the same value, this step ensures the effectiveness of the subsequent calculation. Next, the system performs a feature grouping process, which divides the feature values into four equal frequency groups according to the numerical size. This grouping method can ensure that each group contains a similar number of samples. After grouping, the system verifies that at least two different groups are generated, and if the grouping fails, it directly returns a score of zero. After the establishment of each group, the system performs statistical analysis on the target variable in each group, and calculates the standard deviation of the target variable value in each group. This standard deviation reflects the fluctuation degree of the target variable in the feature value interval. The system only retains groups with more than one sample for calculation to ensure the reliability of the statistical results, and then sums up the standard deviation values of all valid groups for subsequent calculation. In the global statistical calculation stage, the system calculates the global standard deviation of all target variables and the average of the standard deviations in each group. The global standard deviation represents the overall variation of the target variable in the entire data set, while the average group standard deviation represents the average variation within each group after feature grouping. These two statistics constitute the numerator and the denominator of the range score formula The final range score is obtained by dividing the global standard deviation by the average group standard deviation. A small value of 1e-6 is added to the denominator to avoid division by zero errors. The physical meaning of this ratio is the range discrimination ability of the feature: the larger the ratio, the better the feature can group the target variable, i.e., the difference between groups is large and the difference within groups is small, indicating that the feature has good target variable discrimination ability in different numerical intervals.

[0099] Preferably, in the application process of the comprehensive feature scoring CS formula, the comprehensive feature scoring formula Combined_score organically combines the range discrimination ability with the statistical correlation to form a more comprehensive feature evaluation system. In the correlation calculation stage, the system calculates the Spearman rank correlation coefficient for each feature and the target variable. This non-parametric correlation method can capture nonlinear monotonic relationships and is more suitable for the characteristics of chemical engineering data. The system only retains features with a p-value less than 0.1 to ensure statistical significance, and then calculates the absolute value of the correlation coefficient, because both positive and negative correlations indicate the strength of the association between the feature and the target. In the range score calculation stage, the system calls the aforementioned range discrimination ability scoring function to calculate the range_score value for each feature. During this process, the system will perform score standardization to ensure that all scores are scalar values. For non-scalar or abnormal scores that appear during the calculation process, the system will set them to 0.0 to ensure the numerical stability of subsequent calculations. The score integration process embodies the core design concept of this method. The system assigns 70% weight to the range score and 30% weight to the correlation score (the weights are averaged by performing multiple range score comparisons and applying them to model comparisons). This weighting strategy reflects the emphasis on the feature's ability to discriminate across different numerical ranges. The high weighting of the range score emphasizes the feature's discriminatory value in data with narrow distributions, which is particularly important for concentrated distributions common in chemical engineering. The high weighting of the correlation score ensures that the statistical association between the feature and the target variable is not overlooked.

[0100] Furthermore, during the feature sorting and selection phase, the system ranks all features from high to low based on their overall score, then selects the top max_features features with the highest scores (max_features is typically set to 10). To avoid model performance issues caused by too few features, the system ensures that at least three features are selected. During the selection process, the system outputs the top three highest-scoring features and their detailed scoring information, including correlation coefficients, range scores, and p-values, providing users with transparency into feature selection.

[0101] Preferably, the two formulas synergize. The aforementioned two formulas form a multi-level synergistic mechanism in the range-aware two-stage modeling. In the first level application, the range discrimination capability scoring formula runs independently on each feature, systematically identifying those features with good range discrimination capability, effectively filtering out redundant features that cannot distinguish the target variable in different numerical intervals. This screening process is particularly suitable for chemical engineering data, as such data often presents concentrated distribution characteristics, and traditional variance or standard deviation indicators may not effectively identify key features. In the second level application, the comprehensive scoring formula weights and integrates range discrimination capability and statistical correlation, ensuring both the statistical significance of the features and the distinguishing value of the features in the actual numerical distribution. This integration strategy avoids the problem of missing important features by relying solely on correlation, and avoids the risk of selecting statistically insignificant features based solely on range scores. In the third level application, the features selected by the two formulas are input into the subsequent machine learning model for training, and the effectiveness of feature selection is verified by methods such as leave-one-out cross-validation. The system also applies techniques such as target transformation and extreme value resampling during model training to further enhance the model's learning ability for range distribution characteristics. The entire synergistic mechanism ensures consistency from feature selection to model training, enabling the final prediction model to have stronger ability to handle concentrated distribution data and identify extreme value samples. The value of this dual-formula synergistic design lies in its targeted solution to the key challenge in chemical engineering modeling: how to identify features that still have important distinguishing significance in a narrow range in the case of relatively concentrated data distribution, thereby improving the model's sensitivity to small changes in process parameters and prediction accuracy.

[0102] Specifically, feature system construction. According to the workflow shown in part Figure 2 Based on the 19F-NMR spectrum, three types of complementary features are constructed. The specific use process of the three types of features:

[0103] Known functional group features (F-zone features) as prior knowledge carriers of chemical structure, in the feature extraction stage, the F-zone area (f_areas) and intensity (f_intensities) features are grouped into OCF3, CH2-CF2, CF2-CF2, etc. functional group modules through chemical structure knowledge. In the first stage modeling, these features are used as target variables, and an integrated regression model (such as GradientBoosting) driven by monomer ratio and process parameters is established to establish a chemical mechanism-oriented ratio-structure mapping; the second stage is used as a core input feature, and participates in performance prediction together with other features, and through Spearman correlation analysis, the top 30 features most related to performance are selected to strengthen the effectiveness of the physical conduction path;

[0104] Unknown structure fingerprints (U-zone features) are identified by DBSCAN clustering to recognize the high signal but not explicitly attributed structure regions (U1-U15 interval features) in the spectrum, which together with F-zone features constitute the complete NMR feature matrix. In the range-aware feature selection, U-zone features and F-zone features participate in the correlation evaluation and range differentiation ability scoring with equal weight, especially in the prediction strategy for pc performance (compression deformation rate), the implicit structure information of U-zone features is given higher priority, realizing the complementarity of chemical prior knowledge and data-driven discovery;

[0105] Range-aware features are constructed through chemical knowledge-guided feature engineering: on the one hand, continuous value features reflecting the relative content of functional groups are generated (such as OCF3 / CF2-CF2 ratio), and on the other hand, high / low threshold classification indicators (such as ocf3_cf2_high / low) are generated based on quantiles (25% and 75%) to dynamically identify the extreme value region of feature distribution. ocf3_cf2_high / low is a high / low threshold classification indicator generated based on quantiles (25% and 75%), corresponding to OCF3 and CF2-CF2 modules in functional group features. Samples higher than the 75% quantile are labeled as ocf3_cf2_high (representing the relative relationship of the two types of functional groups in the "high extreme interval"), and samples lower than the 25% quantile are labeled as ocf3_cf2_low (representing the "low extreme interval"), which are used to dynamically identify the extreme value region in the data distribution.

[0106] Continuous value features directly participate in the linear regression of the two-stage model, and classification indicators optimize the learning ability of boundary regions by enhancing the sensitivity of the model to extreme samples. In the two-stage transmission path, the first stage takes monomer ratio, process parameters and range-aware features as input to predict F-zone and U-zone features, establishing a "ratio-chemical structure" physical transmission; the second stage fuses NMR features and range-aware features to predict performance indicators such as mv_121 (Mooney ML(1+10)121), ts (tensile strength), pc, elongation (elongation) through four parallel strategies (pure NMR features, monomer+NMR fusion, PLS dimensionality reduction, targeted strategy), among which the PLS dimensionality reduction path introduces a range indicator to strengthen the physical meaning of principal components.

[0107] Preferably, the range-aware mechanism improves model robustness through triple optimization: extreme value enhancement replicates high / low performance samples (±0.8σ outside) through a resampling strategy, combining with range indicator features to strengthen feature learning in extreme value regions; feature selection optimization adopts a comprehensive scoring system of "range score + absolute value of correlation", prioritizing features with strong discrimination ability in different performance ranges; target transformation improves the linearity of the target variable through Yeo-Johnson Power transformation, and reduces the influence of abnormal points through Huber regression. The synergistic application of the three types of features realizes a complete modeling closed loop from chemical structure knowledge embedding (F region), unknown pattern discovery (U region), to data distribution awareness (range features).

[0108] Further, known functional group features: based on literature and expert experience knowledge and nuclear magnetic resonance spectrum experiments, 9 key chemical shift intervals (such as PMVE-OCF3: -52.9ppm to -55.3ppm, TFE main chain: -125.7ppm to -128.4ppm, etc.) and 33 feature peak positions are identified, and the relative peak area (reflecting the functional group content) and signal intensity (reflecting the local structure signal) are extracted. To ensure data quality, invalid features with peak area less than 0.5% of the full spectrum or low signal intensity are removed.

[0109] Unknown structure fingerprint: identify the part not covered by known features in the high signal region through dynamic DBSCAN clustering, remove intervals with peak area less than 0.1% of the full spectrum and intervals with less than 5 data points, and finally extract 15 unknown interval features (U1-U15) as "structure fingerprints" with unclear attribution.

[0110] Range-aware features: continuous and categorical indicators based on key functional group ratios (such as OCF3 / CF2-CF2 ratio and its high / low threshold identification) to enhance model sensitivity to extreme value samples.

[0111] Specifically, a coding feature matrix is constructed. The known region peak area (F1, F2...), known point feature signal intensity (F1*, F2*...), and unknown region peak area (U1, U2...) and range-aware features (R1, R2...) are filtered to construct a unified feature matrix X, which can be represented as:

[0112]

[0113] Where r includes r1, r2, r3, the number of known area features is 6, the number of known point feature intensity features is 32, the number of unknown features is 15, and the number of range-aware features is 9.

[0114] Preferably, StandardScaler normalization is applied to all numerical features before model training, transforming features to standard normal distribution with mean 0 and standard deviation 1. The formula is as follows:

[0115]

[0116] wherein is the normalized data, μ is the feature mean, and σ is the standard deviation.

[0117] Specifically, a two-stage modeling implementation is used to build the base model. Algorithm integration: to reduce model complexity, the same core modeling components are used in both stages to ensure consistency and comparability of the method. This embodiment combines five algorithms that have performed well in various fields, including materials science (Ridge regression, ElasticNet regression, Huber regression, support vector regression, and gradient boosting regression), and uses a grid search algorithm to optimize model parameters and improve model interpretability.

[0118] Optionally, parameter optimization. The Ridge regression regularization parameter a is grid searched in the range [0.1, 1, 10, 100], the ElasticNet regression a parameter is selected in the range [0.1, 1, 10] and l1_ratio is optimized in the range [0.1, 0.5, 0.7, 0.9], the support vector regression uses an RBF kernel function and the C parameter is adjusted in the range [0.1, 1, 10], and the gradient boosting regression uses 50 estimators, a learning rate of 0.1, and a maximum depth of 3. The optimal algorithm configuration is automatically selected for each modeling task through cross-validation. l1_ratio is the ratio that controls the "L1 regularization (makes unimportant feature coefficients zero, achieving feature selection)" and "L2 regularization (makes coefficients small, avoiding overfitting)", and the adjustment range includes [0.1, 0.5, 0.7, 0.9] to find the most balanced ratio.

[0119] Preferably, range-aware technology (RS). This method groups feature values by quartiles, calculates the ratio of global standard deviation to group standard deviation as a range score, and combines a 70% range score and a 30% statistical correlation to balance the distinguishing ability and linear correlation strength of the feature. Effectively addresses the modeling challenge of low variability of fluorine rubber performance data. In the first stage, it is used to evaluate the regulatory ability of monomers on structural features, and in the second stage, it is used to select structural features that are most valuable for performance prediction.

[0120] Reference Figure 3, the model network structure takes two-stage conduction as the core, including the input layer, the data separation layer, the enhanced feature engineering layer, the range-aware feature selection layer, the first-stage encoder (monomer ratio → NMR feature, integrated regression model), the intermediate representation layer, the second-stage decoder (NMR feature → performance index, containing multi-strategy parallel processing, target transformation, extreme value resampling), the enhanced PLS dimensionality reduction branch, the strategy selector and the output layer. In the training process, the original multivariate chemical data is first grouped by the data separation layer according to the chemical meaning, and then the basic ratio maintenance, interactive proportion calculation, functional group aggregation and correlation screening are performed through the feature engineering layer, and then the key features are screened by the range-aware selection layer by fusing the Spearman correlation (0.3 weight) and the range differentiation ability score (0.7 weight, based on the inter-group / global standard deviation ratio of quartile grouping); the first stage establishes the mapping from monomer ratio to NMR feature through the integrated regression model (such as GradientBoosting), and after leave-one-out cross-validation, the predicted NMR feature of the intermediate layer is taken as the input into the second stage, which combines target variable transformation (Yeo-Johnson) and extreme value sample resampling (1.5 times repeated) through four parallel paths (pure NMR, monomer + NMR fusion, PLS dimensionality reduction, customized features), and the strategy selector maximizes the automatic matching of the optimal strategy based on the leave-one-out cross-validation R 2 to generate the performance index prediction distribution. The data flow runs through the whole process of "chemical chunking → knowledge-driven feature expansion → range-aware screening → two-stage physical conduction → multi-strategy adaptive fusion", and each layer realizes the transmission through structured data dictionary, feature vector and predicted value. The core innovative modules (range-aware selection, two-stage conduction, extreme value enhancement) are embedded in the key nodes to improve the prediction accuracy and interpretability of narrow-range data.

[0121] Specifically, the first stage modeling: monomer ratio → NMR feature mapping. The first stage establishes the mapping relationship from 3 monomer ratios to 53 NMR structural features. This stage adopts a single full-feature input strategy, and for each NMR feature, an independent prediction model is constructed, focusing on establishing a stable and reliable ratio-structure basic mapping, without using target transformation, extreme value resampling and other enhancement techniques.

[0122] Further, the second stage modeling: NMR feature → performance mapping. The second stage constructs a performance prediction model based on the 53 NMR features predicted by the first stage, and the output is four key performance indicators (mv_121, ts, pc, elongation). This stage applies target transformation, extreme value resampling and dimensionality reduction enhancement techniques.

[0123] Preferably, four differentiated modeling strategies are adopted: a chemical rule constraint strategy selects functional group features based on domain knowledge, a statistical significance strategy filters relevant features by p-value, a multivariate prediction strategy handles feature redundancy by PLS regression, and a targeted strategy customizes feature combinations according to performance characteristics. Each strategy is independently trained, and the optimal configuration is selected as the prediction best model through cross-validation. The two stages are connected through a strict data transmission process, and the 57 NMR feature prediction values of the first stage are directly used as input features of the second stage. A complete conduction path from monomer ratio to microstructure to macroscopic performance is constructed.

[0124] Further, the training and evaluation strategy. Both stages use leave-one-out cross-validation (LOOCV) to ensure the consistency and reliability of the evaluation. Each iteration uses 51 samples for training and 1 sample for testing, and the cycle is repeated 52 times to obtain an unbiased performance estimate. The evaluation index uniformly uses the coefficient of determination R 2 , root mean square error RMSE and RMSE percentage, range sensitivity (quantile verification) and interpretability (conduction path analysis) three-dimensional index verification method effectiveness, and based on R 2 Select the best model configuration.

[0125] Specifically, group comparison experiment. To verify the effectiveness of the two-stage modeling method, three groups of comparative experiments are designed. The first group is the benchmark comparison, including direct modeling B1 (monomer ratio directly predicts performance without using RS indicators), traditional two-stage model B2 (two-stage modeling without using RS indicators), and range-aware two-stage model RATS of the present embodiment. (For the low variability characteristics of fluoroelastomer performance data) The second group is the quantile stratification verification, which divides the performance data into low (below Q1), medium (Q1-Q3), and high (above Q3) three intervals according to the quantile. Compare the prediction accuracy of each model in different performance intervals, and focus on evaluating the improvement effect of the range-aware technology in the extreme value interval. The third group defines the conduction mechanism verification formula:

[0126]

[0127] In the formula, is the indicator function, ρ(·,·) is the Pearson correlation coefficient, Q q (·) is the quantile function, X represents the monomer ratio, Z is the NMR feature, and Y is the performance index.

[0128] The conduction relationship between monomer ratio, NMR characteristics, and performance indicators is defined, and the verification formula of the conduction mechanism is defined. Among them, DE is the Pearson correlation coefficient of monomer ratio (X) and performance indicator (Y), representing the direct effect of ratio on performance; IE is the product of the correlation coefficient of ratio and NMR characteristics (Z) and the correlation coefficient of NMR characteristics and performance indicator, reflecting the indirect effect mediated by microstructure. CV is an indicator function, which takes 1 when the absolute value of indirect effect is greater than that of direct effect, indicating that the mediation of microstructure dominates, otherwise it takes 0. CS uses quantile function to measure the strength of indirect effect in different data intervals (such as performance extreme interval), which is used to evaluate the improvement effect of range-aware technology in extreme interval, so as to quantify and verify the conduction logic of "formulation adjustment → change microstructure → affect performance" and the difference in different intervals.

[0129] Further, results and discussion, two-stage modeling performance evaluation. The first stage: ratio → NMR characteristics mapping: the monomer ratio → NMR characteristics prediction model based on range-aware method has good performance. Table 1 shows an example of the performance indicators of various NMR characteristics. 2 (R-squared), RMSE (root mean square error), RMSE% (percentage root mean square error) are commonly used performance evaluation indicators in machine learning and regression modeling. 2 Used to measure the degree of model explanation for data, the value range is between 0 and 1, the closer to 1 indicates that the model can explain more part of the dependent variable variation, and the fitting effect is better; RMSE reflects the average error size of predicted value and true value by calculating the square root of the average value of the square sum of the difference between predicted value and true value, the smaller the value, the higher the model prediction accuracy; RMSE% is the normalized processing of RMSE relative to the average value of true value, which presents the error in percentage form, which is convenient for comparing model error level between different dimension or order of magnitude data sets, the lower the percentage, the better the model prediction performance, and the three jointly evaluate the fitting and prediction ability of the model for data from different dimensions.

[0130] Table 1

[0131]

[0132] Among them, the prefix F represents the known functional group characteristics, the suffix _area represents the known functional group area characteristics, and the _intensity represents the signal intensity characteristics of the known chemical shift functional group; the prefix U represents the unknown fingerprint characteristics, and the suffix _area is the area characteristics. The results show that high-precision prediction is achieved even for unknown structure fingerprints, proving that there is a stable and predictable relationship between monomer ratio and microstructure.

[0133] Further, the second stage: NMR feature→performance mapping, the prediction accuracy of the range-aware two-stage modeling and the range-aware direct modeling (directly predicting performance from monomer ratio) was compared at the beginning of modeling, and the results are shown in Table 2.

[0134] Table 2

[0135]

[0136] Two-stage modeling average prediction accuracy R 2 up to 0.91, 0.15 higher than direct modeling, and the average prediction error is reduced by 28.0%.

[0137] Reference Figure 4 , the left side is the narrow distribution feature analysis (CV and concentration degree annotation) of the performance index of fluororubber, and the right side is the comparison of the coefficient of variation (CV) of the performance index of fluororubber and the proportion of concentrated distribution. In this embodiment, all performance data of fluororubber (including mv_121, ts, pc, elongation, etc.) are systematically statistically analyzed and micro-mechanism correlated. Through quantifying the coefficient of variation (CV value is between 0.104 and 0.229) and the proportion of concentrated distribution (75.0% to 82.7% of the data is concentrated in the mean±1σ range), it is found that all performance indicators show significant low variability, which makes the traditional feature selection method ineffective. Using our range-aware method, it is found that the range score successfully identifies some key features that are missed by traditional methods (Pearson correlation analysis) Reference Table 3.

[0138] Table 3

[0139]

[0140]

[0141] The results prove the effectiveness of RS in identifying key features in low variability data, for example, the Pearson correlation coefficient of F7_area is only 0.01 (corr_rank=47), but it ranks first in range-awareness, and also has the highest contribution to the target feature. This difference directly verifies the advantage of the range-aware method (RATS) in capturing key features in low variability data, making up for the shortcomings of traditional linear correlation analysis.

[0142] Specifically, model overall evaluation method, benchmark comparison evaluation. Table 4 summarizes the performance comparison of range-aware two-stage modeling and benchmark methods. The RATS method of this embodiment is significantly better than direct modeling (B1) and traditional two-stage modeling (B2) in all performance indicators, with an average R 2The result shows that the method has greater advantages in solving such low variability data.

[0143] Table 4

[0144]

[0145]

[0146] Further, the quantile stratified performance analysis, Figure 5 The prediction accuracy of the three methods in different performance quantile intervals is shown. RATS performs particularly well in the extreme value interval. The error of the RATS method of the embodiment in the extreme value interval (<Q1 before 25%, >Q3 after 25%) (such as 0.81, 0.67 of mv_121) is significantly better than B1 (0.26, 0.06) and B2 (0.25, 0.06). Figure 5 It is shown that the relative error of RATS in the first 25% (<Q1) is 4.59% (taking mv_121 as an example), which is improved by 49.8% compared with direct modeling (B1) and 49.2% compared with traditional two-stage modeling (B2), and the same is true for the last 25% (>Q3). This improvement plays a positive role in material formula optimization, because the extreme performance (the first 25% determines the lower limit, and the last 25% >Q3 determines the upper limit) often determines the application boundary of the material.

[0147] As an optional implementation, the embodiment also provides an electronic device, comprising: at least one processor, and a memory connected in communication with the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to execute the aforementioned range-aware two-stage modeling method for fluororubber formula regulation.

[0148] As an optional implementation, the embodiment also provides a non-transitory computer readable storage medium storing computer instructions, the computer instructions being used to enable a computer to execute the aforementioned range-aware two-stage modeling method for fluororubber formula regulation.

[0149] The beneficial effects of the present application are as follows:

[0150] The present application decomposes the complex monomer ratio and macroscopic performance mapping problem into two relatively simple and physically meaningful sub-problems, and simultaneously introduces range-aware technology to effectively capture nonlinear relationships and improve the feature mining effect of low variability data; by integrating known functional group characteristics and unknown structure fingerprints, a more complete microstructure characterization system is constructed.

[0151] The various embodiments described in this specification are intended to be exemplary only. The various embodiments were chosen and described in order to best explain the principles of the application and its practical application, to thereby enable others skilled in the art to best utilize the application, and to best enable the present application to be performed with determination by those skilled in the art.

[0152] The principles and implementations of the present application have been described in the specification with specific examples. The above description of the embodiments is only for the purpose of understanding the method of the present application and its core idea; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation and application range will be changed. In view of the above, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A range-aware two-stage modeling method for fluororubber formulation control, characterized in that: include: Build the model network structure; The model network structure includes: an input layer, a data separation layer, an enhanced feature engineering layer, a range-aware feature selection layer, a first-stage encoder, an intermediate representation layer, an enhanced PLS dimensionality reduction branch, a second-stage decoder, and an output layer connected in sequence; the second-stage decoder includes: a strategy selector; receiving raw multivariate chemical data via the input layer; Using the data separation layer to perform grouping processing on the original multiple chemical data to obtain grouped data; Using the enhanced feature engineering layer to perform basic ratio maintenance, interaction ratio calculation, functional group aggregation, and correlation screening on the grouped data to obtain enhanced feature data; The enhanced feature data is subjected to feature screening using the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data; the expression of the comprehensive evaluation formula is: CS(X,y)=0.7×RS(X,y)+0.3×|Corr(X,y)|; wherein, CS(X,y) is the comprehensive score; RS(X,y) is the range score; Corr(X,y) is the absolute value of the correlation; σ global (y) is the global standard deviation of the target variable; Represents σ global (y) performs square operation; σ group (y) is the standard deviation of the target variable y within the i-th quartile group of the feature matrix X; k is the number of quartile groups; In the first-stage encoder, an integrated regression model is used to construct a mapping from monomer ratios to the key feature data and a leave-one-out cross-validation is performed to obtain predicted feature data, and the predicted feature data is transferred to the second-stage decoder using the intermediate representation layer; Performing dimensionality reduction and range indication generation on the predicted feature data using the enhanced PLS dimensionality reduction branch to obtain dimensionality reduction indication data, and passing the dimensionality reduction indication data to the second-stage decoder; According to the dimensionality reduction indication data, the predicted feature data is processed in parallel using the second-stage decoder, the target variable is transformed, and the extreme value samples are resampled to obtain performance indicator data, and the performance indicator data is subjected to leave-one-out cross-validation maximization strategy matching using the strategy selector to obtain the indicator prediction distribution; the parallel processing paths include: full NMR, monomer and NMR fusion, PLS dimensionality reduction, and customized features; Outputting the indicator prediction distribution using the output layer; The constructed model network structure is trained according to the indicator prediction distribution to obtain a fluororubber formula control model.

2. A range-aware two-stage modeling method for regulating fluororubber formulation according to claim 1, characterized in that: The input layer receives raw multi-dimensional chemical data, including: Performing Savitzky-Golay filtering to reduce noise on the original multivariate chemical data; performing baseline correction on the raw multivariate chemical data after noise reduction; The corrected raw multivariate chemical data were intensity normalized.

3. A range-aware two-stage modeling method for regulating fluororubber formulation according to claim 1, characterized in that: The data separation layer is used to perform grouping processing on the original multi-dimensional chemical data to obtain grouped data, including: Extracting functional group characteristic data from the original multi-chemical data according to a preset chemical structure; the functional group characteristic data includes: OCF2, CH2-CF2, CF2-CF2; Using DBSCAN clustering to perform interval feature recognition on the original multivariate chemical data to obtain unknown structural fingerprints; The functional group characteristic data and the unknown structure fingerprint are integrated to obtain the grouping data.

4. A range-aware two-stage modeling method for regulating fluororubber formulation according to claim 1, characterized in that: The enhanced feature engineering layer is used to perform basic ratio maintenance, interaction ratio calculation, functional group aggregation, and correlation screening on the grouped data to obtain enhanced feature data, including: Extracting regional peak areas, known point feature signal intensities, unknown region peak areas, and range perception features of the grouped data after correlation screening; The feature matrix is ​​constructed using the known region peak area, the known point characteristic signal intensity, the unknown region peak area, and the range perception feature; the expression of the feature matrix is: Wherein, r is a monomer ratio set; r1 is the first ratio in the monomer ratio set; F is a set of the peak area of ​​the known region and the characteristic signal intensity of the known point; U is the peak area of ​​the unknown region; R is the range perception feature; F1 is the first parameter in the peak area of ​​the known region; F 1* is the first parameter in the known point characteristic signal intensity; U1 is the first parameter in the unknown region peak area; R1 is the first parameter in the range perception feature; n is the total number of samples; The feature matrix is ​​standardized using a StandardScaler standardization formula to obtain the enhanced feature data.

5. The range-aware two-stage modeling method for regulating fluororubber formulation according to claim 1, characterized in that: In the first-stage encoder, an integrated regression model is used to construct a mapping from monomer ratio to the key feature data and a leave-one-out cross-validation is performed to obtain predicted feature data, and the predicted feature data is transferred to the second-stage decoder using the intermediate representation layer, including: The integrated regression model is pre-set; the integrated regression model includes: Ridge regression model, ElasticNet regression model, Huber regression model, support vector regression model and gradient boosting regression model; The parameter optimization range of the integrated regression model is set; the parameter optimization range of the regularization parameter of the Ridge regression model is [0.1, 1, 10, 100]; the parameter optimization range of the alpha parameter of the ElasticNet regression model is [0.1, 1, 10]; the parameter optimization range of the L1 scale parameter of the ElasticNet regression model is [0.1, 0.5, 0.7, 0.9]; the parameter optimization range of the complexity parameter of the support vector regression model is [0.1, 1, 10]; the learning rate of the gradient boosting regression model is 0.1 and the maximum depth is 3.

6. A range-aware two-stage modeling method for regulating fluororubber formulation according to claim 1, characterized in that: The enhanced feature data is subjected to feature screening using the range-aware feature selection layer according to a comprehensive evaluation formula to obtain key feature data, including: Calculating the enhanced feature data using a comprehensive evaluation formula to obtain the comprehensive score; Sort the comprehensive scores in descending order, select a set number of features with the highest scores, and obtain the key feature data; When the number of the comprehensive scores in the same group is less than the set number, the top three features with the highest scores are selected to obtain the key feature data.

7. A range-aware two-stage modeling method for regulating fluororubber formulation according to claim 3, characterized in that: Extracting functional group characteristic data from the original multi-chemical data according to a preset chemical structure, including: Identifying 9 key chemical shift intervals and 33 characteristic peak positions in the original multivariate chemical data to obtain original extracted data; the key chemical shift intervals include: -52.9 ppm to -55.3 ppm, -125.7 ppm to -128.4 ppm; The features whose peak area accounts for less than 0.5% of the full spectrum or whose signal intensity is lower than a preset signal threshold in the original extracted data are eliminated to obtain the original extracted data.

8. An electronic device, characterized in that: include: At least one processor and a memory communicatively connected to the processor; wherein the memory stores instructions executable by the processor, and the instructions are executed by the processor to enable the processor to perform a range-aware two-stage modeling method for regulating a fluororubber formulation as claimed in any one of claims 1 to 7.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable a computer to execute a range-aware two-stage modeling method for regulating fluororubber formulation according to any one of claims 1 to 7.