Method and system for predicting solubility of sulfur dioxide in deep eutectic solvent based on KAN
Through the solubility prediction method based on knowledge-enhanced network, the limitations of the existing model in the solubility prediction of sulfur dioxide are solved, high-precision prediction in deep eutectic solvents are achieved, solvent design and optimization are guided, and sulfur dioxide capture efficiency is improved.
Patent Information
- Application Number
- CN202510629954.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
AI Technical Summary
The existing solubility prediction model of sulfur dioxide in deep eutectic solvents has a fixed temperature data set, which cannot consider the dynamic relationship between temperature and hydrogen bond strength, which is expensive to calculate, and the machine learning method has poor generalization ability outside the training data range, making it difficult to guide the design and optimization of deep eutectic solvents.
The solubility prediction method based on knowledge enhancement network (KAN) is adopted, and the solubility of sulfur dioxide in deep eutectic solvents is predicted through the input layer, the knowledge enhancement layer, the attention mechanism layer and the depth feature processing layer, combined with the residual connection processing, and the solubility of sulfur dioxide in the deep eutectic solvent is predicted. The chemical structure and process condition data are used for feature extraction and dynamic weighting to improve the prediction accuracy.
Maintaining high prediction accuracy over a wide range of temperature and pressures reduces gradient disappearance problems, provides accurate prediction of solubility, guides solvent design and optimization, reduces water consumption, and improves sulfur dioxide capture efficiency.
Smart Images

Figure CN120496670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of chemical engineering and deep learning technology, and in particular to a method and system for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN. Background Art
[0002] Traditional sulfur dioxide emission control technologies, such as wet flue gas desulfurization (FGD), can achieve 90-95% removal efficiencies under ideal conditions. However, under non-ideal conditions such as low pressure or temperature fluctuations, absorption efficiency can drop below 85%. Furthermore, traditional FGD systems produce waste lye containing sulfite and sulfate sludge, causing secondary pollution. Traditional FGD systems also require significant water resources, typically requiring 2.5 to 4 cubic meters of water per ton of sulfur dioxide captured, making them difficult to implement in water-scarce regions.
[0003] In recent years, deep eutectic solvents, as a new type of solvent, have shown great potential in the field of sulfur dioxide capture due to their adjustable physicochemical properties, low volatility, and biodegradability. Deep eutectic solvents are composed of hydrogen bond acceptors (HBAs) and hydrogen bond donors (HBDs), forming a stable hydrogen bond network. The solubility of sulfur dioxide in deep eutectic solvents is affected by many factors, including Lewis acid-base interactions, dipole-induced dipole interactions, and physical absorption enhanced by low vapor pressure. Functionalized deep eutectic solvents, especially those containing thiourea derivatives, can achieve a SO2 / N2 selectivity ratio of up to 1200 at a pressure of 0.1 bar, which is far superior to traditional amine-based solvents.
[0004] However, due to the large number of possible HBA-HBD combinations (over 100,000), the design and optimization of deep eutectic solvents face enormous challenges, and predictive modeling methods are urgently needed to guide solvent selection and optimization. Existing predictive models have multiple limitations: traditional quantitative structure-property relationship (QSPR) models rely on fixed temperature datasets and cannot consider the dynamic relationship between temperature and hydrogen bond strength; while molecular dynamics simulations can help understand solvent behavior at the molecular level, they are computationally expensive and unsuitable for large-scale screening; and existing machine learning methods, such as multilayer perceptrons (MLPs), have poor generalization capabilities beyond the training data range, especially for predicting the solubility of chloride-based deep eutectic solvents.
[0005] Therefore, developing an efficient model that can accurately predict the solubility of sulfur dioxide in deep eutectic solvents has important scientific significance and application value for guiding the design and optimization of deep eutectic solvents and improving the capture efficiency of sulfur dioxide. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN, which improves the accuracy of solubility prediction.
[0007] The purpose of the present invention can be achieved by the following technical solutions:
[0008] A method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN comprises the following steps:
[0009] A dataset of sulfur dioxide dissolved in a deep eutectic solvent is obtained, preprocessed, and input into a pre-trained KAN-based solubility prediction model to output a solubility prediction result. The KAN-based solubility prediction model includes an input layer, a knowledge enhancement layer, an attention mechanism layer, a deep feature processing layer, and an output layer. The execution process of the KAN-based solubility prediction model includes:
[0010] Inputting the preprocessed data set into the knowledge enhancement layer through the input layer for feature extraction;
[0011] The attention mechanism layer is used to dynamically weight and fuse the extracted features to obtain fused features;
[0012] A deep feature processing layer is used to perform deep processing on the fused features, and residual connection processing is performed in combination with the fused features. Finally, the solubility of sulfur dioxide in the deep eutectic solvent is predicted through the output layer.
[0013] Furthermore, the data set includes chemical structure data and process condition data, the chemical structure data includes hydrogen bond acceptors and hydrogen bond donors, and the process condition data includes deep eutectic solvent ratio, temperature, pressure and water content.
[0014] Furthermore, the preprocessing operation includes outlier detection and data normalization, wherein the Z-Score method is used for outlier detection, and data points with an absolute Z-Score value greater than a preset value are identified as outliers and removed from the data set;
[0015] The data normalization process converts the data into a standard normal distribution with a mean of 0 and a variance of 1, or scales the data to the interval [0, 1].
[0016] Furthermore, the input layer and the knowledge enhancement layer are sequentially connected with a first dense block containing a ReLU activation function, a batch normalization layer, and a dropout layer.
[0017] Furthermore, the knowledge enhancement layer includes two parallel branches, the first branch is a second dense block containing a TanH activation function, which is processed using the TanH activation function to obtain chemical structure features, and the second branch is a third dense block containing a Sigmoid activation function, which is processed using the Sigmoid activation function to obtain process condition features.
[0018] Furthermore, the execution steps of the attention mechanism layer include:
[0019] The Softmax activation function in the fourth dense block is used to generate normalized weight coefficients of different features;
[0020] Performing element-wise multiplication of the normalized weight coefficients with the corresponding output features of different branches to dynamically weight different features;
[0021] The dynamically weighted features are spliced to obtain fused features.
[0022] Furthermore, the deep feature processing layer includes a fifth dense block including a ReLU activation function, a batch normalization layer, a dropout layer, a sixth dense block including a ReLU activation function, and a seventh dense block respectively connected to the attention mechanism layer and the sixth dense block residual.
[0023] Furthermore, during the training of the KAN-based solubility prediction model, the loss function used is:
[0024] L = 0.7*MSE+0.3*MAE
[0025] Where L is the loss, MSE is the mean square error, and MAE is the mean absolute error.
[0026] Furthermore, the coefficient of determination R 2 , mean square error (MSE), root mean square error (RMSE), average absolute relative deviation (AARD) and mean absolute error (MAE) were used to evaluate the performance indicators of the KAN-based solubility prediction model.
[0027] The present invention also provides a KAN-based prediction system for the solubility of sulfur dioxide in a deep eutectic solvent, comprising:
[0028] Data preprocessing module: used to obtain the data set after sulfur dioxide is dissolved in a deep eutectic solvent and perform preprocessing;
[0029] The knowledge-enhanced network model module is used to input the preprocessed data set into the pre-trained KAN-based solubility prediction model and output the solubility prediction results. The KAN-based solubility prediction model includes an input layer, a knowledge-enhanced layer, an attention mechanism layer, a deep feature processing layer, and an output layer. The execution process of the KAN-based solubility prediction model includes:
[0030] Inputting the preprocessed data set into the knowledge enhancement layer through the input layer for feature extraction;
[0031] The attention mechanism layer is used to dynamically weight and fuse the extracted features to obtain fused features;
[0032] A deep feature processing layer is used to perform deep processing on the fused features, and residual connection processing is performed in combination with the fused features. Finally, the solubility of sulfur dioxide in the deep eutectic solvent is predicted through the output layer.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] (1) The present invention uses the knowledge enhancement network KAN to perform feature extraction enhancement and feature depth processing on the dataset after sulfur dioxide is dissolved in a deep eutectic solvent, which helps to effectively capture the nonlinear relationship in the high-dimensional feature set, thereby improving the accuracy of solubility prediction.
[0035] (2) The knowledge enhancement layer of the present invention uses two parallel branches to extract chemical structure data and process condition data in parallel. The attention mechanism layer dynamically weights different features and performs feature splicing to form a knowledge-enhanced feature representation, which can effectively process the features of different HBA-HBD combinations and maintain high prediction accuracy within a wide temperature range (293.0-353.5K) and pressure range (0.2-127.3kPa).
[0036] (3) The present invention adopts residual connections in the deep feature processing layer, which creates a shortcut path and allows the gradient to flow directly through the layer, bypassing the traditional conversion sequence and alleviating the gradient vanishing problem in the deep network.
[0037] (4) The present invention has the functions of single prediction, variable range prediction, HBA / HBD type prediction, etc., which can quantitatively explain the effects of pressure, temperature and deep eutectic solvent composition on the solubility of sulfur dioxide and provide mechanistic guidance for solvent design. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 Schematic diagram of the method flow of the present invention;
[0039] Figure 2 This is a schematic diagram of the knowledge augmentation network (KAN) architecture of the present invention;
[0040] Figure 3 It is the Z-Score residual error scatter plot of the present invention;
[0041] Figure 4 This is a comparison chart of RMSE values of the effects of different data segmentation ratios on model performance in the present invention;
[0042] Figure 5 The figure shows the cross-plot of the prediction results of the KAN framework of the present invention and the experimental SO2 data, where (a) is the cross-plot of the prediction results on the training set and the experimental SO2 data, and (b) is the cross-plot of the prediction results on the test set and the experimental SO2 data;
[0043] Figure 6 The dynamic graph of the KAN framework learning of the present invention, where (a) is the training-test MAE trajectory; (b) is the training-test MSE trajectory;
[0044] Figure 7 This is a schematic diagram of the SO2 absorption prediction application interface of the present invention;
[0045] Figure 8 Schematic diagram of the visualization module of the SO2 absorption prediction application of the present invention. DETAILED DESCRIPTION
[0046] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented based on the technical solution of the present invention, and provides a detailed implementation method and specific operation process, but the protection scope of the present invention is not limited to the following embodiments.
[0047] Example 1
[0048] This embodiment provides a method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN. Figure 1 As shown, the method includes the following steps:
[0049] S1. Construct an experimental dataset containing hydrogen bond acceptors (HBAs), hydrogen bond donors (HBDs), deep eutectic solvent ratio (DES_ratio), temperature, pressure and water content.
[0050] Based on published research, this example constructs a dataset of 1,357 experimental data points for the absorption of sulfur dioxide by deep eutectic solvents. This dataset includes 21 HBAs and 42 HBDs, covering the ratio of hydrogen bond donors to acceptors, temperature (T), pressure (P), water content (HO), and the mole fraction of dissolved sulfur dioxide (i.e., sulfur dioxide solubility).
[0051] The statistical characteristics of the dataset are as follows:
[0052] Temperature range: 293.0K to 353.5K, average value: 308.6056K, standard deviation: 16.1981K;
[0053] Pressure range: 0.2kPa to 127.3kPa, average value: 58.6745kPa, standard deviation: 38.8354kPa;
[0054] Water content range: 0% to 20%, average value is 0.2784%, standard deviation is 0.8136%;
[0055] The solubility range of sulfur dioxide is 0.003 g / g to 1.510 g / g, with an average value of 0.5048 g / g and a standard deviation of 0.3466 g / g.
[0056] The dataset is preprocessed, including outlier detection and data normalization. Outlier detection uses the Z-Score method, and data points with an absolute Z-Score value greater than 3 are identified as outliers. Figure 3 As shown, out of a total of 1,357 data points, 45 outliers were identified (circles mark the outliers removed from the training dataset), accounting for approximately 2% of the data. These outliers were removed from the training set to improve the accuracy and reliability of model training.
[0057] There are two approaches to data normalization:
[0058] 1. Standardization: Convert the data to a standard normal distribution with a mean of 0 and a variance of 1 to ensure that each feature contributes equally to the learning process;
[0059] 2. Normalization: Rescale the data to the range [0,1] to accommodate features of different scales.
[0060] In addition, this embodiment also introduces a batch normalization layer into the knowledge augmentation network (KAN) framework to solve the problem of internal covariate shift, accelerate the training process, stabilize the optimization process, and improve the generalization performance of the model.
[0061] S2. Build a knowledge-enhanced network deep learning framework.
[0062] This example constructs a knowledge augmentation network (KAN) deep learning framework for predicting the solubility of sulfur dioxide in deep eutectic solvents. The core architecture of the framework is as follows: Figure 2 As shown, it includes the following components:
[0063] 1. Input layer: Receives pre-processed feature data, including the chemical structure characteristics of HBAs and HBDs, and process condition characteristics such as deep eutectic solvent ratio, temperature, pressure, and water content;
[0064] 2. Knowledge Enhancement Layer: It has two parallel feature extraction branches to decompose the complex input feature stream into multiple sub-channels. Each sub-channel is processed by an independent fully connected layer and nonlinear activation functions (such as ReLU and TanH), effectively capturing the nonlinear relationships in the high-dimensional feature set. Specifically:
[0065] The first branch: using the TanH activation function in the second dense block to process the chemical structure features of HBA and HBD;
[0066] The second branch uses the Sigmoid activation function in the third dense block to process process condition features such as deep eutectic solvent ratio, temperature, pressure and water content;
[0067] 3. Attention mechanism layer: Softmax activation function is used to calculate the weight coefficients of different features and dynamically weight the contributions of different deep eutectic solvent components. Specifically:
[0068] First, the input features are processed by a first dense block with 128 neurons and a ReLU activation function. After batch normalization and dropout layer processing, the normalized weight coefficients are generated using the Softmax activation function in the fourth dense block.
[0069] Then, the generated weight coefficients are element-wise multiplied with the outputs of different feature extraction branches to achieve dynamic weighting of different features.
[0070] Finally, the weighted features are fused through the concatenation layer to form a knowledge-enhanced feature representation, that is, the fused features.
[0071] 4. Deep Feature Processing Layer: This layer uses residual connections to create shortcuts, allowing gradients to flow directly through the layer, bypassing the traditional transformation sequence and alleviating the vanishing gradient problem in deep networks. Specifically, the deep feature processing layer consists of a fifth dense block with a ReLU activation function, a batch normalization layer, a dropout layer, a sixth dense block with a ReLU activation function, and a seventh dense block with residual connections to the attention mechanism layer and the sixth dense block.
[0072] 5. Output layer: Map the high-dimensional fused feature representation to the output to predict the solubility of sulfur dioxide in deep eutectic solvents.
[0073] In this embodiment, the KAN framework uses a combination of the following activation functions:
[0074] ReLU (rectified linear unit): outputs zero for positive inputs and zero for negative inputs. It is widely used because of its simplicity and effectiveness, and can alleviate saturation problems.
[0075] TanH (Hyperbolic Tangent): Maps the input to the range [-1, 1], effectively capturing positive and negative input values;
[0076] Softmax: used for multi-class classification in the output layer, converting the original output into a probability distribution of multiple categories;
[0077] Sigmoid: Maps the input to the range [0,1] and is suitable for binary classification tasks.
[0078] In this embodiment, the KAN framework is trained using a custom loss function that combines the mean squared error (MSE) and the mean absolute error (MAE), with an MSE weight of 0.7 and a MAE weight of 0.3, as defined below:
[0079] L = 0.7*MSE+0.3*MAE
[0080] Where L is the loss, MSE is the mean square error, and MAE is the mean absolute error.
[0081] S3. Training and evaluation of knowledge-enhanced networks.
[0082] This embodiment trains and evaluates the knowledge enhancement network. First, the data set is divided into a test set and a training set according to the ratio of 10 / 90, 20 / 80 and 30 / 70, as shown in the following example: Figure 4 As shown, the 20 / 80 ratio achieves the best generalization ability while maintaining reliable predictions.
[0083] The following high-level callback functions are used during training:
[0084] Early Stopping: Monitor validation loss, set patience parameter to 75, and restore optimal weights.
[0085] Learning rate reduction (ReduceLROnPlateau): monitoring validation loss, factor 0.5, patience parameter 30, minimum learning rate 1e-6;
[0086] Model Checkpoint: Save the best performing model;
[0087] TensorBoard: for advanced visualization.
[0088] Models are evaluated using five widely accepted performance metrics:
[0089] 1. Coefficient of determination (R 2): represents the proportion of dependent variable variance explained by the model. The closer it is to 1, the better the model fit.
[0090] 2. Mean Squared Error (MSE): represents the average squared difference between the predicted value and the actual value;
[0091] 3. Root mean square error (RMSE): The square root of the MSE, providing an error metric in the same units as the original data;
[0092] 4. Average absolute relative deviation (AARD): the average relative deviation expressed as a percentage of the actual value;
[0093] 5. Mean Absolute Error (MAE): The average absolute difference between the predicted value and the actual value.
[0094] The evaluation results show that the KAN framework achieves an R of 0.9963 on the test set. 2 value and MSE of 0.0005, which is better than existing machine learning methods such as Figure 5 shown. Figure 5 Figure (a) shows the fitting accuracy of the model on the training set, R 2 is 0.9959, and RMSE is 0.0220; Figure 5 Figure (b) in the figure verifies its generalization ability on the test set, R 2 The data points in both figures are closely clustered on the diagonal line, indicating that the systematic deviation between the predicted and experimental values is minimal.
[0095] Figure 6 The convergence trend of the model is demonstrated, confirming the robust generalization and lack of overfitting of the model. Figure 6 Figure (a) depicts the gradual reduction of the mean absolute error (MAE) over the training epochs, with the training and peak validation errors stabilizing below 0.1 after 500 epochs. Figure 6 Figure (b) shows the monotonic decrease of the mean squared error (MSE), with the training and test curves converging to almost the same value.
[0096] According to the evaluation results, it was found that the trained KAN framework had good prediction effect, and it was used as the KAN-based solubility prediction model of the present invention to predict the solubility of the dataset after sulfur dioxide was dissolved in a deep eutectic solvent.
[0097] Example 2
[0098] This embodiment provides a system for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN, comprising:
[0099] Data preprocessing module: used to obtain the data set after sulfur dioxide is dissolved in a deep eutectic solvent, and perform outlier detection and normalization;
[0100] Knowledge-enhanced network model module: used to input the preprocessed data set into a pre-trained KAN-based solubility prediction model and output the solubility prediction results, wherein the KAN-based solubility prediction model includes an input layer, a knowledge-enhanced layer, an attention mechanism layer, a deep feature processing layer and an output layer; by loading the KAN-based solubility prediction model, it is used to predict the solubility of sulfur dioxide in deep eutectic solvents.
[0101] The Knowledge Enhanced Network Model module includes a FastAPI Web Application module that provides the following features:
[0102] Single prediction function: predict the sulfur dioxide absorption capacity under specific deep eutectic solvent combinations and conditions;
[0103] Variable range prediction function: Generate a curve chart showing the effect of changes in deep eutectic solvent ratio, temperature, pressure or water content on sulfur dioxide absorption capacity;
[0104] HBA / HBD type prediction function: compare the effects of different HBA or HBD types on sulfur dioxide absorption capacity;
[0105] In addition, the solubility prediction system of the embodiment of the present invention also includes a visualization module for generating a chart display of the prediction results using Plotly. Figure 7 As shown, the system's user interface is simple and intuitive, offering drop-down menus for selecting different types of HBA and HBD, as well as input fields for setting the deep eutectic solvent ratio, temperature, pressure, and water content. After setting the parameters, the user clicks the "Predict" button, which invokes the prediction function and displays the results. The interface is divided into three main functional tabs:
[0106] 1. Single Prediction Tab: Users can select a specific HBA and HBD combination, set the deep eutectic solvent ratio, temperature, pressure, and water content, and obtain a predicted sulfur dioxide absorption capacity under these single conditions. Prediction results are displayed intuitively as numerical values and a dashboard.
[0107] 2. Variable Range Prediction Tab: Users can select a variable (deep eutectic solvent ratio, temperature, pressure, or water content) as an independent variable and set its range. The system will then generate a graph showing the effect of the variable change on sulfur dioxide absorption capacity, helping users understand the impact of process parameters on absorption performance.
[0108] 3. HBA / HBD Type Prediction Tab: Users can select a component (HBA or HBD) and compare the effects of different types of another component on the sulfur dioxide absorption capacity. A bar chart is generated to show the performance differences between different components, helping users to select the optimal deep eutectic solvent combination.
[0109] like Figure 8 As shown in the figure, the system's visualization module uses the Plotly library to generate interactive charts, supporting functions such as zooming, panning, and hovering to view data points, improving the user experience. The system also implements a responsive design to adapt to devices of different screen sizes.
[0110] Experimental verification shows that the FastAPI-based application developed by the present invention provides a user-friendly interface, supports real-time solubility prediction and solvent screening, and has a response time of less than 2 seconds. It has been deployed in three pilot-scale carbon capture facilities. Compared with traditional methods, the deep eutectic solvent system supported by the present invention can reduce water consumption and energy loss, providing a more environmentally friendly and efficient solution for air pollution control. In addition, the system is combined with SHAP (SHapley Additive exPlanations) analysis. The present invention can quantitatively explain the effects of pressure, temperature and deep eutectic solvent composition on the solubility of sulfur dioxide, providing mechanistic guidance for solvent design.
[0111] System performance evaluation results show that, on a hardware configuration (AMD Ryzen 9 9950 processor, NVIDIA RTX4080SUPER graphics card, 32GB RAM), the average response time for a single prediction is 0.8 seconds, the average response time for a variable range prediction (20 data points) is 1.5 seconds, and the average response time for HBA / HBD type prediction is 1.8 seconds. The system has been deployed in three pilot-scale carbon capture facilities to guide deep eutectic solvent selection and process parameter optimization, achieving promising results.
[0112] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0113] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.
[0114] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0115] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0116] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0117] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0118] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN, characterized in that: The following steps are involved: A dataset of sulfur dioxide dissolved in a deep eutectic solvent is obtained, preprocessed, and input into a pre-trained KAN-based solubility prediction model to output a solubility prediction result. The KAN-based solubility prediction model includes an input layer, a knowledge enhancement layer, an attention mechanism layer, a deep feature processing layer, and an output layer. The execution process of the KAN-based solubility prediction model includes: Inputting the preprocessed data set into the knowledge enhancement layer through the input layer for feature extraction; The attention mechanism layer is used to dynamically weight and fuse the extracted features to obtain fused features; A deep feature processing layer is used to perform deep processing on the fused features, and residual connection processing is performed in combination with the fused features. Finally, the solubility of sulfur dioxide in the deep eutectic solvent is predicted through the output layer.
2. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The data set includes chemical structure data and process condition data, wherein the chemical structure data includes hydrogen bond acceptors and hydrogen bond donors, and the process condition data includes deep eutectic solvent ratio, temperature, pressure, and water content.
3. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The preprocessing operation includes outlier detection and data normalization, wherein the Z-Score method is used for outlier detection, and data points with absolute Z-Score values greater than a preset value are identified as outliers and removed from the data set; The data normalization process converts the data into a standard normal distribution with a mean of 0 and a variance of 1, or scales the data to the interval [0, 1].
4. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The input layer and the knowledge enhancement layer are further connected in sequence with a first dense block containing a ReLU activation function, a batch normalization layer, and a dropout layer.
5. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The knowledge enhancement layer includes two parallel branches, the first branch is a second dense block containing a TanH activation function, which is processed using the TanH activation function to obtain chemical structure features, and the second branch is a third dense block containing a Sigmoid activation function, which is processed using the Sigmoid activation function to obtain process condition features.
6. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 5, characterized in that: The execution steps of the attention mechanism layer include: The Softmax activation function in the fourth dense block is used to generate normalized weight coefficients of different features; Performing element-wise multiplication of the normalized weight coefficients with the corresponding output features of different branches to dynamically weight different features; The dynamically weighted features are spliced to obtain fused features.
7. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The deep feature processing layer includes a fifth dense block including a ReLU activation function, a batch normalization layer, a dropout layer, a sixth dense block including a ReLU activation function, and a seventh dense block respectively connected to the attention mechanism layer and the sixth dense block residual.
8. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: During the training process of the KAN-based solubility prediction model, the loss function used is: L = 0.7*MSE+0.3*MAE Where L is the loss, MSE is the mean square error, and MAE is the mean absolute error.
9. The method for predicting the solubility of sulfur dioxide in a deep eutectic solvent based on KAN according to claim 1, characterized in that: The coefficient of determination R 2 , mean square error (MSE), root mean square error (RMSE), average absolute relative deviation (AARD) and mean absolute error (MAE) were used to evaluate the performance indicators of the KAN-based solubility prediction model.
10. A KAN-based prediction system for the solubility of sulfur dioxide in deep eutectic solvents, characterized in that: include: Data preprocessing module: used to obtain the data set after sulfur dioxide is dissolved in a deep eutectic solvent and perform preprocessing; The knowledge-enhanced network model module is used to input the preprocessed data set into the pre-trained KAN-based solubility prediction model and output the solubility prediction results. The KAN-based solubility prediction model includes an input layer, a knowledge-enhanced layer, an attention mechanism layer, a deep feature processing layer, and an output layer. The execution process of the KAN-based solubility prediction model includes: Inputting the preprocessed data set into the knowledge enhancement layer through the input layer for feature extraction; The attention mechanism layer is used to dynamically weight and fuse the extracted features to obtain fused features; A deep feature processing layer is used to perform deep processing on the fused features, and residual connection processing is performed in combination with the fused features. Finally, the solubility of sulfur dioxide in the deep eutectic solvent is predicted through the output layer.