Water-soluble bag solubility detection method based on big data

Through the water-soluble bag solubility detection method based on big data and the use of random forest regression analysis to build a prediction model, the problems of low efficiency and high subjectivity of traditional detection methods were solved, and the efficiency and accuracy of water-soluble bag solubility detection were achieved.

CN120727136APending Publication Date: 2025-09-30BEIJING QINCHENG YIXIN TECH DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510868464.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Traditional solubility testing methods for water-soluble bags are inefficient and highly subjective, making it difficult to comprehensively and accurately evaluate the solubility of water-soluble bags.

Method used

A big data-based detection method was used to collect and clean data, perform feature selection and random forest regression analysis, build a prediction model, and optimize and evaluate the model to achieve accurate prediction of the solubility of water-soluble bags.

Benefits of technology

The accuracy and efficiency of detection are improved, and the test results can be obtained quickly without manual experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120727136A_ABST
    Figure CN120727136A_ABST
Patent Text Reader

Abstract

The invention discloses a water-soluble bag solubility detection method based on big data, and relates to the technical field of water-soluble bag solubility detection.The method comprises the steps that S1, solubility performance data of a water-soluble bag under different environmental conditions and raw material information data of the water-soluble bag are collected; s2, key features influencing the solubility of the water-soluble bag are extracted from the collected original data, and feature selection is carried out; s3, constructing a prediction model based on a random forest regression analysis algorithm; s4, training and optimizing the prediction model; s5, predicting the solubility of the water-soluble bag sample; by collecting a large amount of source data and operating an advanced data processing and analysis technology, the relationship between the solubility of the water-soluble bag and various influencing elements can be reflected more comprehensively and accurately by constructing the prediction model, and compared with a traditional detection means, the accuracy of detection data is greatly improved, and the detection efficiency is improved. The method is used for developing water-soluble packaging bag products and safely applying the water-soluble packaging bag products.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of water-soluble bag solubility detection technology, and more specifically, to a water-soluble bag solubility detection method based on big data. Background Art

[0002] Water-soluble bags, also known as water-soluble packaging bags, are multi-layer composite packaging materials made of polymers, polysaccharides, etc. They are waterproof and have the characteristics of self-disintegration and dissolution in water at a specific temperature. They are disposable plastic packaging bags, and their solubility directly affects the use effect and quality of the product. Traditional water-soluble bag solubility detection methods mainly rely on manual operation and simple physical and chemical analysis, such as observing the dissolution time of water-soluble bags in specific solvents and the clarity of the solution after dissolution. These traditional detection methods have many limitations, such as low detection efficiency, strong subjectivity, and difficulty in comprehensively and accurately evaluating the solubility of water-soluble bags.

[0003] Therefore, in response to the above-mentioned related technical issues, a water-soluble bag solubility detection method based on big data is proposed. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, the present application provides a water-soluble bag solubility detection method based on big data to solve the problems raised in the above-mentioned background technology.

[0005] To achieve the above objectives, the present application provides the following technical solutions: a method for detecting the solubility of water-soluble bags based on big data, the method comprising:

[0006] S1. Collect the solubility performance data of water-soluble bags under different environmental conditions and the raw material information data of water-soluble bags, and then perform data cleaning and data standardization on the collected raw data;

[0007] S2. Extract key features that affect the solubility of water-soluble bags from the collected raw data and perform feature selection;

[0008] S3. Build a prediction model based on the random forest regression analysis algorithm and calculate the out-of-bag error;

[0009] S4. Train and optimize the prediction model;

[0010] S5. Evaluate the prediction model and use it to predict the solubility of new water-soluble bag samples.

[0011] Furthermore, the solubility data in S1 includes but is not limited to temperature, water and solvent types and solvent pH value, and the raw material information data of the water-soluble bag includes material composition, thickness and production process parameters;

[0012] In the data cleaning, duplicate, erroneous and incomplete data records are removed;

[0013] The data standardization process uses normalization or one-hot encoding to standardize data of different types and magnitudes so that the data has a unified dimension and value range.

[0014] Furthermore, the feature selection in S2 uses an impact analysis method and a machine learning algorithm to screen out several features that affect the solubility of the water-soluble bag. Through correlation analysis, the correlation coefficient between each feature and the dissolution time or dissolution degree of the water-soluble bag is calculated, and features with low correlation are removed. Then, the random forest algorithm model is used to select features with high scores as key features based on the importance scores of the features in the random forest algorithm model.

[0015] Furthermore, the specific method of constructing the prediction model is:

[0016] The environmental condition data after feature selection is used as the feature set, and the solubility data of the water-soluble bag is used as the label set to form the training data set;

[0017] Multiple sample subsets are extracted from the training data set with replacement through bootstrap sampling, and a regression decision tree is trained for each sample subset;

[0018] In the process of building a regression decision tree, for each split node, a global search is performed in the feature subspace to screen out the optimal split point that minimizes the node mean square error;

[0019] Suppose the sample set of the current node is , the feature set is , for each feature , try different split points t, and divide the sample set D into the left child node and right child node , calculate the mean square error MSE after segmentation: ,in, represents the number of samples in the sample set D, and Represent the mean square error of the left child node and the right child node respectively, and select the feature and split point that minimizes the mean square error as the splitting condition of the current node.

[0020] Furthermore, the specific method for calculating the out-of-bag error is:

[0021] The out-of-bag error value of each regression decision tree is calculated separately through out-of-bag data, and the arithmetic average of the out-of-bag errors of all regression decision subtrees is taken as the overall out-of-bag error evaluation indicator of the random forest model. Among them, the out-of-bag data is: In the process of constructing the random forest model in the random forest regression analysis algorithm, the bootstrap sampling method will produce a unique data distribution mechanism. Each time a single regression decision tree is trained, 30%-35% of the data will not participate in the training process of the current tree. These unused data are called out-of-bag data;

[0022] Specifically: Assume that there are T regression decision trees in the random forest heavy chemical industry, and the out-of-bag error of the i-th tree is , then the overall out-of-bag error of random forest is for: .

[0023] Furthermore, when optimizing the prediction model in S4, the parameters of the random forest model are optimized by taking the out-of-bag error and the prediction model running time as indicators. The specific steps are as follows:

[0024] Select the number of features m when the initial split node is selected, fix the value of m, obtain the change of out-of-bag error and prediction model training time with the number of regression decision subtrees n, and find the number of regression decision subtrees that minimizes the out-of-bag error ;

[0025] Select , obtain the out-of-bag error and the prediction model training time as the number of features m when splitting the node changes, and find the number of features when splitting the node that minimizes the out-of-bag error .

[0026] Furthermore, when the prediction model is trained in S4, the optimized parameters are first used to train the random forest model based on the entire training data set to obtain the final water-soluble bag solubility prediction model. During the training process, the parameters of the regression decision tree are continuously iterated and updated.

[0027] Furthermore, when evaluating the prediction model in S5, the collected data is first divided into a training set, a validation set, and a test set. The trained prediction model is verified using the validation set, and the performance of the prediction model is evaluated by calculating the root mean square error and the mean absolute error. The calculation formula of the root mean square error RMSE is: , the calculation formula of mean absolute error MAE is: , where N is the number of test set samples, is the true value, is the model's predicted value.

[0028] Furthermore, the prediction model that has been trained and evaluated effectively in S5 is used to predict the solubility of new water-soluble bag samples. The environmental condition data of the new sample and the key feature data after feature extraction and selection are input, and the prediction model outputs the predicted dissolution time and dissolution degree.

[0029] The technical effects and advantages of this application are:

[0030] Compared with existing technologies, this big data-based water-soluble bag solubility detection method, by collecting a large amount of source data and using advanced data processing and analysis technologies, constructs a prediction model that can more comprehensively and accurately reflect the relationship between the solubility of water-soluble bags and various influencing elements. Compared with traditional detection methods, it greatly improves the accuracy of detection data.

[0031] Detection based on big data and predictive models only requires relevant data and does not require manual experiments and relationships to quickly obtain test results, saving time and effort and improving detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a flowchart of the method of this application. DETAILED DESCRIPTION

[0033] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0034] Example 1

[0035] As attached Figure 1 A water-soluble bag solubility detection method based on big data is shown, and the detection method includes:

[0036] S1. Collect the solubility performance data of water-soluble bags under different environmental conditions and the raw material information data of water-soluble bags, and then perform data cleaning and data standardization on the collected raw data;

[0037] The solubility data in S1 include but are not limited to temperature, water and solvent types and solvent pH value; the raw material information data of the water-soluble bag includes material composition, thickness and production process parameters;

[0038] During data cleaning, duplicate, erroneous, and incomplete data records are removed;

[0039] Data standardization uses normalization or one-hot encoding to standardize data of different types and magnitudes so that the data has a unified dimension and value range.

[0040] S2. Extract key features that affect the solubility of the water-soluble bag from the collected raw data and perform feature selection. In S2, feature selection uses impact analysis and machine learning algorithms to screen out several features that affect the solubility of the water-soluble bag. Through correlation analysis, the correlation coefficient between each feature and the dissolution time or dissolution degree of the water-soluble bag is calculated, and features with low correlation are removed. Then, using the random forest algorithm model, based on the feature importance score in the random forest algorithm model, features with high scores are selected as key features.

[0041] S3. Build a prediction model based on the random forest regression analysis algorithm and calculate the out-of-bag error. The specific method for building the prediction model is as follows:

[0042] The environmental condition data after feature selection is used as the feature set, and the solubility data of the water-soluble bag is used as the label set to form the training data set;

[0043] Multiple sample subsets are extracted from the training data set with replacement through bootstrap sampling, and a regression decision tree is trained for each sample subset;

[0044] In the process of building a regression decision tree, for each split node, a global search is performed in the feature subspace to screen out the optimal split point that minimizes the node mean square error;

[0045] Suppose the sample set of the current node is , the feature set is , for each feature , try different split points t, and divide the sample set D into the left child node and right child node , calculate the mean square error MSE after segmentation: ,in, represents the number of samples in the sample set D, and Represent the mean square error of the left child node and the right child node respectively, and select the feature and split point that minimizes the mean square error as the splitting condition of the current node.

[0046] The specific method for calculating the out-of-bag error is:

[0047] The out-of-bag error value of each regression decision tree is calculated separately through out-of-bag data, and the arithmetic average of the out-of-bag errors of all regression decision subtrees is taken as the overall out-of-bag error evaluation indicator of the random forest model. Among them, the out-of-bag data is: In the process of constructing the random forest model in the random forest regression analysis algorithm, the bootstrap sampling method will produce a unique data distribution mechanism. Each time a single regression decision tree is trained, 30%-35% of the data will not participate in the training process of the current tree. These unused data are called out-of-bag data;

[0048] Specifically: Assume that there are T regression decision trees in the random forest heavy chemical industry, and the out-of-bag error of the i-th tree is , then the overall out-of-bag error of random forest is for: ;

[0049] S4. Train and optimize the prediction model. When optimizing the prediction model, the parameters of the random forest model are optimized using the out-of-bag error and the prediction model running time as indicators. The specific steps are as follows:

[0050] Select the number of features m when the initial split node is selected, fix the value of m, obtain the change of out-of-bag error and prediction model training time with the number of regression decision subtrees n, and find the number of regression decision subtrees that minimizes the out-of-bag error ;

[0051] Select , obtain the out-of-bag error and the prediction model training time as the number of features m when splitting the node changes, and find the number of features when splitting the node that minimizes the out-of-bag error .

[0052] When training the prediction model, the optimized parameters are first used to train the random forest model based on the entire training data set to obtain the final water-soluble bag solubility prediction model. During the training process, the parameters of the regression decision tree are continuously iteratively updated;

[0053] S5. Evaluate the prediction model and use it to predict the solubility of new water-soluble bag samples.

[0054] When evaluating the prediction model, the collected data is first divided into a training set, a validation set, and a test set. The validation set is used to verify the trained prediction model, and the performance of the prediction model is evaluated by calculating the root mean square error and the mean absolute error. The calculation formula for the root mean square error RMSE is: , the calculation formula of mean absolute error MAE is: , where N is the number of test set samples, is the true value, is the model's predicted value.

[0055] A well-trained and well-evaluated prediction model is used to predict the solubility of new water-soluble bag samples. The environmental condition data of the new sample and the key feature data after feature extraction and selection are input, and the prediction model outputs the predicted dissolution time and dissolution degree.

[0056] Example 2

[0057] Take the solubility test of pesticides in water-soluble bags as an example:

[0058] Prepare a variety of pesticide water-soluble bag samples of different materials and thicknesses, and conduct dissolution tests at different temperatures (15°C, 25°C, and 30°C) and different pesticide aqueous solvents (pH values ​​of 5, 6, and 7). Record the dissolution time and dissolution degree of each sample. A total of 100 sets of experiments were conducted to obtain 100 pieces of dissolution performance data.

[0059] Collect the production process parameters of these water-soluble bags, including raw material suppliers, material composition ratios, film blowing temperature, blow-up ratio, etc.

[0060] After checking the data records, we found that 3 data records had temperature record errors, with temperature values ​​outside the reasonable range, and 5 data records lacked humidity information. After removing these data, we obtained 92 valid data records.

[0061] Normalize the temperature data and normalize the range of 20℃-30℃ to interval, i.e. ,in, is the actual temperature, is the normalized temperature. For the material composition ratio data, one-hot encoding is used;

[0062] Extract the temperature change rate (the ratio of the temperature difference between two adjacent experiments to the time interval) and humidity change rate from the environmental condition data; extract the proportion of the main components in the material, the difference between the film blowing temperature and the average film blowing temperature and other characteristics from the production process parameters;

[0063] Through correlation analysis, we found that temperature, solvent pH value, and the proportion of a key component in the material had a high correlation coefficient with dissolution time, so we retained these features. We used the random forest algorithm to evaluate the importance of these features and finally selected temperature, humidity, solvent pH value, and the proportion of a key component in the material as key features.

[0064] The environmental condition data and key feature data after feature selection were used as the feature set, and the dissolution time was used as the label set to form the training data set. Then, 50 sample subsets were extracted from the training data set using the bootstrap sampling method. A regression decision tree was trained for each sample subset. When constructing the regression decision tree, the features and split points of the split nodes were selected according to the above-mentioned mean square error minimization principle;

[0065] The out-of-bag error of each regression decision tree is calculated using out-of-bag data, and the overall out-of-bag error of the random forest is 0.25 (unit: minutes);

[0066] Then optimize the parameters, first fix the number of features when splitting the node , adjust the number of regression decision subtrees , found that when When the out-of-bag error is minimum, then fix , adjust the number of features m when splitting the node, and find that when 2, the out-of-bag error is the smallest;

[0067] Using the optimized parameters ( , 2) Retrain the random forest model based on the entire training data set;

[0068] The 92 data points were divided into training, validation, and test sets at a ratio of 70%, 15%, and 15%. The trained model was evaluated using the test set, and the root mean square error (RMSE) was 0.3 (unit: minute) and the mean absolute error (MAE) was 0.2 (unit: minute), indicating that the model has good prediction performance.

[0069] For newly produced pesticide water-soluble bag samples, the environmental condition data and key characteristic data were input. The model predicted that the dissolution time in the specific pesticide water solvent was 5.2 minutes. The actual tested dissolution time was 5.5 minutes. The predicted result was close to the actual one, meeting the testing needs in production.

[0070] Finally: The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should be included in the scope of protection of the present application.

Claims

1. A method for detecting the solubility of water-soluble bags based on big data, characterized in that: The test method includes: S1. Collect the solubility performance data of water-soluble bags under different environmental conditions and the raw material information data of water-soluble bags, and then perform data cleaning and data standardization on the collected raw data; S2. Extract key features that affect the solubility of water-soluble bags from the collected raw data and perform feature selection; S3. Build a prediction model based on the random forest regression analysis algorithm and calculate the out-of-bag error; S4. Train and optimize the prediction model; S5. Evaluate the prediction model and use it to predict the solubility of new water-soluble bag samples.

2. The method for detecting the solubility of a water-soluble bag based on big data according to claim 1, characterized in that: The solubility data in S1 include but are not limited to temperature, water and solvent types and solvent pH value, and the raw material information data of the water-soluble bag include material composition, thickness and production process parameters; In the data cleaning, duplicate, erroneous and incomplete data records are removed; The data standardization process uses normalization or one-hot encoding to standardize data of different types and magnitudes so that the data has a unified dimension and value range.

3. The method for detecting the solubility of a water-soluble bag based on big data according to claim 1, wherein: The feature selection in S2 uses an impact analysis method and a machine learning algorithm to screen out several features that affect the solubility of the water-soluble bag. Through correlation analysis, the correlation coefficient between each feature and the dissolution time or dissolution degree of the water-soluble bag is calculated, and features with low correlation are removed. Then, the random forest algorithm model is used to select features with high scores as key features based on the importance scores of the features in the random forest algorithm model.

4. The method for detecting the solubility of a water-soluble bag based on big data according to claim 1, wherein: The specific method of constructing the prediction model is: The environmental condition data after feature selection is used as the feature set, and the solubility data of the water-soluble bag is used as the label set to form the training data set; Multiple sample subsets are extracted from the training data set with replacement through bootstrap sampling, and a regression decision tree is trained for each sample subset; In the process of building a regression decision tree, for each split node, a global search is performed in the feature subspace to screen out the optimal split point that minimizes the node mean square error; Suppose the sample set of the current node is , the feature set is , for each feature , try different split points t, and divide the sample set D into the left child node and right child node , calculate the mean square error MSE after segmentation: ,in, represents the number of samples in the sample set D, and Represent the mean square error of the left child node and the right child node respectively, and select the feature and split point that minimizes the mean square error as the splitting condition of the current node.

5. The method for detecting the solubility of a water-soluble bag based on big data according to claim 4, characterized in that: The specific method for calculating the out-of-bag error is: The out-of-bag error value of each regression decision tree is calculated separately through out-of-bag data, and the arithmetic average of the out-of-bag errors of all regression decision subtrees is taken as the overall out-of-bag error evaluation indicator of the random forest model. Among them, the out-of-bag data is: In the process of constructing the random forest model in the random forest regression analysis algorithm, the bootstrap sampling method will produce a unique data distribution mechanism. Each time a single regression decision tree is trained, 30%-35% of the data will not participate in the training process of the current tree. These unused data are called out-of-bag data; Specifically: Assume that there are T regression decision trees in the random forest heavy chemical industry, and the out-of-bag error of the i-th tree is , then the overall out-of-bag error of random forest is for: .

6. The method for detecting the solubility of a water-soluble bag based on big data according to claim 1, wherein: When optimizing the prediction model in S4, the parameters of the random forest model are optimized by taking the out-of-bag error and the prediction model running time as indicators. The specific steps are as follows: Select the number of features m when the initial split node is selected, fix the value of m, obtain the change of out-of-bag error and prediction model training time with the number of regression decision subtrees n, and find the number of regression decision subtrees that minimizes the out-of-bag error ; Select , obtain the out-of-bag error and the prediction model training time as the number of features m when splitting the node changes, and find the number of features when splitting the node that minimizes the out-of-bag error .

7. The method for detecting the solubility of a water-soluble bag based on big data according to claim 6, characterized in that: When the prediction model is trained in S4, the optimized parameters are first used to train the random forest model based on the entire training data set to obtain the final water-soluble bag solubility prediction model. During the training process, the parameters of the regression decision tree are continuously iterated and updated.

8. The method for detecting the solubility of a water-soluble bag based on big data according to claim 7, characterized in that: When evaluating the prediction model in S5, the collected data is first divided into a training set, a validation set, and a test set. The trained prediction model is verified using the validation set, and the performance of the prediction model is evaluated by calculating the root mean square error and the mean absolute error. The calculation formula of the root mean square error RMSE is: , the calculation formula of mean absolute error MAE is: , where N is the number of test set samples, is the true value, is the model's predicted value.

9. The method for detecting the solubility of a water-soluble bag based on big data according to claim 8, wherein: The prediction model that has been well trained and evaluated is used in S5 to predict the solubility of the new water-soluble bag sample. The environmental condition data of the new sample and the key feature data after feature extraction and selection are input, and the prediction model outputs the predicted dissolution time and dissolution degree.