Element to phase rapid analysis method based on classification model
Patent Information
- Application Number
- CN202410590944.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-13
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-05-13
AI Technical Summary
但在大多数实际生产过程中,原料来源广,成分复杂,需要快速对样品进行检测并对生产过程进行调整,采用人工经验难以实现
在本发明中,一方面,以元素种类及含量为输入,通过元素信息演算物相信息,弥补了现有技术中物相检测速度慢、复杂样品定量分析难的缺点;另一方面,基于分类模型建立元素种类及含量与样品种类之间的关系,可采用不同的样品数据集对分类模型进行训练,适用性强。
Smart Images

Figure CN118430681B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of phase detection technology, specifically to a rapid element-to-phase analysis method based on a classification model. Background Technology
[0002] A phase is a substance with a specific physicochemical structure; the same element can exist in different compound states within a substance. Phase detection is an important means of understanding the phase composition of a sample. Currently, conventional phase detection methods include X-ray diffraction analysis, ultraviolet-visible spectrophotometry, and gas chromatography-mass spectrometry. However, these methods are relatively slow, and for samples with complex compositions, or those containing amorphous or unknown crystalline phases, quantitative analysis is difficult.
[0003] Currently, elemental detection technology is mature, and methods such as X-ray fluorescence spectroscopy can quickly detect the types and contents of elements in a sample. In some production processes with fixed raw materials and stable chemical composition, elemental detection methods are used to detect elemental information, and the approximate compound composition and content can be estimated using human experience. However, in most actual production processes, raw materials are diverse and complex, requiring rapid sample detection and adjustments to the production process, which is difficult to achieve using human experience alone. Classification models, on the other hand, can establish a functional model between input features and output labels using large amounts of user-provided data, thereby enabling the prediction of unknown data categories.
[0004] Therefore, a rapid element-to-phase analysis method based on a classification model is provided. Summary of the Invention
[0005] Therefore, in order to overcome the above-mentioned defects in the prior art, the present invention provides a rapid element-to-phase analysis method based on a classification model to improve the phase detection speed in industrial production and thereby achieve rapid feedback in the production process.
[0006] This invention discloses a rapid element-to-phase analysis method based on a classification model, comprising the following steps: S1. Use elemental detection methods and phase detection methods to detect known samples and establish a database of known sample types, element types and contents, and compound types; S2. The classification model is trained using the known element types and contents of the sample as input features and the known sample types as output labels. S3. Use elemental detection methods to detect unknown samples and obtain the types and contents of elements in the unknown samples; S4. Using the element types and contents of unknown samples as input, predict the type of unknown samples through a trained classification model; S5. Using the predicted types of unknown samples as input, search the database to predict the types of compounds in the unknown samples; S6. Using the element types and contents of the unknown sample and the predicted types of compounds as input, calculate the compound contents of the unknown sample based on the law of conservation of elements.
[0007] In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into a training dataset and a test dataset; S26. Calculate the original predicted value, calculate the predicted probability, calculate the loss function, calculate the gradient of the loss function with respect to the parameter matrix, and update the parameter matrix; S27. Determine whether the iteration count and the change in the loss function satisfy the iteration termination condition; if they do, stop the iteration and output the parameter matrix; otherwise, return to step S26.
[0008] In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into a training dataset and a test dataset; S26. Calculate the distance, find the K nearest neighbors, calculate the probability of each category, and output the predicted category.
[0009] In S23, the transformation function for normalization is: (1); in, Data before normalization; The data is after normalization; The mean of all sample data. represents the standard deviation of all sample data.
[0010] In S26, given the analysis model is a multinomial logistic regression classification model, the training phase first initializes a parameter matrix of size N·(n+1). The initial value is set to 0, where N is the number of sample categories and n is the number of elements; S261. Calculate the original predicted value : ; in, The input feature array for sample i in the training dataset; S262. Calculate the predicted probability: ; in, For sample i, which is of category j, the predicted value is... Let i be the predicted value for sample i belonging to category k; S263. Calculate the loss function J(Θ): ; Where m is the number of samples; For sample i, the actual value is for category j; S264. Calculate the gradient ∇J(Θ) of the loss function with respect to the parameter matrix Θ: ; in, Input the k-th feature value from the feature array for sample i; Let θ be the weight parameter of sample i corresponding to category k in the parameter matrix Θ; S265. Update parameter matrix Θ: ; in, This is the learning rate.
[0011] In S26, the distance is calculated, the K nearest neighbors are found, the probability of each category is calculated, and the predicted category is output; the given analysis model is the K nearest neighbor classification model. S261. Calculate Euclidean Distance ; ; in, Input the k-th feature value from the feature array for the unknown sample; S262. Based on the calculated distances, find the K training samples that are closest to the unknown sample: ; Where M is the total number of samples in the training set; S263. Calculate the probability of each category appearing in K neighbors: ; in, The number of times category k appears among the K neighbors; S264. Output the category with the highest probability among N categories as the predicted category. If multiple categories have the same probability, select the first one as the predicted category.
[0012] In S27, the changes in the number of iterations n and the loss function are determined. Check if the iteration termination condition is met; if it is, stop the iteration and output the parameter matrix Θ; otherwise, return to step S26; where the iteration termination condition is: ; in, This represents the maximum number of iterations. This is the threshold for the change in the loss function.
[0013] In S6, the following are included: S61. Calculate the mass fraction of each element in each compound based on its elemental composition and relative atomic mass, as shown below: ; in, It is a compound; For compounds Elements contained therein; For compounds middle The relative atomic mass of an element; For compounds middle The number of atoms of an element; For compounds The total number of elements; S62. Establish a nonlinear optimization model based on the law of conservation of mass, as shown below: ; Its constraints are: ; in, The sum of squared errors for each element; Unknown sample The mass percentage of the element; For compounds middle Mass fraction of the element; For compounds Percentage of mass; This represents the total number of compounds in the unknown sample.
[0014] The technical solution of this invention has the following advantages: In this invention, on the one hand, the phase information is calculated by taking the element type and content as input, which makes up for the shortcomings of the existing technology that the phase detection speed is slow and the quantitative analysis of complex samples is difficult; on the other hand, the relationship between the element type and content and the sample type is established based on the classification model, and different sample datasets can be used to train the classification model, which has strong applicability.
[0015] In other words, this invention uses a rapid elemental detection method to obtain the types and contents of elements in a sample. The classification model uses the types and contents of elements to predict the phase composition of the sample, and the phase content is indirectly calculated based on the law of conservation of mass. This analytical method is of great significance for improving the speed of phase detection in industrial production and realizing rapid feedback in the production process. Attached Figure Description
[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating the analytical method described in this invention. Detailed Implementation
[0018] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0020] Example 1: This example provides a rapid element-to-phase analysis method based on a classification model, including the following steps: S1. Use elemental detection methods and phase detection methods to detect known samples and establish a database of known sample types, element types and contents, and compound types; S2. The classification model is trained using the known element types and contents of the sample as input features and the known sample types as output labels. S3. Use elemental detection methods to detect unknown samples and obtain the types and contents of elements in the unknown samples; S4. Using the element types and contents of unknown samples as input, predict the type of unknown samples through a trained classification model; S5. Using the predicted types of unknown samples as input, search the database to predict the types of compounds in the unknown samples; S6. Using the element types and contents of the unknown sample and the predicted types of compounds as input, calculate the compound contents of the unknown sample based on the law of conservation of elements.
[0021] In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; Specifically, this embodiment uses CaSO4 compounds as an example and employs X-ray fluorescence spectroscopy to detect electroplating sludge. The detection results are shown in Table 1. Table 1 Elemental Analysis Information of CaSO4 Compounds
[0022] Create a one-dimensional array of size 118, and input the corresponding element content into the array in atomic number order. The element input feature array for CaSO4 compounds is as follows: ; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; Specifically, the transformation function for normalization is: (1); in, Data before normalization; The data is after normalization; The mean of all sample data. The standard deviation of all sample data; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. For example: if the electroplating sludge is numbered 1, the circuit board is numbered 2, and the soot is numbered 3, and the input feature dataset is merged in the order of electroplating sludge 1, electroplating sludge 2, circuit board 1, soot 1, circuit board 2, then the output labels are: ; S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into a training dataset and a test dataset, with the ratio of the two datasets being 9:1, 8:2, 7:3, 6:4, or 5:5. S26. Calculate the original predicted value, calculate the predicted probability, calculate the loss function, calculate the gradient of the loss function with respect to the parameter matrix, and update the parameter matrix; Specifically, in this embodiment, the preferred analysis model is a multinomial logistic regression classification model. During the training phase, a parameter matrix Θ of size N·(n+1) is first initialized with an initial value of 0, where N is the number of sample categories and n is the number of elements, which is 118. S261. Calculate the original predicted value : ; in, The input feature array for sample i in the training dataset; S262. Calculate the predicted probability: ; in, For sample i, which is of category j, the predicted value is... Let i be the predicted value for sample i belonging to category k; S263. Calculate the loss function J(Θ): ; Where m is the number of samples; For sample i, the actual value is for category j; S264. Calculate the gradient ∇J(Θ) of the loss function with respect to the parameter matrix Θ: ; in, Input the k-th feature value from the feature array for sample i; Let θ be the weight parameter of sample i corresponding to category k in the parameter matrix Θ; S265. Update parameter matrix Θ: ; in, This is the learning rate.
[0023] Further, in step S27, determine whether the changes in the number of iterations and the loss function satisfy the iteration termination condition; if they do, stop the iteration and output the parameter matrix; otherwise, return to step S26.
[0024] Specifically, the changes in the number of iterations n and the loss function are determined. Check if the iteration termination condition is met; if it is, stop the iteration and output the parameter matrix Θ; otherwise, return to step S26; where the iteration termination condition is: ; in, This represents the maximum number of iterations. This is the threshold for the change in the loss function.
[0025] Based on the above, using the elemental types and contents of the unknown sample and the predicted compound types as inputs, the compound contents are calculated according to the law of conservation of elements. The specific process is as follows (i.e., S6 includes): S61. Calculate the mass fraction of each element in each compound based on its elemental composition and relative atomic mass, as shown below: ; in, It is a compound; For compounds Elements contained therein; For compounds middle The relative atomic mass of an element; For compounds middle The number of atoms of an element; For compounds The total number of elements; S62. Establish a nonlinear optimization model based on the law of conservation of mass, as shown below: ; Its constraints are: ; in, This is the sum of squared errors for each element; Unknown sample The mass percentage of the element; For compounds middle Mass fraction of the element; For compounds Percentage of mass; This represents the total number of compounds in the unknown sample.
[0026] Example 2: The difference between this example and Example 1 is that the analysis model is different.
[0027] The specific analysis method in this embodiment is as follows: S1. Use elemental detection methods and phase detection methods to detect known samples and establish a database of known sample types, element types and contents, and compound types; S2. The classification model is trained using the known element types and contents of the sample as input features and the known sample types as output labels. S3. Use elemental detection methods to detect unknown samples and obtain the types and contents of elements in the unknown samples; S4. Using the element types and contents of unknown samples as input, predict the type of unknown samples through a trained classification model; S5. Using the predicted types of unknown samples as input, search the database to predict the types of compounds in the unknown samples; S6. Using the element types and contents of the unknown sample and the predicted types of compounds as input, calculate the compound contents of the unknown sample based on the law of conservation of elements.
[0028] In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; Specifically, this embodiment uses CaSO4 compounds as an example and employs X-ray fluorescence spectroscopy to detect electroplating sludge. The detection results are shown in Table 1. Table 1 Elemental Analysis Information of CaSO4 Compounds
[0029] Create a one-dimensional array of size 118, and input the corresponding element content into the array in atomic number order. The element input feature array for CaSO4 compounds is as follows: ; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; Specifically, the transformation function for normalization is: (1); in, Data before normalization; The data is after normalization; The mean of all sample data. The standard deviation of all sample data; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. For example: if the electroplating sludge is numbered 1, the circuit board is numbered 2, and the soot is numbered 3, and the input feature dataset is merged in the order of electroplating sludge 1, electroplating sludge 2, circuit board 1, soot 1, circuit board 2, then the output labels are: ; S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into training dataset and test dataset according to the ratio of 9:1, 8:2, 7:3, 6:4 or 5:5. S26. Calculate the distance, find the K nearest neighbors, calculate the probability of each category, and output the predicted category; Specifically, the preferred analysis model in this embodiment is the K-nearest neighbor classification model; S261. Calculate the distance. Distance measurement methods include Euclidean distance, Manhattan distance, etc. This embodiment prefers Euclidean distance. ; ; in, Input the k-th feature value from the feature array for the unknown sample; S262. Based on the calculated distances, find the K training samples closest to the unknown sample. In this example, K is preferably: ; Where M is the total number of samples in the training set; S263. Calculate the probability of each category appearing in K neighbors. : ; in, The number of times category k appears among the K neighbors; S264. Output the category with the highest probability among N categories as the predicted category. If multiple categories have the same probability, select the first one as the predicted category.
[0030] Based on the above, using the elemental types and contents of the unknown sample and the predicted compound types as inputs, the compound contents are calculated according to the law of conservation of elements. The specific process is as follows (i.e., S6 includes): S61. Calculate the mass fraction of each element in each compound based on its elemental composition and relative atomic mass, as shown below: ; in, It is a compound; For compounds Elements contained therein; For compounds middle The relative atomic mass of an element; For compounds middle The number of atoms of an element; For compounds The total number of elements; S62. Establish a nonlinear optimization model based on the law of conservation of mass, as shown below: ; Its constraints are: ; in, This is the sum of squared errors for each element; Unknown sample The mass percentage of the element; For compounds middle Mass fraction of the element; For compounds Percentage of mass; This represents the total number of compounds in the unknown sample.
[0031] Example 3: Based on Example 1, this example takes the element-to-phase calculation of an unknown sample as an example, and uses the rapid element-to-phase analysis method based on the classification model described in this invention for calculation, as detailed below.
[0032] The XRF elemental detection information of the unknown samples is shown in Table 1.
[0033] Table 1. XRF elemental detection information of unknown samples ; The above detection data was input into the classification model of the analytical method described in this invention, which predicted that the sample was electroplating sludge. After searching the database, the compound composition of the electroplating sludge is shown in Table 2.
[0034] Table 2 Predicted Composition of Unknown Sample Compounds
[0035] Furthermore, by using this analytical method, the compound content of the unknown sample can be obtained as shown in Table 3.
[0036] Table 3. Calculation results of elements to phases of unknown samples
[0037] In summary, this application uses element types and contents as inputs to calculate phase information, overcoming the shortcomings of existing technologies such as slow phase detection speed and difficulty in quantitative analysis of complex samples. Furthermore, by establishing a classification model to establish the relationship between element types and contents and sample types, the classification model can be trained using different sample datasets, demonstrating strong applicability. In other words, this invention employs a rapid element detection method to obtain the element types and contents of a sample, and the classification model uses element types and contents to predict the phase composition of the sample, indirectly calculating the phase content based on mass conservation. This analytical method is of great significance for improving the speed of phase detection in industrial production and achieving rapid feedback in the production process.
[0038] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A rapid element-to-phase analysis method based on a classification model, characterized in that, Includes the following steps: S1. Use elemental detection methods and phase detection methods to detect known samples and establish a database of known sample types, element types and contents, and compound types; S2. The classification model is trained using the known element types and contents of the sample as input features and the known sample types as output labels. S3. Use elemental detection methods to detect unknown samples and obtain the types and contents of elements in the unknown samples; S4. Using the element types and contents of unknown samples as input, predict the type of unknown samples through a trained classification model; S5. Using the predicted types of unknown samples as input, search the database to predict the types of compounds in the unknown samples; S6. Using the element types and contents of the unknown sample and the predicted compound types as input, calculate the compound contents of the unknown sample based on the law of conservation of elements; In S6, the following are included: S61. Calculate the mass fraction of each element in each compound based on its elemental composition and relative atomic mass, as shown below: ; in , It is a compound; For compounds Elements contained therein; For compounds middle The relative atomic mass of an element; For compounds middle The number of atoms of an element; For compounds The total number of elements; S62. Establish a nonlinear optimization model based on the law of conservation of mass, as shown below: ; Its constraints are: ; in, The sum of squared errors for each element; Unknown sample The mass percentage of the element; For compounds middle Mass fraction of the element; For compounds Percentage of mass; This represents the total number of compounds in the unknown sample.
2. The rapid element-to-phase analysis method based on a classification model according to claim 1, characterized in that, In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into a training dataset and a test dataset; S26. Calculate the original predicted value, calculate the predicted probability, calculate the loss function, calculate the gradient of the loss function with respect to the parameter matrix, and update the parameter matrix; S27. Determine whether the iteration count and the change in the loss function satisfy the iteration termination condition; if they do, stop the iteration and output the parameter matrix; Otherwise, return to step S26.
3. The rapid element-to-phase analysis method based on a classification model according to claim 1, characterized in that, In S2, the training process of the classification model is as follows: S21. Establish an element input feature array based on the element detection results; S22. Merge the samples that need to be classified into a two-dimensional array. Each element in the two-dimensional array represents the element input feature of a sample. This two-dimensional array is the element input feature data of the classification model. S23. Normalize the element input feature data so that the input feature data satisfies a normal distribution; S24. Number the different types of samples with numbers and form a one-dimensional array according to the order of merging the input feature data. This one-dimensional array is the sample type output label data of the classification model. S25. Combine the input feature data and output label data according to their correspondence to form the total dataset, and divide the total dataset into a training dataset and a test dataset; S26. Calculate the distance, find the K nearest neighbors, calculate the probability of each category, and output the predicted category.
4. The rapid element-to-phase analysis method based on a classification model according to claim 2, characterized in that, In S23, the transformation function for normalization is: (1); in, Data before normalization; The data is after normalization; The mean of all sample data. represents the standard deviation of all sample data.
5. The rapid element-to-phase analysis method based on a classification model according to claim 3, characterized in that, In S23, the transformation function for normalization is: (1); in, Data before normalization; The data is after normalization; The mean of all sample data. represents the standard deviation of all sample data.
6. The rapid element-to-phase analysis method based on a classification model according to claim 4, characterized in that, In S26, given that the analysis model is a multinomial logistic regression classification model, the training phase first initializes the parameter matrix Θ of size N·(n+1) with an initial value of 0, where N is the number of sample categories and n is the number of elements; S261. Calculate the original predicted value : ; in, The input feature array for sample i in the training dataset; S262. Calculate the predicted probability: ; in, For sample i, which is of category j, the predicted value is... Let i be the predicted value for sample i belonging to category k; S263. Calculate the loss function: ; Where m is the number of samples; For sample i, the actual value for category j; S264. Calculate the gradient of the loss function with respect to the parameter matrix Θ. : ; in, Input the k-th feature value from the feature array for sample i; Let θ be the weight parameter of sample i corresponding to category k in the parameter matrix Θ; S265. Update parameter matrix Θ: ; in, This is the learning rate.
7. The rapid element-to-phase analysis method based on a classification model according to claim 5, characterized in that, In S26, the distance is calculated, the K nearest neighbors are found, the probability of each category is calculated, and the predicted category is output; the given analysis model is the K nearest neighbor classification model. S261. Calculate Euclidean Distance ; ; in, Input the k-th feature value from the feature array for the unknown sample; S262. Based on the calculated distances, find the K training samples that are closest to the unknown sample: ; Where M is the total number of samples in the training set; S263. Calculate the probability of each category appearing in K neighbors: ; in, The number of times category k appears among the K neighbors; S264. Output the category with the highest probability among N categories as the predicted category. If multiple categories have the same probability, select the first one as the predicted category.
8. The rapid element-to-phase analysis method based on a classification model according to claim 6, characterized in that, In S27, the changes in the number of iterations n and the loss function are determined. Check if the iteration termination condition is met; if it is, stop the iteration and output the parameter matrix Θ; otherwise, return to step S26; where the iteration termination condition is: ; in, This represents the maximum number of iterations. This is the threshold for the change in the loss function.