Micro-plastic infrared characteristic spectrum extraction and efficient and accurate identification method

By combining equal-interval sampling and competitive adaptive reweighted sampling to extract microplastic feature spectra, and using a genetic algorithm to optimize the artificial neural network model, the problems of high computational cost and easy overfitting in microplastic identification and detection are solved, achieving efficient and accurate microplastic identification.

CN121997011APending Publication Date: 2026-05-08XUZHOU UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XUZHOU UNIV OF TECH
Filing Date
2024-04-02
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing microplastic identification and detection methods are computationally intensive, inefficient, and prone to overfitting, making it difficult to achieve efficient and accurate identification.

Method used

Feature spectra are extracted using a progressive two-step method combining equal-interval sampling and competitive adaptive reweighted sampling. By constructing an artificial neural network model and combining it with a genetic algorithm to optimize hyperparameters, overfitting is prevented and the model's generalization ability is improved.

Benefits of technology

It effectively reduces computational load, improves the efficiency of feature spectrum extraction, prevents model overfitting, and achieves efficient and accurate identification of microplastics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121997011A_ABST
    Figure CN121997011A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of environmental protection and micro-plastic identification, and discloses a micro-plastic infrared characteristic spectrum extraction and efficient and accurate identification method, which comprises the following steps of: 1, acquiring infrared spectrum data of a micro-plastic sample, and constructing a micro-plastic spectrum database; 2, extracting a characteristic spectrum from the infrared spectrum data in the step 1 by adopting a progressive two-step method combining equal-interval sampling and a competitive self-adaptive reweighted sampling algorithm; step 3, performing standard normal transformation on the transmittance of the characteristic spectrum in the step 2; and 4, constructing a feature training set and a feature test set by taking the transmittance of the infrared spectrum transformed in the step 3 as input and the micro-plastic type as output. Training the artificial neural network model by adopting the feature training set, and optimizing hyper-parameters in the model through the feature test set by applying a genetic algorithm to form a final micro-plastic identification artificial neural network model; and step 5, for micro-plastics to be identified, obtaining the characteristic transmittance of the micro-plastics to be identified through the spectral data obtained in the step 1 and the characteristic wave number extracted in the step 2, performing standard normal transformation on the micro-plastics to be identified through the step 3, inputting the transformed infrared spectral transmittance into the micro-plastic identification artificial neural network model, and giving the type of the micro-plastics. And identification of the micro-plastic is realized. According to the method, traversal sampling of the whole spectrum data set is effectively avoided, the extraction efficiency of the characteristic spectrum is improved, hyper-parameter optimization is performed on the artificial neural network model by adopting the genetic algorithm, overfitting of the model is effectively prevented, and the generalization ability of the model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of environmental protection and microplastic identification technology, specifically a method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics. Background Technology

[0002] Plastic products are widely used in cosmetics and personal care products, textiles and clothing, packaging materials, agriculture, construction and building materials, and other fields due to their lightweight, durability, and versatility in color, texture, and shape. However, most consumer plastics are single-use, and a large portion accumulates in natural systems. Influenced by natural forces such as sunlight, oxidation, and weathering, these large plastics entering the environment gradually decompose into tiny plastic pollutants (microplastics are generally defined as those with a diameter of less than 5 mm). Experiments have shown that exposure to microplastics can cause various toxic effects, including oxidative stress, metabolic disorders, immune responses, neurotoxicity, and reproductive and developmental toxicity. Therefore, the environmental and health risks posed by plastics must be given sufficient attention, and microplastic pollution has become one of the most serious environmental challenges of the 21st century.

[0003] Establishing efficient and accurate methods for plastic identification and detection is fundamental to plastic recycling and the development of microplastic pollution control measures. Currently, the main methods for microplastic identification and detection include visual methods, thermal analysis, and spectroscopic methods. Visual methods are highly subjective, while thermal analysis requires measuring the thermal properties of materials at a specific temperature to determine the chemical composition of a sample, which can easily damage the sample. With the development of computer science, the combined use of spectroscopic technology and machine learning has been widely applied in the field of microplastic identification and detection. Infrared spectroscopy has advantages such as simple operation, high sensitivity, accurate wavenumbers, and good repeatability, and is widely used in microplastic identification and detection. However, when using spectroscopic methods to build identification models, full-spectrum modeling is often used, resulting in high computational load and low efficiency. Furthermore, preventing model overfitting and improving its generalization ability are also problems that need to be solved in microplastic identification and detection. Summary of the Invention

[0004] The purpose of this invention is to provide a method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics. This method can extract characteristic spectra while retaining the main information of the infrared spectrum, and achieve efficient and accurate identification of microplastics by constructing an Artificial Neural Network (ANN) model. This method can prevent the loss of characteristic spectra and improve the modeling efficiency of the ANN model.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for extracting and efficiently and accurately identifying infrared feature spectra of microplastics, comprising:

[0006] Step 1: Obtain infrared spectral data of microplastic samples and construct a microplastic spectral database;

[0007] Step 2: Extract feature spectra from the infrared spectral data in Step 1 using a progressive two-step method that combines Equal Interval Sampling (EIS) and Competitive Adaptive Reweighted Sampling (CARS) algorithms.

[0008] Step 3: Perform a standard normal transformation on the transmittance of the characteristic spectra from Step 2;

[0009] Step 4: Using the transmittance of the transformed infrared spectrum from Step 3 as input and the microplastic type as output, construct a feature training set and a feature test set. Train the ANN model using the feature training set, and then use a genetic algorithm to optimize the hyperparameters of the model using the feature test set, forming the final microplastic recognition ANN model.

[0010] Step 5: For the microplastic to be identified, obtain spectral data through Step 1, obtain the characteristic transmittance of the microplastic to be identified through the characteristic wavenumber extracted in Step 2, perform a standard normal transformation on it through Step 3, and input the transformed infrared spectral transmittance into the microplastic identification ANN model to give the type of microplastic and realize the identification of microplastic.

[0011] Preferably, in step 1, infrared spectral data of microplastics are obtained using a Fourier transform infrared spectrometer, and the number of scans for each microplastic sample can be determined through literature review and preliminary experiments.

[0012] Preferably, step 2, which extracts the feature spectrum, is divided into two steps: (1) EIS of the original spectrum; (2) Extracting the feature spectrum based on the CARS algorithm.

[0013] (1) The EIS process for the original spectrum is as follows:

[0014] The abscissa wavenumber vector of the original infrared spectrum is The transmittance vector on the ordinate is in This represents a wavenumber point on the horizontal axis of the infrared spectrum, in cm. -1 ; Indicates the spectrum with The corresponding transmittance, %. When taking points at the wavenumber interval, assume the interval value is N, that is, the first wavenumber point. Starting from point N, take... For the second wavenumber point, points are selected sequentially to construct a new wavenumber vector r'. To determine the optimal interval value N, based on the microplastic spectral database (c groups of samples) constructed in step 1, a partial least squares (PLS) model is established with the transmittance corresponding to r' as input and the microplastic type as output. K-fold cross-validation is performed, and the root mean square error of cross-validation (RMSECV) is used as the evaluation metric to determine the optimal interval value. The transmittance vector corresponding to the minimum RMSECV is denoted as x' = [x'1, x'2, ..., x']. q ].

[0015] k-fold cross-validation refers to randomly and uniformly dividing c groups of sample data into k subsets (denoted as C1, C2, ..., C6). k ), sequentially select one subset as the test set and the remaining subsets as the training set, and build a k-fold PLS model. The expression for calculating RMSECV for k-fold cross-validation is:

[0016]

[0017] When the test set is C i When (i=1,2,…,k), the model output is denoted as (u1) i u2 i ,…,u c / k i ) T The correct output is denoted as (d1) i ,d2 i ,…,d c / k i ) T u j i With d j i Both represent m-dimensional column vectors.

[0018] (2) The process of extracting feature spectra based on the CARS algorithm is as follows:

[0019] ① Initialize weights: Based on the transmittance x' determined in (1), initialize the absolute value weights of the transmittance regression coefficients in x' through PLS modeling;

[0020] ② Feature selection loop: In each iteration, the following operations are performed:

[0021] a. EDF sampling: For the current weight, the Exponentially Decreasing Function (EDF) method is used to remove transmittance variables with small absolute values.

[0022] b. ARS Sampling: For the EDF sampling results, the feature transmittance is selected using the Adaptive Reweighted Sampling (ARS) algorithm based on Monte Carlo sampling (MCS).

[0023] c. PLS modeling: Using the feature transmittance after ARS sampling as input and the microplastic type as output, PLS modeling is adopted.

[0024] ③ Output results: Set the number of iterations, repeat steps a, b, and c, and use the transmittance combination corresponding to the model with the smallest RMSECV as the final feature transmittance variable.

[0025] The characteristic wavenumber corresponding to the minimum value of RMSECV is denoted as r”=[r'1',r'2',…,r' n The feature transmittance vector is denoted as x”=[x'1',x'2',…,x']. n ').

[0026] Preferably, in step 3, the characteristic spectral transmittance vector x”=[x'1',x'2',…,x' n The formula for performing the standard normal transformation is as follows:

[0027] x=(x”-μ) / σ

[0028] In the formula, x is the transformed characteristic transmittance vector, μ is the mean of x”, and σ is the standard deviation of x”.

[0029] After transformation, the mean of x is 0 and the standard deviation is 1. After processing, the characteristic transmittance of different samples exhibits similar distribution characteristics, facilitating comparison and analysis.

[0030] Preferably, in step 4, the ratio of the number of samples in the feature training set to the number of samples in the feature test set is generally 1:2-8:2, depending on the amount of sample data. ANN models include single-layer neural network (SNN) models, multi-layer neural network (MNN) models, and convolutional neural network (CNN) models.

[0031] The hyperparameters of the ANN model are optimized using a genetic algorithm, and the steps are as follows:

[0032] ①Based on the feature wavenumber r”, feature training set and feature test set are selected from the original training set and test set respectively, and then standard normal transformation is performed;

[0033] ②Use the feature training set to build an ANN model;

[0034] ③ Determine the hyperparameters of the model that need to be optimized based on the ANN model, use them as optimization variables z, and give them a range of values ​​[A,B];

[0035] ④ Establish an optimization model In the formula, f(z) represents the correct recognition rate of the ANN model in the feature test set when the hyperparameter is z;

[0036] ⑤ Use a genetic algorithm to solve the optimization model in ④ to obtain the optimal solution z. * ;

[0037] ⑥ z * As a hyperparameter in the ANN model, an optimized ANN model is built using the feature training set.

[0038] Preferably, the performance evaluation metrics for the ANN model in step 5 include Root Mean Squared Error (RMSE) and Accuracy. RMSE is used to evaluate the accuracy of the ANN model on both the training and test sets, while Accuracy is used to evaluate its recognition performance. The calculation formulas for each metric are given below, using the training set as an example:

[0039]

[0040] In the formula, Y=(y1,y2,…,y a ) T Let D = (d1, d2, ..., dn) represent the model output for a samples in the training set. a ) T y represents the correct output of the training set. i With d i Both represent m-dimensional column vectors;

[0041]

[0042] In the formula, TP and FP represent the number of correctly identified samples and the number of incorrectly identified samples in the training set, respectively.

[0043] Compared with the prior art, the beneficial effects of the present invention are:

[0044] This invention provides a method for extracting and efficiently and accurately identifying infrared feature spectra of microplastics. It combines the EIS and CARS algorithms to construct the EIS-CARS algorithm, which effectively avoids traversing and sampling the entire spectral dataset while preserving the main information of the original spectra, reducing computational load and improving the extraction efficiency of feature spectra. A genetic algorithm is used to optimize the hyperparameters of the ANN model, effectively preventing overfitting and improving its generalization ability. Attached Figure Description

[0045] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0046] Figure 2 The average mid-infrared spectrum of the microplastic samples in the examples.

[0047] Figure 3 The variation of RMSECV with interval value during the EIS process in the example embodiment.

[0048] Figure 4 The following are examples of RMSECV changes (a) and extracted feature spectra (b) during the CARS algorithm sampling process after EIS implementation.

[0049] Figure 5 The following are examples of RMSECV changes (a) and extracted feature spectra during the CARS algorithm sampling process in this embodiment (b).

[0050] Figure 6 Example 1: ANN convergence rate (a) and changes in the correct microplastic recognition rate on the feature test set (b).

[0051] Figure 7 The correct recognition rate of the test samples in the example Detailed Implementation

[0052] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0053] This invention proposes a method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics. The method flowchart is as follows: Figure 1 The specific steps include:

[0054] Step 1: Obtain infrared spectral data of microplastic samples and construct a microplastic spectral database;

[0055] Step 2: Extract feature spectra from the infrared spectral data in Step 1 using a progressive two-step method that combines equal-interval sampling (EIS) with competitive adaptive reweighted sampling (CARS) algorithms;

[0056] Step 3: Perform a standard normal transformation on the transmittance of the characteristic spectra from Step 2;

[0057] Step 4: Using the transmittance of the transformed infrared spectrum from Step 3 as input and the microplastic type as output, construct a feature training set and a feature test set. Train the ANN model using the feature training set, and then use a genetic algorithm to optimize the hyperparameters of the model using the feature test set, forming the final microplastic recognition ANN model.

[0058] Step 5: For the microplastic to be identified, obtain spectral data through Step 1, obtain the characteristic transmittance of the microplastic to be identified through the characteristic wavenumber extracted in Step 2, perform a standard normal transformation on it through Step 3, and input the transformed infrared spectral transmittance into the microplastic identification ANN model to give the type of microplastic and realize the identification of microplastic.

[0059] In step 1, infrared spectral data of microplastics are obtained using a Fourier transform infrared spectrometer. The number of scans for each microplastic sample can be determined through literature review and preliminary experiments.

[0060] Step 2 involves two steps to extract the feature spectrum: (1) EIS of the original spectrum; (2) Extracting the feature spectrum based on the CARS algorithm.

[0061] (1) The EIS process for the original spectrum is as follows:

[0062] The abscissa wavenumber vector of the original infrared spectrum is The transmittance vector on the ordinate is in This represents a wavenumber point on the horizontal axis of the infrared spectrum, in cm. -1 ; Indicates the spectrum with The corresponding transmittance, %. When taking points at the wavenumber interval, assume the interval value is N, that is, the first wavenumber point. Starting from point N, take... For the second wavenumber point, points are selected sequentially to construct a new wavenumber vector r'. To determine the optimal interval value N, based on the microplastic spectral database (c groups of samples) constructed in step 1, a partial least squares (PLS) model is established with the transmittance corresponding to r' as input and the microplastic type as output. K-fold cross-validation is performed, and the root mean square error of cross-validation (RMSECV) is used as the evaluation index to determine the optimal interval value. The interval transmittance vector corresponding to the minimum RMSECV is denoted as x' = [x'1, x'2, ..., x']. q ].

[0063] k-fold cross-validation refers to randomly and uniformly dividing c groups of sample data into k subsets (denoted as C1, C2, ..., C6). k), sequentially select one subset as the test set and the remaining subsets as the training set, and build a k-fold PLS model. The expression for calculating RMSECV for k-fold cross-validation is:

[0064]

[0065] When the test set is C i When (i=1,2,…,k), the model output is denoted as (u1) i u2 i ,…,u c / k i ) T The correct output is denoted as (d1) i ,d2 i ,…,d c / k i ) T u j i With d j i Both represent m-dimensional column vectors.

[0066] (2) The process of extracting feature spectra based on the CARS algorithm is as follows:

[0067] ① Initialize weights: Based on the transmittance x' determined in (1), initialize the absolute value weights of the transmittance regression coefficients in x' through PLS modeling;

[0068] ② Feature selection loop: In each iteration, the following operations are performed:

[0069] a. EDF sampling: For the current weight, the transmittance variable with a small absolute value is removed using the exponential decay method (EDF).

[0070] b. ARS Sampling: Based on the EDF sampling results, the adaptive reweighting algorithm (ARS) based on Monte Carlo sampling (MCS) is used to select the feature transmittance.

[0071] c. PLS modeling: Using the feature transmittance after ARS sampling as input and the microplastic type as output, PLS modeling is adopted.

[0072] ③ Output results: Set the number of iterations, repeat steps a, b, and c, and use the transmittance combination corresponding to the model with the smallest RMSECV as the final feature transmittance variable.

[0073] The characteristic wavenumber corresponding to the minimum value of RMSECV is denoted as r”=[r'1',r'2',…,r' n The feature transmittance vector is denoted as x”=[x'1',x'2',…,x']. n ').

[0074] In step 3, the characteristic spectral transmittance vector x”=[x'1',x'2',…,x' n The formula for performing the standard normal transformation is as follows:

[0075] x=(x”-μ) / σ

[0076] In the formula, x is the transformed characteristic transmittance vector, μ is the mean of x”, and σ is the standard deviation of x”.

[0077] After transformation, the mean of x is 0 and the standard deviation is 1. After processing, the characteristic transmittance of different samples exhibits similar distribution characteristics, facilitating comparison and analysis.

[0078] In step 4, the ratio of the number of samples in the feature training set to the number of samples in the feature test set is generally 1:2 to 8:2, depending on the amount of sample data. ANN models include single-layer neural network (SNN) models, multi-layer neural network (MNN) models, and convolutional neural network (CNN) models.

[0079] The hyperparameters of the ANN model are optimized using a genetic algorithm, and the steps are as follows:

[0080] ①Based on the feature wavenumber r”, feature training set and feature test set are selected from the original training set and test set respectively, and then standard normal transformation is performed;

[0081] ②Use the feature training set to build an ANN model;

[0082] ③ Determine the hyperparameters of the model that need to be optimized based on the ANN model, use them as optimization variables z, and give them a range of values ​​[A,B];

[0083] ④ Establish an optimization model In the formula, f(z) represents the correct recognition rate of the ANN model in the feature test set when the hyperparameter is z;

[0084] ⑤ Use a genetic algorithm to solve the optimization model in ④ to obtain the optimal solution z*;

[0085] ⑥ z * As a hyperparameter in the ANN model, an optimized ANN model is built using the feature training set.

[0086] In step 5, the performance evaluation metrics for the ANN model include root mean square error (RMSE) and accuracy. RMSE is used to evaluate the accuracy of the ANN model on both the training and test sets, while accuracy is used to evaluate its recognition performance. The calculation formulas for each metric are given below, using the training set as an example:

[0087]

[0088] In the formula, Y=(y1,y2,…,y a ) T Let D = (d1, d2, ..., dn) represent the model output for a samples in the training set. a ) T y represents the correct output of the training set. i With d i Both represent m-dimensional column vectors;

[0089]

[0090] In the formula, TP and FP represent the number of correctly identified samples and the number of incorrectly identified samples in the training set, respectively.

[0091] Specific

[0092] Obtain 20 common typical microplastic types, including acrylonitrile butadiene styrene (ABS), ethylene vinyl acetate (EVA), polybutylene terephthalate (PBT), polycarbonate (PC), polycaprolactone (PCL), polyethylene (PE), polyethersulfone (PES), polyethylene terephthalate (PET), polylactic acid (PLA), polymethyl methacrylate (PMMA), polyoxymethylene (POM), polypropylene (PP), polyphenylene oxyether (PPO), and polyphenylene sulfide (P2P). Sulfide (PPS), polystyrene (PS), polytetrafluoroethylene (PTFE), polyvinyl alcohol (PVA), polyvinyl chloride (PVC), styrene-butadiene block copolymer (SBS), and thermoplastic polyurethane (TPU).

[0093] Microplastics were scanned using a Fourier transform infrared spectrometer (Nicolet iS10, Nicolet Corporation, USA), with a test range of 4000 cm⁻¹. -1 ~400cm -1 The resolution is 4cm. -1 Each sample was scanned nine times, resulting in 180 sets of spectral data. The average mid-infrared spectrum of the 20 microplastic samples is shown below. Figure 2 .

[0094] (1) Feature spectral extraction

[0095] Three sets of transmittance data from nine mid-infrared spectral analyses of each microplastic were randomly selected, resulting in a training set of 60 data points. The microplastic type was used as the output. The remaining 120 sets of transmittance data and microplastic types constituted the test set. In this test set, the transmittance of each sample was used as input, and the microplastic type as the output. To calculate the error, the 20 microplastics were sequentially numbered, and the microplastic type names were converted into type numbers. Taking EVA as an example, the microplastic type output is a 20×1 column vector, where the second component is 1, and the remaining components are 0.

[0096] In less than 500cm -1 and greater than 3500cm -1 Within the wavenumber range, the spectra of each sample contain relatively little information, while in the 500-3500 cm⁻¹ range... -1 The wavenumber range contains rich spectral information, therefore only within 500-3500 cm⁻¹... -1 Feature spectra were extracted within the wavenumber range. The wavenumber interval within this spectral range is 1.4464 cm⁻¹. -1 A total of 2074 wavenumber points were obtained, and the EIS-CARS algorithm was used to extract the characteristic spectra.

[0097] Step 1: Equal-Interval Sampling (EIS)

[0098] Within 2074 wavenumber points, with intervals N ranging from 0 to 19, 20 new wavenumber vectors r'0, r'1, r'2, ..., r' are constructed respectively. 19 Based on the training and test sets (a total of 180 samples), PLS models were built using the transmittance corresponding to the new wavenumber vector as input and the microplastic type as output. Five-fold cross-validation was used at different interval values. The RMSECV results are shown below. Figure 3 When N=8, RMSECV is minimized (0.0954), therefore the optimal interval is 8, and a total of 231 transmittance data points are obtained.

[0099] Step 2: Extract feature spectra based on CARS algorithm

[0100] Based on the 231 transmittance data points obtained in step 1, feature transmittance was extracted using the CARS algorithm. MCS was used for 60 samplings, and the changes in RMSECV during this process are as follows: Figure 4 (a) At the 22nd sampling, the RMSECV was the lowest (0.0904), and a total of 44 characteristic wavenumbers were extracted. The characteristic infrared spectrum is as follows: Figure 4 (b). During the extraction of feature spectra, the time taken for equally spaced sampling was 4.190s, and the time taken for feature spectrum extraction based on CARS was 5.906s, for a total of 10.096s (computer configuration as follows: Intel Core i7-6700 CPU@3.40GHz, RAM 8GB).

[0101] To verify the efficiency and modeling accuracy of the EIS-CARS algorithm in extracting feature spectra, the original infrared spectrum (2074 wavenumber points) was used as a basis, and the CARS algorithm was directly employed to extract feature spectra. The changes in RMSECV during 60 MCS samplings are shown below. Figure 5 (a) At the 25th sampling, the RMSECV was the lowest (0.0933), and a total of 130 characteristic wavenumbers were extracted. Its characteristic infrared spectrum is as follows: Figure 5 (b) The time taken to extract the feature spectrum based on the CARS algorithm was 367.570s, which is 36.41 times that of the EIS-CARS algorithm.

[0102] Both the EIS-CARS and CARS algorithms achieve an RMSECV of less than 0.095 under optimal sampling conditions, with the EIS-CARS algorithm showing a lower RMSECV. This indicates that the feature spectra extracted by both methods are relatively reliable. However, the CARS algorithm extracts nearly three times the number of feature wavenumbers compared to the EIS-CARS algorithm. This means that using the feature spectra extracted by the EIS-CARS algorithm is more efficient for ANN modeling. Therefore, considering both RMSECV and computational efficiency, the EIS-CARS algorithm is superior in extracting feature spectra.

[0103] Regarding the feature wavenumber distribution, the CARS algorithm extracts a relatively uniform feature wavenumber distribution, especially in the 500-1700 cm⁻¹ range. -1 The distribution is dense. The feature wavenumbers extracted by the EIS-CARS algorithm are mainly concentrated in the 500-774 cm⁻¹ range. -1 956-1178cm -1 1216-1738cm -1 2960-3013cm -1 Four wavenumber bands, and Figure 2In comparison, these four wavenumber bands exhibit a rich variety and quantity of functional groups, containing the most information. Comparison with the functional groups in Table 1 reveals that these four wavenumber bands primarily contain aromatic CH deformation, CH deformation, COC bending, aromatic CO, aromatic S=O, COC stretching, CO bending, CSC, CF, and C-Cl functional groups, all characteristic functional groups of the 20 microplastics. Therefore, the spectral information extracted by the EIS-CARS algorithm is more targeted and specific.

[0104] Table 1 Comparison of Feature Spectra Extracted by Different Algorithms

[0105]

[0106]

[0107] (2) Artificial Neural Network Modeling

[0108] Feature transmittance was obtained by extracting 44 feature wavenumbers using the EIS-CARS algorithm, and then subjected to a standard normal transformation. The transformed transmittance was used as the input to an ANN model (input layer n = 44 nodes), and the output consisted of 20 types of microplastics (output layer m = 20 nodes). Feature training and test sets were constructed. Microplastic recognition models were established using SNN, MNN, and CNN, respectively.

[0109] ① Single-layer neural network modeling (SNN)

[0110] The parameter values ​​of the ANN model have a certain impact on model performance and training speed. During the SNN learning process, the initial values ​​of the weight matrix are generated by Matlab's random number generator, with the seed being a key parameter. The mini-batch size (bsize) also affects the model's performance. During parameter tuning, the two most important parameters of the Adam algorithm, β² and ε, have a minimal impact on model performance; recommended values ​​are β² = 0.999 and ε = 10. -8 α and β1 need to be optimized. Therefore, the parameters that need to be optimized in SNN include seed, bsize, α, and β1.

[0111] The SNN parameter optimization model is as follows: maxf(z), where z = (seed, bsize, α, β1) TThe parameters were set to the following ranges: 1 ≤ seed ≤ 100, 1 ≤ 60 / bsize ≤ 6, 0 < α ≤ 1, 0 < β1 ≤ 1, with seed and 60 / bsize being integers. f(z) represents the correct recognition rate of the SNN on the feature test set after training, when the parameter z is taken. In the genetic algorithm, the population size was set to 20, the number of generations to 100, the crossover probability to 0.8, and the mutation probability to 0.3. The optimal parameter combination was calculated to be: seed = 54, bsize = 12, α = 0.7718, β1 = 0.0662. An SNN model was built using these parameters. The feature training set was used to test the convergence rate of the ANN model, and the feature test set was used to test the correct recognition rate of the SNN model for microplastics after hyperparameter optimization. The results are shown in […]. Figure 6 With 200 training epochs, the optimized ANN model achieved a 100% accuracy rate in recognizing microplastics.

[0112] ② Multilayer Neural Network Modeling (MNN)

[0113] Assuming that MNN has two hidden layers, compared with the parameters that SNN needs to optimize, the parameters that MNN needs to optimize include the number of nodes in the hidden layers, m1 and m2. Therefore, the parameters that MNN needs to optimize include six parameters: seed, m1, m2, bsize, α, and β1.

[0114] The parameter optimization model for a fully connected MNN is as follows: maxf(z), where z = (seed, m1, m2, bsize, α, β1) T The parameters are set to the following ranges: 1 ≤ seed ≤ 100, 1 ≤ m1 ≤ 200, 1 ≤ m2 ≤ 200, 1 ≤ 60 / bsize ≤ 6, 0 < α ≤ 1, 0 < β1 ≤ 1, and seed, m1, m2, and 60 / bsize are integers. f(z) represents the correct recognition rate of the MNN on the feature test set after training when the parameter z is taken. Due to the increased number of parameters to be optimized, the genetic algorithm is set to a population size of 40, a generation count of 200, a crossover probability of 0.8, and a mutation probability of 0.3. The optimal parameter combination is calculated as: seed = 7, m1 = 94, m2 = 73, bsize = 15, α = 0.0297, β1 = 0.0934. Using these parameters, an MNN model is established. The model convergence rate and the correct recognition rate of microplastics on the feature test set are shown in [link to relevant documentation]. Figure 6 With 200 training epochs, the optimized MNN model achieved a 100% accuracy rate in recognizing microplastics.

[0115] ③ Convolutional Neural Network (CNN) Modeling

[0116] Assuming a CNN has one convolutional layer, the output of which passes through a sigmoid function before entering a pooling layer using 2×1 average pooling, followed by one hidden layer and an output layer. The number of convolutional filters is set to nF, the kernel size to ks, and the number of hidden layer nodes to m. Therefore, the CNN needs to optimize seven parameters: seed, nF, ks, m, bsize, α, and β1.

[0117] The CNN parameter optimization model is as follows: maxf(z), where z = (seed, nF, ks, m, bsize, α, β1) T The value range of each parameter is set as 1≤seed≤100, 1≤nF≤20. 1≤m≤200, 1≤60 / bsize≤6, 0<α≤1, 0<β1≤1, and seed, nF, m and 60 / bsize are rounded to integers. f(z) represents the correct recognition rate of the CNN on the feature test set after training, when the parameter z is set. In the genetic algorithm, the population size is set to 40, the number of generations to 300, the crossover probability to 0.8, and the mutation probability to 0.3. The optimal parameter combination is calculated as: seed = 89, nF = 19, ks = 21, m = 59, bsize = 20, α = 0.0861, β1 = 0.7929. A CNN model is built using this set of parameters. The model convergence rate and the correct recognition rate of microplastics on the feature test set are shown in [link to data]. Figure 6 With 200 training epochs, the CNN model, after hyperparameter optimization, achieved a 100% accuracy rate in recognizing microplastics.

[0118] (3) Identification of microplastics

[0119] The aim is to identify the aforementioned 20 typical microplastics. The microplastics were scanned using a Fourier transform infrared spectroscopy (Nicolet iS10, Nicolet Inc., USA), with a testing range of 4000 cm⁻¹. -1 ~400cm -1 The resolution is 4cm. -1 Each sample was scanned six times, resulting in 120 sets of spectral data. Feature transmittance was obtained using the feature wavenumbers extracted based on the EIS-CARS algorithm (Table 1), and then subjected to a standard normal transformation. The transformed transmittance (considered as the sample to be tested) was used as input to the aforementioned ANN model. Running the model yielded the corresponding microplastic types, which were then compared with the actual microplastic types to calculate the correct identification rate. Figure 7 ).

[0120] Table 2 shows the RMSE and correct recognition rate of different ANN models on the feature training set and feature test set, as well as the correct recognition rate for the test samples. After optimization, the SNN model achieved a correct recognition rate of 98.33% for the test samples, while the MNN and CNN models achieved a correct recognition rate of 99.17%, all demonstrating excellent recognition performance. Among them, the MNN and CNN models showed better recognition performance.

[0121] Table 2 shows the performance of the ANN model in recognizing microplastics.

[0122]

[0123] The above specific embodiments are merely several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics, characterized in that, include: Step 1: Obtain infrared spectral data of microplastic samples and construct a microplastic spectral database; Step 2: Extract feature spectra from the infrared spectral data in Step 1 using a progressive two-step method that combines Equal Interval Sampling (EIS) and Competitive Adaptive Reweighted Sampling (CARS) algorithms. Step 3: Perform a standard normal transformation on the transmittance of the characteristic spectra from Step 2; Step 4: Using the transmittance of the transformed infrared spectrum from Step 3 as input and the microplastic type as output, construct a feature training set and a feature test set. The feature training set is used to train the Artificial Neural Network (ANN) model, and a genetic algorithm is used to optimize the hyperparameters of the model through the feature test set, forming the final microplastic recognition ANN model. Step 5: For the microplastic to be identified, obtain spectral data through Step 1, obtain the characteristic transmittance of the microplastic to be identified through the characteristic wavenumber extracted in Step 2, perform a standard normal transformation on it through Step 3, and input the transformed infrared spectral transmittance into the microplastic identification ANN model to give the type of microplastic and realize the identification of microplastic.

2. The method for extracting and efficiently and accurately identifying infrared features of microplastics according to claim 1, characterized in that: In step 1, infrared spectral data of microplastics are obtained using a Fourier transform infrared spectrometer. The number of scans for each microplastic sample can be determined through literature review and preliminary experiments.

3. The method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics according to claim 1, characterized in that: Step 2 involves two steps to extract the feature spectrum: (1) EIS of the original spectrum; (2) Extracting the feature spectrum based on the CARS algorithm. (1) The EIS process for the original spectrum is as follows: The abscissa wavenumber vector of the original infrared spectrum is The transmittance vector on the ordinate is in This represents a wavenumber point on the horizontal axis of the infrared spectrum, in cm. -1 ; Indicates the spectrum with The corresponding transmittance, %. When taking points at the wavenumber interval, assume the interval value is N, that is, the first wavenumber point. Starting from point N, take... For the second wavenumber point, points are selected sequentially to construct a new wavenumber vector r'. To determine the optimal interval value N, based on the microplastic spectral database (c groups of samples) constructed in step 1, a partial least squares (PLS) model is established with the transmittance corresponding to r' as input and the microplastic type as output. K-fold cross-validation is performed, and the root mean square error of cross-validation (RMSECV) is used as the evaluation metric to determine the optimal interval value. The transmittance vector corresponding to the minimum RMSECV is denoted as x' = [x'1, x'2, ..., x']. q ]. k-fold cross-validation refers to randomly and uniformly dividing c sets of sample data into k subsets (denoted as C1, C2, ..., C5). k ), sequentially select one subset as the test set and the remaining subsets as the training set, and build a k-fold PLS model. The expression for calculating RMSECV for k-fold cross-validation is: When the test set is C i When (i=1,2,…,k), the model output is denoted as (u1) i u2 i ,…,u c / k i ) T The correct output is denoted as (d1) i ,d2 i ,…,d c / k i ) T u j i With d j i Both represent m-dimensional column vectors. (2) The process of extracting feature spectra based on the CARS algorithm is as follows: ① Initialize weights: Based on the transmittance x' determined in (1), initialize the absolute value weights of the transmittance regression coefficients in x' through PLS modeling; ② Feature selection loop: In each iteration, the following operations are performed: a. EDF sampling: For the current weights, the Exponentially Decreasing Function (EDF) method is used to remove transmittance variables with small absolute values. b. ARS Sampling: For the EDF sampling results, the feature transmittance is selected using the Adaptive Reweighted Sampling (ARS) algorithm based on Monte Carlo sampling (MCS). c. PLS modeling: Using the feature transmittance after ARS sampling as input and the microplastic type as output, PLS modeling is adopted. ③ Output results: Set the number of iterations, repeat steps a, b, and c, and use the transmittance combination corresponding to the model with the smallest RMSECV as the final feature transmittance variable. The characteristic wavenumber corresponding to the minimum value of RMSECV is denoted as r″=[r″1,r″2,…,r″ n The feature transmittance vector is denoted as x″ = [x″1, x″2, ..., x″]. n ].

4. The method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics according to claim 1, characterized in that: In step 3, the characteristic spectral transmittance vector x”=[x″1,x″2,…,x″ n The formula for performing the standard normal transformation is as follows: x=(x”-μ) / σ In the formula, x is the transformed characteristic transmittance vector, μ is the mean of x”, and σ is the standard deviation of x”. After transformation, the mean of x is 0 and the standard deviation is 1. After processing, the characteristic transmittance of different samples exhibits similar distribution characteristics, facilitating comparison and analysis.

5. The method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics according to claim 1, characterized in that: In step 4, the ratio of the number of samples in the feature training set to the number of samples in the feature test set is generally 1:2 to 8:2, depending on the amount of sample data. ANN models include single-layer neural network (SNN) models, multi-layer neural network (MNN) models, and convolutional neural network (CNN) models. The hyperparameters of the ANN model are optimized using a genetic algorithm, and the steps are as follows: ①Based on the feature wavenumber r”, feature training set and feature test set are selected from the original training set and test set respectively, and then standard normal transformation is performed; ②Use the feature training set to build an ANN model; ③ Determine the hyperparameters of the model that need to be optimized based on the ANN model, use them as optimization variables z, and give them a range of values ​​[A,B]; ④ Establish an optimization model In the formula, f(z) represents the correct recognition rate of the ANN model in the feature test set when the hyperparameter is z; ⑤ Use a genetic algorithm to solve the optimization model in ④ to obtain the optimal solution z. * ; ⑥ z * As a hyperparameter in the ANN model, an optimized ANN model is built using the feature training set.

6. The method for extracting and efficiently and accurately identifying infrared characteristic spectra of microplastics according to claim 1, characterized in that: Step 5 evaluates the recognition performance of the ANN model using the Root Mean Squared Error (RMSE) and accuracy. RMSE is used to evaluate the accuracy of the ANN model on both the training and test sets, while accuracy is used to evaluate its recognition performance. The calculation formulas for each indicator are given below, using the training set as an example: In the formula, Y=(y1,y2,…,y a ) T Let D = (d1, d2, ..., dn) represent the model output for a samples in the training set. a ) T y represents the correct output of the training set. i With d i Both represent m-dimensional column vectors; In the formula, TP and FP represent the number of correctly identified samples and the number of incorrectly identified samples in the training set, respectively.