Dimensionality Reduction Method and System for Near-Infrared Spectroscopy Quantitative Analysis Based on Random Projection Algorithm

Through the combination of Gaussian random projection algorithm and two-dimensional convolutional neural network, the problem of time-consuming and labor-intensive wavelength selection and nonlinear relationships in quantitative analysis of near-infrared spectral is solved, and fast and accurate quantitative analysis of near-infrared spectral is achieved.

CN114676792BActive Publication Date: 2025-07-18EAST CHINA UNIV OF SCI & TECH
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210385752.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-13
Publication Date
2025-07-18
Estimated Expiration
2042-04-13

AI Technical Summary

Technical Problem

The existing near-infrared spectral quantitative analysis technology requires cumbersome wavelength selection process, which makes modeling time-consuming and laborious, and traditional methods are difficult to reflect the nonlinear relationship between the spectrum and the oil product attributes to be analyzed.

Method used

The Gaussian random projection algorithm is used to reduce the dimensionality, combined with the artificial neural network model, and the spectral matrix is reduced by the Gaussian random projection algorithm, and a prediction model is established using a two-dimensional convolutional neural network to avoid wavelength selection and improve modeling efficiency.

Benefits of technology

Fast and accurate quantitative analysis of near-infrared spectral spectroscopy is achieved, reducing modeling complexity and time, and can effectively reflect the nonlinear relationship between the spectrum and the oil product attributes to be analyzed, improving the stability and accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114676792B_ABST
    Figure CN114676792B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of near-infrared modeling data processing, and more specifically, to a near-infrared spectrum quantitative analysis dimensionality reduction method and system based on a random projection algorithm. The present invention includes: Step S1, obtaining near-infrared spectra x val samples and corresponding physical and chemical property values y val as a sample set; Step S2, dividing the sample set into a calibration set and a validation set, and calculating the average spectrum x avg ; Step S3, respectively preprocessing the near-infrared spectra x val and the average spectrum x avg to obtain a spectral matrix X val and an average spectral matrix X avg ; Step S4, performing random dimensionality reduction projection on the spectral matrix X val based on the Gaussian random projection algorithm to obtain a dimensionality-reduced spectral matrix X valRed ; Step S5, establishing an artificial neural network prediction model; Step S6, using the validation set to test the model; Step S7, performing quantitative analysis on the input near-infrared spectra and outputting corresponding predicted physical and chemical property values. The present invention does not require wavelength selection of the spectra, reduces the modeling difficulty, and shortens the modeling time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of near-infrared modeling data processing, and more specifically, to a near-infrared spectroscopy quantitative analysis dimensionality reduction method and system based on a random projection algorithm. Background Art

[0002] Near-infrared analysis technology is a method for analyzing the absorption characteristics of a certain chemical component in a detected sample in the near-infrared spectral region. Through chemometric multivariate calibration methods, qualitative and quantitative analysis of samples is carried out relying on the subtle differences in spectral information between samples. Due to the influence of unstable factors such as spectral instrument noise and external environment changes, the signal-to-noise ratio of some bands in the near-infrared spectrum is low and the spectral quality is poor. These bands will cause model instability, and there is multiple correlation between the sample spectral wavelengths, and there is redundant information in the spectral information, making the calculation of the near-infrared analysis model complex.

[0003] Therefore, wavelength selection is often required when establishing a near-infrared analysis model. At present, the methods of wavelength selection include the correlation coefficient method, genetic algorithm, simulated annealing algorithm, interval partial least squares method, etc.

[0004] However, the process of wavelength selection is very cumbersome and is the most time-consuming and laborious process before modeling.

[0005] At present, the commonly used methods for near-infrared modeling mainly include multiple linear regression, partial least squares method, artificial neural network, support vector machine, etc.

[0006] The partial least squares method is one of the most commonly used modeling methods for near-infrared modeling. It can effectively reduce the dimension, extract the effective information of the independent variable matrix, reflect the linear relationship between the near-infrared spectral wave numbers and the properties of the oil products to be analyzed, and the modeling is reliable and accurate. However, the partial least squares method cannot effectively reflect the non-linear relationship between the near-infrared spectrum and the properties of the oil products to be analyzed.

[0007] Therefore, it is urgent to improve and solve the above deficiencies of the existing near-infrared spectroscopy quantitative analysis technology. Summary of the Invention

[0008] The purpose of the present invention is to provide a near-infrared spectroscopy quantitative analysis dimensionality reduction method and system based on a random projection algorithm, and solve the problem that the existing technology is time-consuming and laborious for near-infrared analysis and requires wavelength selection.

[0009] To achieve the above purpose, the present invention provides a near-infrared spectroscopy quantitative analysis dimensionality reduction method based on a random projection algorithm, including the following steps:

[0010] Step S1, obtaining the near-infrared spectrum x val samples and the corresponding physical and chemical property values y val as a sample set;

[0011] Step S2: Divide the sample set into a calibration set and a validation set, and calculate the average spectrum x based on the near-infrared spectra of the calibration set val Calculate the average spectrum x avg ;

[0012] Step S3: Preprocess the near-infrared spectra x of the calibration set to obtain a spectral matrix X val , and preprocess the average spectrum x of the calibration set to obtain an average spectral matrix X val ; avg ; avg ;

[0013] Step S4: Perform random dimensionality reduction projection on the spectral matrix X of the calibration set based on the Gaussian random projection algorithm to obtain a dimensionality-reduced spectral matrix X val ; valRed ;

[0014] Step S5: Establish an artificial neural network prediction model based on the dimensionality-reduced spectral matrix X valRed ;

[0015] Step S6: Use the validation set to test the artificial neural network prediction model established in Step S5

[0016] Step S7: Based on the artificial neural network prediction model after being tested in Step S6, perform quantitative analysis on the input near-infrared spectra and output the corresponding predicted values of physicochemical properties

[0017] In one embodiment, in Step S2, the average spectrum x avg corresponds to the following expression:

[0018]

[0019] where n is the number of near-infrared spectra, and x vali is the i-th spectrum

[0020] In one embodiment, in Step S2, dividing the sample set into a calibration set and a validation set further includes:

[0021] Use the K-S algorithm based on Euclidean distance or the SPXY algorithm based on property variables to select m spectra from the sample set as the calibration set, and use the remaining samples as the validation set

[0022] In one embodiment, the preprocessing methods in Step S3 include: first derivative, second derivative, and maximum-minimum normalization

[0023] In one embodiment, Step S3 further includes:

[0024] Preprocess the near-infrared spectra x valPerform multiple preprocessings simultaneously to obtain the spectral matrix X val ;

[0025] For the average spectrum x avg Perform multiple preprocessings simultaneously to obtain the average spectral matrix X avg .

[0026] In one embodiment, step S4 further includes:

[0027] Step S41: Obtain the Gaussian random projection transition matrix P according to the average spectral matrix X avg ;

[0028] Step S42: Based on the Gaussian random projection transition matrix P, perform random dimensionality reduction projection on the spectral matrix X val to obtain the spectral matrix X after dimensionality reduction valRed .

[0029] In one embodiment, step S41 further includes:

[0030] Perform random dimensionality reduction projection on the average spectral matrix X of p wavelength points avg to obtain the average spectral matrix X after dimensionality reduction of q wavelength points avgRed ;

[0031] According to the expression X avgRed = P * X avg , solve for the Gaussian random projection transition matrix P

[0032] In one embodiment, the average spectral matrix X avg and the average spectral matrix X avgRed satisfy the following inequality:

[0033] (1 - eps) ||X avg - X avgRed || 2 < ||X avg - X avgRed || 2 < (1 + eps) ||X avg - X avgRed || 2 ;

[0034] The p wavelength points and the q wavelength points after dimensionality reduction satisfy the following inequality:

[0035]

[0036] where eps is the dimensionality reduction error

[0037] In one embodiment, the artificial neural network prediction model is a two-dimensional convolutional prediction model

[0038] Step S5 further includes:

[0039] Step S51: Import the dimension-reduced spectral matrix into the input layer of the two-dimensional convolution. After calculations through two convolutional layers, weights and activation functions, and one pooling layer, it is passed to the output layer after calculations through multiple convolutional layers and pooling layers;

[0040] Step S52: Compare the predicted value obtained from the output layer with the sample expected value. If there is an error between the two, return to Step S51 to adjust the weights until the difference between the predicted value and the sample expected value reaches the first threshold.

[0041] In one embodiment, Step S6 further includes:

[0042] Use the validation set to test the artificial neural network prediction model established in Step S5, and calculate the prediction standard deviation Rmsep. The corresponding expression is:

[0043]

[0044] where m is the number of spectra in the validation set, y i,actual1 is the measured value of the i-th spectrum in the validation set, and y i,predicted1 is the predicted value of the i-th spectrum in the validation set.

[0045] In one embodiment, Step S6 further includes:

[0046] Use the calibration set to perform cross-validation on the artificial neural network prediction model established in Step S5, and calculate the cross-validation standard deviation Rmsecv. The corresponding expression is:

[0047]

[0048] where n is the number of spectra in the calibration set, y i,actual2 is the measured value of the i-th spectrum in the calibration set, and y i,predicted2 is the predicted value of the i-th spectrum in the calibration set.

[0049] To achieve the above object, the present invention provides a near-infrared spectroscopy quantitative analysis dimension reduction system based on a random projection algorithm, including:

[0050] A memory for storing instructions executable by a processor;

[0051] A processor for executing the instructions to implement the method as described in any one of the above.

[0052] To achieve the above object, the present invention provides a computer-readable medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the method as described in any one of the above is executed.

[0053] A dimensionality reduction method and system for near-infrared spectroscopy quantitative analysis based on the random projection algorithm provided by the present invention uses Gaussian random projection for dimensionality reduction, does not require wavelength selection of the spectrum, reduces the modeling difficulty, shortens the modeling time, and can perform concise and rapid modeling for near-infrared analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] The above and other features, properties, and advantages of the present invention will become more apparent through the following description in conjunction with the drawings and embodiments. In the drawings, the same reference numerals always represent the same features, where:

[0055] Figure 1 Discloses a flowchart of a dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to an embodiment of the present invention;

[0056] Figure 2 Discloses an original sample spectrogram according to an embodiment of the present invention;

[0057] Figure 3 Discloses a schematic diagram of a dimensionality reduction system for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the invention and are not used to limit the invention.

[0059] Aiming at the deficiencies of existing near-infrared technologies, the present invention provides a dimensionality reduction method and system for near-infrared spectroscopy quantitative analysis based on the random projection algorithm, which can be widely applied to industries such as petrochemical, agriculture, and food.

[0060] Figure 1 Discloses a flowchart of a dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to an embodiment of the present invention. As Figure 1 shown, the dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm proposed by the present invention specifically includes the following steps:

[0061] Step S1, obtain the near-infrared spectrum x val samples and the corresponding physical and chemical property values y val as a sample set;

[0062] Step S2, divide the sample set into a calibration set and a validation set, and calculate the average spectrum x val according to the near-infrared spectrum x avg of the calibration set;

[0063] Step S3. Preprocess the near-infrared spectrum x of the calibration set val to obtain the spectral matrix X val , and preprocess the average spectrum x avg of the calibration set to obtain the average spectral matrix X avg ;

[0064] Step S4. Perform random dimensionality reduction projection on the spectral matrix X of the calibration set based on the Gaussian random projection algorithm val to obtain the dimensionality-reduced spectral matrix X valRed ;

[0065] Step S5. Establish an artificial neural network prediction model based on the dimensionality-reduced spectral matrix X valRed ;

[0066] Step S6. Use the validation set to test the artificial neural network prediction model established in Step S5

[0067] Step S7. Based on the artificial neural network prediction model tested in Step S6, perform quantitative analysis on the input near-infrared spectrum and output the corresponding predicted values of physical and chemical properties.

[0068] These steps will be described in detail below. It should be understood that within the scope of the present invention, the above technical features of the present invention and the technical features specifically described below (such as in the embodiments) can be combined with each other and are interrelated, thereby constituting a preferred technical solution.

[0069] Step S1. Obtain the near-infrared spectrum x val samples and the corresponding physical and chemical property values y val as the sample set.

[0070] Obtain the near-infrared spectra of a batch of samples and their corresponding physical and chemical property values for modeling.

[0071] The physical and chemical properties include physical properties and chemical properties.

[0072] Optionally, the physical properties include, but are not limited to, density, freezing point, viscosity, distillation range, etc.;

[0073] Optionally, the chemical properties include component composition, element content, etc.

[0074] Optionally, the physical and chemical property values can be measured by laboratory methods as the measured values.

[0075] In this embodiment, a batch of near-infrared spectra x val and physical and chemical property values y val are obtained for modeling;

[0076] wherein, the near-infrared spectrum x val includes n spectra xvali , where i ranges from 1 to n, and x vali represents the i-th spectrum;

[0077] The i-th spectrum x vali The physicochemical property value y corresponding to the label attribute vali ;

[0078] The near-infrared spectrum has p wavelength points.

[0079] Step S2: Divide the sample set into a calibration set and a validation set, and calculate the average spectrum x according to the near-infrared spectra x of the calibration set val Calculate the average spectrum x avg .

[0080] Furthermore, the average spectrum x of the n near-infrared spectra avg , and the calculation formula is as shown in (1):

[0081]

[0082] where n is the number of near-infrared spectra, and x vali is the i-th spectrum.

[0083] Furthermore, use the K-S algorithm based on Euclidean distance or the SPXY algorithm based on property variables to select m spectra with strong representativeness from the sample set as the calibration set, and use the remaining samples as the validation set.

[0084] The principle of the K-S (Kennard-Stone) algorithm is to regard all samples as candidate samples for the training set and select samples from them in turn. First, select the two samples with the farthest Euclidean distance to enter the training set. By calculating the Euclidean distance from each of the remaining samples to each known sample in the training set, find the two samples with the farthest and nearest distances to the selected samples, and select these two samples into the training set. Repeat the above steps until the number of samples reaches the requirement.

[0085] The SPXY (sample set partitioning based on joint x-y distance) algorithm is developed on the basis of the K-S algorithm. The SPXY algorithm takes both the x variable and the y variable into account when calculating the distance between samples.

[0086] Step S3: Preprocess the near-infrared spectra x of the calibration set val to obtain the spectral matrix X val , and preprocess the average spectrum x of the calibration set avg to obtain the average spectral matrix X avg .

[0087] Near-infrared spectroscopy is vulnerable to interference from some environmental factors during measurement, generating noise and making the spectrum contain some unusable wavelength points. Therefore, in this step, the spectrum is preprocessed.

[0088] Spectrum preprocessing can amplify the effective information of the spectrum, filter out the noise information in the spectrum, thereby reducing the modeling complexity and improving the robustness of the model.

[0089] The preprocessing methods include, but are not limited to, first derivative, second derivative, and maximum-minimum normalization.

[0090] Furthermore, the step S3 further includes:

[0091] Performing multiple preprocessings on the near-infrared spectrum x val simultaneously to obtain the spectral matrix X val ;

[0092] Performing multiple preprocessings on the average spectrum x avg simultaneously to obtain the average spectral matrix X avg .

[0093] Ordinary preprocessing methods only use one of the above. In this embodiment, the above three preprocessing methods are used simultaneously, so that a sample spectrum x val and the average spectrum x avg become the spectral matrix X val containing three preprocessing methods and the average spectral matrix X avg .

[0094] Performing multiple preprocessings on a sample spectrum, and combining the data obtained from the multiple preprocessings into one matrix.

[0095] Step S4: Based on the Gaussian random projection algorithm, perform random dimensionality reduction projection on the spectral matrix X val of the calibration set to obtain the dimensionality-reduced spectral matrix X valRed .

[0096] Perform Gaussian random projection on each sample matrix to a low-dimensional matrix.

[0097] Furthermore, the step S4 further includes:

[0098] Step S41: Obtain the Gaussian random projection transition matrix P according to the average spectral matrix X avg ;

[0099] Step S42: Based on the Gaussian random projection transition matrix P, perform random dimensionality reduction projection on the near-infrared spectrum x val to obtain the dimensionality-reduced spectral matrix X valRed .

[0100] More specifically, the Gaussian random projection transition matrix P in step S41 is obtained as follows:

[0101] The average spectral matrix X of p wavelength points avg is subjected to random dimensionality reduction projection to obtain the reduced average spectral matrix X of q wavelength points avgRed , and the reduced average spectral matrix X avgRed satisfies the following formula (2):

[0102] X avgRed = P * X avg (2)

[0103] where P is the Gaussian random projection transition matrix (subject to Gaussian distribution) for dimensionality reduction using the average spectral matrix, and the transition matrix P can be solved according to formula (2).

[0104] where the average spectral matrix X avg and the reduced average spectral matrix X avgRed satisfy the following inequality (3):

[0105] (1 - eps) ||X avg - X avgRed || 2 < ||X avg - X avgRed || 2 < (1 + eps) ||X avg - X avgRed || 2 (3)

[0106] The p wavelength points and the reduced q wavelength points satisfy the following inequality (4):

[0107]

[0108] eps is the dimensionality reduction error.

[0109] In this embodiment, eps uses the default value 0.1.

[0110] Random dimensionality reduction is performed on each spectrum, and its calculation is as shown in formula (5):

[0111] X valRedi = P * X vali (5)

[0112] where: X valRedi is the i-th element of the reduced spectral matrix X valRed and X vali is the i-th element of the spectral matrix X val to be reduced.

[0113] The dimensionality reduction from p wavelength points to q wavelength points can be achieved through formula (5), and the characteristics of the data are greatly preserved.

[0114] Step S5. Based on the dimension-reduced spectral matrix X valRed , an artificial neural network prediction model is established.

[0115] Using the calibration set, an artificial neural network prediction model is established. The artificial neural network model includes, but is not limited to, a multi-layer perceptron prediction model, a backpropagation neural network prediction model, a convolutional neural network prediction model, etc.

[0116] A multi-layer perceptron (MLP, Multilayer Perceptron) is a feedforward artificial neural network model that maps multiple input data sets to a single output data set.

[0117] A convolutional neural network has extremely strong feature extraction capabilities. Using a convolutional neural network can achieve any non-linear mapping from input to output, thereby overcoming the problem that the partial least squares method cannot reflect non-linear relationships.

[0118] In this embodiment, a two-dimensional convolutional neural network is used to establish an analysis model to obtain the predicted value of quantitative analysis.

[0119] Thus, step S5 further includes:

[0120] Step S51. Import the dimension-reduced spectral matrix X valRedi into the input layer of the two-dimensional convolution. After calculations through two convolutional layers, weights and activation functions, and one pooling layer, and after calculations through multiple convolutional layers and pooling layers, it is transmitted to the output layer;

[0121] Step S52. Compare the predicted value obtained from the output layer with the sample expected value. If there is an error between the two, return to step S51 to adjust the weights, and continuously adjust the weights until the difference between the predicted value and the sample expected value reaches the first threshold.

[0122] Among them, the first threshold is a preset minimum value or extremely small value.

[0123] Step S6. Use the validation set to test the artificial neural network prediction model established in step S5.

[0124] Use the validation set to perform a validation set test on the artificial neural network prediction model established in step S5. Calculate the prediction standard deviation Rmsep according to formula (6), and the corresponding expression is:

[0125]

[0126] Among them, m is the number of spectra in the validation set, y i,actual1The measured value of the i-th spectrum in the validation set, y i,predicted1 is the predicted value of the i-th spectrum in the validation set.

[0127] The artificial neural network prediction model established in step S5 is cross-checked using the calibration set, and the cross-validation standard deviation Rmsecv is calculated according to formula (7), and the corresponding expression is:

[0128]

[0129] where n is the number of spectra in the calibration set, y i,actual2 is the measured value of the i-th spectrum in the calibration set, y i,predicted2 is the predicted value of the i-th spectrum in the calibration set.

[0130] The model is tested using the true measured values of the spectra in the validation set to determine whether the prediction accuracy requirement is met. If not, return to step S1 to restart the modeling process. If it is met, enter step S7 for model application.

[0131] Step S7: Based on the artificial neural network prediction model after being tested in step S6, perform quantitative analysis on the input near-infrared spectrum and output the corresponding predicted values of physical and chemical properties.

[0132] The model after being tested in step S6 already meets the prediction accuracy requirement and can be used for quantitative analysis of actual infrared spectra. Importing the near-infrared spectrum data to be analyzed into the model can output the corresponding predicted values of physical and chemical properties.

[0133] Although the above methods are illustrated and described as a series of actions to simplify the explanation, it should be understood and appreciated that these methods are not limited by the order of the actions, because according to one or more embodiments, some actions may occur in a different order and / or occur concurrently with other actions that are illustrated and described herein or that are not illustrated and described herein but are understood by those skilled in the art.

[0134] The near-infrared spectrum quantitative analysis dimensionality reduction method based on the random projection algorithm proposed by the present invention uses Gaussian random projection for dimensionality reduction without wavelength selection, solving the problems of the prior art that a large amount of manual experience intervention and a large amount of time are required in the wavelength selection process, simplifying the complexity.

[0135] At the same time, by virtue of the excellent non-linear fitting ability of the neural network model and the powerful feature extraction ability of the two-dimensional convolutional neural network, the difficulty that the traditional method cannot fit non-linearity is overcome.

[0136] Therefore, the near-infrared spectrum quantitative analysis dimensionality reduction method based on the random projection algorithm proposed by the present invention can, while ensuring the modeling accuracy, greatly reduce the modeling time and establish a fast and accurate near-infrared quantitative analysis model.

[0137] The following uses the data of the near-infrared spectrum of aviation kerosene and the corresponding laboratory analysis reports as the experimental objects to specifically and detailedly describe the near-infrared spectrum quantitative analysis dimensionality reduction method based on the random projection algorithm proposed by the present invention.

[0138] Step S1: Obtain a batch of near-infrared spectra and their corresponding physical and chemical property values.

[0139] The specific process is as follows:

[0140] Obtain the spectral diagrams of 52 samples by a near-infrared spectrometer. The number of wavelength points p of the spectrum is 2074, and the sample spectra are corresponded with the physical and chemical property values in the laboratory analysis report. The original sample spectra are as Figure 2 shown.

[0141] Step S2: Use the K-S algorithm based on the Euclidean distance to select 36 spectra with strong spectral representativeness from the sample set as the calibration set, and calculate the average spectrum x avg .

[0142] Step S3: Perform three kinds of preprocessing on the near-infrared spectrum simultaneously to obtain a spectral matrix.

[0143] The specific process is as follows:

[0144] Convert the near-infrared spectrum into a row matrix, and at the same time use maximum-minimum preprocessing, first-order derivative and second-order derivative preprocessing to form a spectral row matrix with 3 rows.

[0145] Step S4: Perform random Gaussian random projection on the spectral matrix X val to obtain the spectral matrix X valRed of each sample after dimensionality reduction;

[0146] The specific process is as follows:

[0147] Calculate the average spectrum X val through the spectral matrix X avg , obtain the transition matrix P of the Gaussian distribution through Gaussian random projection, and get the spectral matrix X avlRed after dimensionality reduction, that is, a spectral row matrix with 3 rows and 941 wavelength points after dimensionality reduction.

[0148] Step S5: Establish a two-dimensional convolutional prediction model;

[0149] The specific process is as follows:

[0150] Import the input spectral matrix sample X valRedi after dimensionality reduction into the input layer of the two-dimensional convolution, and through convolution, pooling and output layers, fit the physical and chemical property (density), and the generated two-dimensional convolutional prediction model is called the "model of this method".

[0151] Specifically, as a control, the non-dimensionality-reduced spectral matrix X val is imported into the input layer of the two-dimensional convolution. After convolution, pooling, and the output layer, the density is fitted to generate a two-dimensional convolution prediction model for control, which is called the "non-dimensionality-reduced model".

[0152] Step S6: Verify the established model;

[0153] The specific process is as follows:

[0154] Import the 16 verification samples of the verification set into the "model of the present method" and the "non-dimensionality-reduced model" respectively, and predict the density of the 16 samples respectively.

[0155] The measured values and predicted values of the verification sets of the "model of the present method" and the "non-dimensionality-reduced model" are shown in Table 1, and the training time required for generating the model and the model evaluation results are shown in Table 2.

[0156] Table 1 Results of measured and predicted values of sample density in the verification sets of the two models (unit: kg / m 3 )

[0157]

[0158]

[0159] Table 2 Comparison table of model evaluation data of the two models

[0160] Method Training time (s) <![CDATA[Rmsep (kg / m 3 )]]> <![CDATA[Rmsecv(kg / m 3 )]]> The model of this method 65 2.90 2.38 The non-dimensionality-reduced model 446 2.80 2.22

[0161] As can be seen from Table 1 and Table 2, the model of the present method saves 85% of the training time compared with the non-dimensionality-reduced model, and only increases the prediction deviation by no more than 10%.

[0162] Figure 3 Fig. shows a block diagram of a near-infrared spectroscopy quantitative analysis dimensionality reduction system based on a random projection algorithm according to an embodiment of the present invention. The near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm may include an internal communication bus 301, a processor 302, a read-only memory (ROM) 303, a random access memory (RAM) 304, a communication port 305, and a hard disk 307. The internal communication bus 301 can realize data communication between components of the near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm. The processor 302 can make judgments and give prompts. In some embodiments, the processor 302 may be composed of one or more processors.

[0163] The communication port 305 can enable data transmission and communication between the near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm and external input / output devices. In some embodiments, the near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm can send and receive information and data from the network through the communication port 305. In some embodiments, the near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm can perform data transmission and communication with external input / output devices in a wired form through the input / output terminal 306.

[0164] The near-infrared spectroscopy quantitative analysis dimensionality reduction system based on the random projection algorithm may further include program storage units and data storage units in different forms, such as a hard disk 307, a read-only memory (ROM) 303, and a random access memory (RAM) 304, which can store various data files used for computer processing and / or communication, as well as possible program instructions executed by the processor 302. The processor 302 executes these instructions to implement the main part of the method. The results processed by the processor 302 are transmitted to an external output device through the communication port 305 and displayed on the user interface of the output device.

[0165] For example, the implementation process file of the above-mentioned near-infrared spectroscopy quantitative analysis dimensionality reduction method based on the random projection algorithm can be a computer program, stored in the hard disk 307 and can be recorded into the processor 302 for execution to implement the method of this application.

[0166] When the implementation process file of the near-infrared spectroscopy quantitative analysis dimensionality reduction method based on the random projection algorithm is a computer program, it can also be stored in a computer-readable storage medium as an article of manufacture. For example, the computer-readable storage medium may include, but is not limited to, magnetic storage devices (such as hard disks, floppy disks, magnetic strips), optical discs (such as compact discs (CDs), digital versatile discs (DVDs)), smart cards, and flash memory devices (such as electrically erasable programmable read-only memories (EPROMs), cards, sticks, key drives). In addition, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media (and / or storage media) that can store, contain, and / or carry code and / or instructions and / or data.

[0167] Compared with the prior art, the present invention provides a near-infrared spectroscopy quantitative analysis dimensionality reduction method and system based on the random projection algorithm, which specifically has the following beneficial effects:

[0168] 1) It neither requires complex wavelength selection nor sacrifices the reliability and accuracy of the model. At the same time, it reduces the requirements for technicians and the complexity of modeling, promoting the improvement and enhancement of near-infrared quantitative modeling methods.

[0169] 2) Use the Gaussian random projection method to reduce the dimension. While ensuring sufficient spectral information is extracted, it avoids information loss, reduces the spatial dimension, and thus decreases the amount of data processed for modeling.

[0170] 3) Simultaneously use multiple typical preprocessing methods, effectively reducing spectral noise and laying a solid foundation for subsequent modeling.

[0171] As shown in this application and the claims, unless the context clearly indicates an exceptional situation, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified steps and elements, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0172] Those skilled in the art will understand that information, signals, and data can be represented using any of a variety of different technologies and techniques. For example, the data, instructions, commands, information, signals, bits, symbols, and chips described throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or optical particles, or any combination thereof.

[0173] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, boxes, modules, circuits, and steps are described above in terms of their functional form. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Skilled artisans can implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as causing a departure from the scope of the present invention.

[0174] The various illustrative logical modules and circuits described in conjunction with the embodiments disclosed herein can be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. The general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors cooperating with a DSP core, or any other such configuration.

[0175] The steps of the methods or algorithms described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read from, and write to, the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0176] In one or more exemplary embodiments, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. The computer-readable medium includes both a computer storage medium and a communication medium including any medium that facilitates transfer of a computer program from one place to another. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable medium can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a web site, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable medium.

[0177] The above embodiments are provided to those skilled in the art to implement or use the present invention. Those skilled in the art can make various modifications or variations to the above embodiments without departing from the inventive concept of the present invention. Therefore, the protection scope of the present invention is not limited by the above embodiments, but should be the maximum scope that conforms to the innovative features mentioned in the claims.

Claims

1. A dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on a random projection algorithm, characterized in that, Including the following steps: Step S1: Obtain the near-infrared spectrum x val and the corresponding physical and chemical property value y val as a sample set; Step S2: Divide the sample set into a calibration set and a validation set, and calculate the average spectrum x based on the near-infrared spectra x of the calibration set val Calculate the average spectrum x avg ; Step S3: Preprocess the near-infrared spectrum x of the calibration set val to obtain the spectral matrix X val , and preprocess the average spectrum x of the calibration set avg to obtain the average spectral matrix X avg ; Step S4: Based on the Gaussian random projection algorithm, perform random dimensionality reduction projection on the spectral matrix X of the calibration set val to obtain the spectral matrix X after dimensionality reduction valRed ; Step S5: Based on the dimension-reduced spectral matrix X valRed , establish an artificial neural network prediction model; Step S6: Use the validation set to test the artificial neural network prediction model established in step S5; Step S7: Based on the artificial neural network prediction model after being tested in step S6, perform quantitative analysis on the input near-infrared spectrum and output the corresponding predicted value of the physicochemical property; The step S4 further includes: Step S41: Obtain the Gaussian random projection transition matrix P according to the average spectral matrix X avg ; Step S42: Based on the Gaussian random projection transition matrix P, perform random dimensionality reduction projection on the spectral matrix X val to obtain the spectral matrix X after dimensionality reduction valRed .

2. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, wherein In the said step S2, the average spectrum x avg The corresponding expression is: where n is the number of near-infrared spectra, and x vali is the i-th spectrum.

3. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, characterized in that, In the step S2, dividing the sample set into a calibration set and a validation set further includes: Using the K-S algorithm based on Euclidean distance or the SPXY algorithm based on property variables, select m spectra from the sample set as the calibration set, and use the remaining samples as the validation set.

4. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, wherein The preprocessing methods in the step S3 include: first derivative, second derivative, and maximum-minimum normalization.

5. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, characterized in that, The step S3 further includes: For near-infrared spectrum x val Perform multiple pre-treatments simultaneously to obtain the spectral matrix X val ; For the average spectrum x avg Perform multiple preprocessings simultaneously to obtain the average spectrum matrix X avg .

6. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, characterized in that The step S41 further includes: The average spectral matrix X of p wavelength points avg is subjected to random dimensionality reduction projection to obtain the reduced-dimension average spectral matrix X of q wavelength points avgRed ; According to the expression X avgRed = P * X avg , solve for the Gaussian random projection transition matrix P.

7. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 6, characterized in that Average spectral matrix X avg and the average spectral matrix X after dimensionality reduction avgRed satisfy the following inequality: The p wavelength points and the q wavelength points after dimensionality reduction satisfy the following inequality: where eps is the dimensionality reduction error.

8. The near-infrared spectroscopy quantitative analysis dimensionality reduction method based on the random projection algorithm according to claim 1, characterized in that The artificial neural network prediction model in the step S5 further includes: Multi-layer perceptron prediction model, backpropagation neural network prediction model, and convolutional neural network prediction model.

9. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, characterized in that The artificial neural network prediction model is a two-dimensional convolutional prediction model; The step S5 further includes: Step S51: Import the spectrum matrix after dimensionality reduction into the input layer of the two-dimensional convolution, perform calculations through two convolutional layers, weights and activation functions, and one pooling layer, and pass it to the output layer after multiple calculations of convolutional layers and pooling layers; Step S52: Compare the predicted value obtained from the output layer with the sample expected value. If there is an error between the two, return to step S51 to adjust the weights until the difference between the predicted value and the sample expected value reaches the first threshold.

10. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, wherein The step S6 further includes: Use the validation set to perform validation set test on the artificial neural network prediction model established in step S5, and calculate the prediction standard deviation Rmsep. The corresponding expression is: where m is the number of spectra in the validation set, and \(y_{ i,actual1}\) is the measured value of the \(i\)-th spectrum in the validation set, and \(y_{ i,predicted1}\) is the predicted value of the \(i\)-th spectrum in the validation set.​​​​ 11. The dimensionality reduction method for near-infrared spectroscopy quantitative analysis based on the random projection algorithm according to claim 1, wherein The step S6 further includes: Use the calibration set to perform cross-validation on the artificial neural network prediction model established in step S5, and calculate the cross-validation standard deviation Rmsecv. The corresponding expression is: where n is the number of spectra in the calibration set, and y i,actual2 is the measured value of the i-th spectrum in the calibration set, and y i,predicted2 is the predicted value of the i-th spectrum in the calibration set.

12. A near-infrared spectrum quantitative analysis dimensionality reduction system based on a random projection algorithm, including: A memory for storing instructions executable by a processor; A processor for executing the instructions to implement the method according to any one of claims 1-11.

13. A computer-readable medium having computer instructions stored thereon, wherein when the computer instructions are executed by a processor, the method according to any one of claims 1-11 is executed.

Citation Information

Patent Citations

  • Cream brilliant blue pigment detection method and device and storage medium

    CN114112992A