Method and device for improving spectrum data quality based on spectrum reconstruction
By constructing a spectral reconstruction model, the low-quality spectral data of the portable spectrometer is improved to high-quality spectral data, which solves the problem of insufficient analysis performance of the portable spectrometer and improves the analysis accuracy of the miniaturized spectrometer.
Patent Information
- Application Number
- CN202410104606.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-25
- Publication Date
- 2025-07-25
AI Technical Summary
The portable spectrometer has low resolution, narrow spectral range and poor signal-to-noise ratio, resulting in insufficient qualitative and quantitative analysis performance during real-time, on-site, and in-situ analysis.
By building a model based on spectral reconstruction, using low-quality spectral data and high-quality spectral data of the calibration sample, data dimensionality reduction processing is carried out, spectral reconstruction model is established, and spectral data quality is improved.
Reconstructing low-quality spectral data into high-quality spectral data improves the analytical performance of miniaturized spectrometers and improves the accuracy of qualitative and quantitative analysis.
Smart Images

Figure CN120369646A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spectrometers, and particularly to a method and device for improving the quality of spectral data based on spectral reconstruction. Background Art
[0002] Spectral technology is an analytical technique that uses the optical properties of substances to determine their chemical composition or structure. With the miniaturization of spectral measurement devices, portable spectrometers have shown advantages such as low cost, convenient operation, and no need for complex sample preparation, and are widely used in scenarios of real-time, on-site, and in-situ analysis. Compared with bench-top spectrometers, portable spectrometers have low resolution, narrow spectral range, and poor signal-to-noise ratio, making it difficult to effectively characterize some key chemical composition or structural information, resulting in problems of insufficient qualitative analysis performance and quantitative analysis performance when using portable spectrometers for real-time, on-site, and in-situ analysis. Summary of the Invention
[0003] In view of the above-mentioned drawbacks of the prior art, the present application provides a method and device for improving the quality of spectral data based on spectral reconstruction to improve the analysis performance of various miniaturized spectrometers including portable spectrometers.
[0004] The first aspect of the present application provides a method for improving the quality of spectral data based on spectral reconstruction, including:
[0005] Obtaining first low-quality spectral data and first high-quality spectral data corresponding to a calibration sample;
[0006] Performing data dimensionality reduction processing on the first low-quality spectral data to obtain low-dimensional feature data of the first low-quality spectral data;
[0007] Constructing a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data;
[0008] Processing second low-quality spectral data of a sample to be measured based on the spectral reconstruction model to obtain second high-quality spectral data, and the second high-quality spectral data is used as a basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
[0009] Optionally, the constructing a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data includes:
[0010] Obtaining a model to be trained;
[0011] Determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data;
[0012] In the case that the model loss does not meet the convergence condition, update the model to be trained according to the model loss until the model loss meets the convergence condition;
[0013] In the case that the model loss meets the convergence condition, determine the model to be trained as the spectral reconstruction model.
[0014] Optionally, determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data includes:
[0015] Compare the differences in the low-dimensional feature data of a pair of calibration samples output by the siamese neural network of the model to be trained to obtain a comparison loss;
[0016] Use the decoder of the model to be trained to process the low-dimensional feature data to obtain the simulated hyperspectral data corresponding to the low-dimensional feature data, and determine the reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data;
[0017] Use the fully connected layer of the model to be trained to process the low-dimensional feature data to obtain the predicted label corresponding to the low-dimensional feature data, and determine the label loss according to the difference between the predicted label and the true label corresponding to the calibration sample;
[0018] Combine the comparison loss, the reconstruction loss and the label loss to obtain the model loss of the model to be trained.
[0019] Optionally, obtaining the first low-quality spectral data includes:
[0020] Detect the calibration sample by at least one of a portable spectrometer, a handheld spectrometer, and a micro spectrometer to obtain the first low-quality spectral data, where the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, and laser-induced breakdown spectral data;
[0021] Obtain the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample.
[0022] Optionally, obtaining the first high-quality spectral data includes:
[0023] Detect the calibration sample by a bench spectrometer to obtain the first high-quality spectral data, where the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, and X-ray fluorescence spectral data.
[0024] The second aspect of the present application provides a device for improving the quality of spectral data based on spectral reconstruction, including:
[0025] An acquisition unit for acquiring first low-quality spectral data and first high-quality spectral data corresponding to a calibration sample;
[0026] A dimensionality reduction unit for performing dimensionality reduction processing on the first low-quality spectral data to obtain low-dimensional feature data of the first low-quality spectral data;
[0027] A construction unit for constructing a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data;
[0028] A processing unit for processing second low-quality spectral data of a sample to be measured based on the spectral reconstruction model to obtain second high-quality spectral data, and using the second high-quality spectral data as a basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
[0029] Optionally, when the construction unit constructs a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data, it is specifically used for:
[0030] Obtaining a model to be trained;
[0031] Determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data;
[0032] In the case where the model loss does not meet the convergence condition, updating the model to be trained according to the model loss until the model loss meets the convergence condition;
[0033] In the case where the model loss meets the convergence condition, determining the model to be trained as a spectral reconstruction model.
[0034] Optionally, when the construction unit determines the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data, it is specifically used for:
[0035] Comparing the differences in the low-dimensional feature data of a pair of calibration samples output by the siamese neural network of the model to be trained to obtain a comparison loss;
[0036] Processing the low-dimensional feature data by using the decoder of the model to be trained to obtain simulated hyperspectral data corresponding to the low-dimensional feature data, and determining a reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data;
[0037] Processing the low-dimensional feature data by using the fully connected layer of the model to be trained to obtain a predicted label corresponding to the low-dimensional feature data, and determining a label loss according to the difference between the predicted label and the true label corresponding to the calibration sample;
[0038] Combine the contrast loss, the reconstruction loss, and the label loss to obtain the model loss of the model to be trained.
[0039] Optionally, when the obtaining unit obtains the first low-quality spectral data, it is specifically used for:
[0040] Detect the calibration sample by at least one of a portable spectrometer, a handheld spectrometer, and a micro spectrometer to obtain the first low-quality spectral data, where the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, and laser-induced breakdown spectral data;
[0041] Obtain the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample.
[0042] Optionally, when the obtaining unit obtains the first high-quality spectral data, it is specifically used for:
[0043] Detect the calibration sample by a bench-top spectrometer to obtain the first high-quality spectral data, where the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, and X-ray fluorescence spectral data.
[0044] The beneficial effects of this application are as follows:
[0045] This solution constructs a spectral reconstruction model using the low-quality spectral data and high-quality spectral data of the calibration sample. Then, when obtaining the low-quality spectral data of the sample to be measured, it can use this spectral reconstruction model to reconstruct the low-quality spectral data of the sample to be measured into high-quality spectral data, improving the analysis performance of the miniaturized spectrometer by improving the quality of the spectral data. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to the provided drawings.
[0047] Figure 1 It is a schematic structural diagram of a miniaturized spectrometer provided by an embodiment of this application;
[0048] Figure 2 It is a schematic structural diagram of a bench-top spectrometer provided by an embodiment of this application;
[0049] Figure 3A flowchart of a method for improving spectral data quality based on spectral reconstruction provided by an embodiment of the present application;
[0050] Figure 4 A framework diagram of a spectral reconstruction model based on a twin neural network provided by an embodiment of the present application;
[0051] Figure 5 An effect evaluation diagram for improving the qualitative analysis performance of a portable spectrometer based on spectral reconstruction provided by an embodiment of the present application;
[0052] Figure 6 A structural schematic diagram of a device for improving spectral data quality based on spectral reconstruction provided by an embodiment of the present application. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] Spectral signal quality such as spectral resolution, range, and signal-to-noise ratio is a key factor affecting the qualitative / quantitative analysis performance of spectral technology. At present, the research on reconstructing low-quality spectral data into high-quality spectral data to improve the signal quality of portable spectrometers is still blank. The main reasons are as follows: Spectral acquisition depends on portable spectrometers and bench spectrometers, and the equipment cost is high; constructing a mapping relationship model between low-quality spectral data and high-quality spectral data to solve the ill-posed problem, that is, there are countless solutions and the solutions are unstable, and a large number of calibration samples are required to ensure the rationality of the solutions.
[0055] Spectral reconstruction and spectral super-resolution provide a reference for solving the ill-posed problem. These methods usually adopt a deep neural network structure to extract features from images (such as RGB images, hyperspectral images) to establish a mapping relationship model between the image and the hyperspectral image, thereby improving the spectral resolution and spatial resolution of the image. However, spectral reconstruction and spectral super-resolution are currently only applied to generate reflection spectra in the visible light range (400–700nm), and have not been applied to generate other types of spectra, such as Raman spectra, near-infrared spectra, and laser-induced breakdown spectra. In addition, hyperspectral images generated based on spectral reconstruction and spectral super-resolution are rarely applied to qualitative / quantitative analysis.
[0056] In view of the above problems, the present application provides a method for improving spectral data quality based on spectral reconstruction.
[0057] In this application, the low-quality spectral data of a sample can be obtained by detecting the sample with any miniaturized spectrometer.
[0058] As an example, the structure of the miniaturized spectrometer can be as Figure 1 shown. It can be seen that the miniaturized spectrometer can include a light source 1 and a spectral signal collection device 3. Below the right of the spectral signal collection device 3 is the sample 2.
[0059] Among them, the light source can be a laser source, a near-infrared light source or a visible light source.
[0060] The working principle of this spectrometer is that the light source 1 irradiates on the surface of the sample 2, the light interacts with the sample surface, is absorbed, reflected or transmitted by the sample surface, generating a spectral signal. The spectral signal is collected by the spectral signal collection device to obtain the low-quality spectral data of the sample.
[0061] Optionally, a collection device for transmitted light can also be provided below the sample to collect the spectral signal generated by the transmission of the sample.
[0062] The high-quality spectral data of the sample can be obtained by detecting the sample with any bench-top spectrometer.
[0063] As an example, the structure of the bench-top spectrometer can be as Figure 2 shown. It can be seen that the bench-top spectrometer includes an energy source 4, a signal source 6, a lens 7, a timing control device (DDG) 8, a signal receiving device 9, a computer 10, and the signal source 6 is generated by the sample 5.
[0064] The working principle of this spectrometer is that the energy source 4 emits a beam of light under the control of the timing control device 8, which is focused on the surface of the sample 5 and interacts with the sample to generate a signal source. The radiation light emitted by the signal source is transmitted to the signal receiving device 9 through the lens 7. The signal receiving device 9 receives the radiation light under the control of the timing control device, converts the radiation light into a spectral signal and then transmits the spectral signal to the computer 10. The computer 10 obtains the first high-quality spectral data according to the spectral signal.
[0065] According to the above devices, an embodiment of this application provides a method for improving the quality of spectral data based on spectral reconstruction. Please refer to Figure 3 , which is the flowchart of this method. This method can include the following steps.
[0066] S301, obtain the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample.
[0067] A calibration sample refers to a sample with a pre-determined true label.
[0068] In this embodiment, any material of the sample that needs to be spectroscopically analyzed. As an example, the sample can be intumescent fire retardant coatings of different brands and batches.
[0069] The true label of the calibration sample can be the category of the calibration sample determined by qualitative analysis of the spectral data of the calibration sample, or the composition and proportion of the calibration sample obtained by quantitative analysis of the spectral data of the calibration sample.
[0070] In S301, the first low-quality spectral data and the first high-quality spectral data of multiple calibration samples can be obtained. Among them, for each calibration sample, there is a corresponding first low-quality spectral data of the calibration sample and a corresponding first high-quality spectral data of the calibration sample.
[0071] Exemplarily, the calibration samples can include calibration sample a and calibration sample b. The spectral data obtained in S301 can include the first low-quality spectral data a and the first high-quality spectral data a corresponding to calibration sample a, and the first low-quality spectral data b and the first high-quality spectral data b corresponding to calibration sample b.
[0072] Among them, for any calibration sample, the obtaining method of the first low-quality spectral data corresponding to the calibration sample can include:
[0073] Detecting the calibration sample by at least one of a portable spectrometer, a handheld spectrometer, and a micro spectrometer to obtain the first low-quality spectral data, and the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, and laser-induced breakdown spectral data.
[0074] Portable spectrometers, handheld spectrometers, and micro spectrometers have the characteristics of low equipment cost, low resolution, narrow spectral range, low signal-to-noise ratio, and are suitable for rapid on-site in-situ measurement.
[0075] The low-quality spectral data can be obtained based on various molecular spectroscopy and atomic spectroscopy techniques.
[0076] For any calibration sample, the obtaining method of the first high-quality spectral data corresponding to the calibration sample can include:
[0077] Detecting the calibration sample by a bench-top spectrometer to obtain the first high-quality spectral data, and the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, and X-ray fluorescence spectral data.
[0078] The high-quality spectral data can be obtained based on various molecular spectroscopy and atomic spectroscopy techniques.
[0079] The bench-top spectrometer has the characteristics of high equipment cost, high resolution, wide spectral range, high signal-to-noise ratio, and is difficult to be applied to rapid on-site in-situ measurement.
[0080] As an example, the first low-quality spectral data can be Raman spectral data, and the first high-quality spectral data can be laser-induced breakdown spectral data.
[0081] Assume there are n calibration samples, and the first low-quality spectral data of each calibration sample can include l variables. Then, the total first low-quality spectral data of n calibration samples can be represented by X(n*l).
[0082] Correspondingly, the total first high-quality spectral data of n calibration samples can be represented by Y(n*h), where h is the number of variables of a first high-quality spectral data. And, l is less than h. As an example, h can be equal to 423.
[0083] In this embodiment, the number of variables of the spectral data can be understood as the number of wavelengths included in the spectral data. The quality of the spectral data can be characterized by the number of variables of the spectral data. Under the condition that other parameters are the same, it can be considered that the spectral data with more variables has higher quality. Therefore, the method provided in this embodiment is essentially to reconstruct the spectral data with fewer wavelengths into the spectral data with more wavelengths to achieve the effect of improving the quality of the spectral data.
[0084] S302, perform data dimensionality reduction processing on the first low-quality spectral data to obtain the low-dimensional feature data of the first low-quality spectral data.
[0085] For the first low-quality spectral data corresponding to each calibration sample, after dimensionality reduction processing, the number of variables can be reduced from l to d. Correspondingly, the low-dimensional feature data obtained after dimensionality reduction processing of the first low-quality spectral data of n calibration samples can be represented by P(n*d).
[0086] As an example, l can be equal to 260 and d can be equal to 32.
[0087] The above data dimensionality reduction processing can be implemented in various ways.
[0088] As an example, please refer to Figure 4 the framework of the spectral reconstruction model shown. The implementation method of the above data dimensionality reduction processing can be that for each pair of calibration samples, the two first low-quality spectral data corresponding to the pair of calibration samples are input into the encoder, and the encoder outputs the two low-dimensional feature data corresponding to the pair of calibration samples.
[0089] Among them, each pair of calibration samples includes two calibration samples randomly selected from the aforementioned n calibration samples.
[0090] TakeFigure 4 For example, a pair of calibration samples may include calibration samples a and b. After being processed by the encoder, low-dimensional feature data a corresponding to calibration sample a and low-dimensional feature data b corresponding to calibration sample b can be obtained.
[0091] The encoder may have a twin neural network structure (TSR-Net). This structure specifically includes two neural networks that share weights and parameters. Each neural network includes three layers, and from top to bottom, each layer includes 260, 96, and 32 neurons in sequence. When two first low-quality spectral data are input into the encoder, one first low-quality spectral data is input into one neural network, and the other first low-quality spectral data is input into the other neural network.
[0092] Optionally, the first layer of the neural network in the encoder (the layer with 260 neurons) can adopt the Tanh activation function, and the Dropout of the neural network is 0.2. Dropout represents the probability that the activation value of a certain neuron stops working during the forward propagation of the neural network. Setting this parameter can prevent the neural network from overfitting.
[0093] S303. Construct a spectral reconstruction model based on the low-dimensional feature data and the first high-quality spectral data.
[0094] The process of constructing the spectral reconstruction model may include:
[0095] Obtain the model to be trained;
[0096] Determine the model loss of the model to be trained based on the low-dimensional feature data and the first high-quality spectral data;
[0097] In the case where the model loss does not meet the convergence condition, update the model to be trained according to the model loss until the model loss meets the convergence condition;
[0098] In the case where the model loss meets the convergence condition, determine the model to be trained as the spectral reconstruction model.
[0099] Among them, the structure of the model to be trained is the same as the structure of the spectral reconstruction model. The parameters and weights in the model to be trained can be generated by random initialization. Updating the model to be trained during the above training process means updating the parameters and weights of the model to be trained. When the model loss meets the convergence condition, the model to be trained is the spectral reconstruction model to be constructed in this embodiment.
[0100] When updating the model to be trained according to the model loss, the backpropagation algorithm can be used to calculate the model loss to obtain the update amounts of the parameters and weights in the model to be trained, and then update the parameters and weights of the model to be trained one by one according to the update amounts.
[0101] The convergence condition can be that the model loss is less than a certain loss threshold; it can also be that the model loss tends to be stable, and the difference between the model losses of two consecutive iterations is small enough.
[0102] Among them, the model loss of the model to be trained can be determined as follows:
[0103] Compare the low-dimensional feature data differences of a pair of calibrated samples output by the siamese neural network of the model to be trained to obtain a contrast loss;
[0104] Use the decoder of the model to be trained to process the low-dimensional feature data to obtain the simulated hyperspectral data corresponding to the low-dimensional feature data, and determine the reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data;
[0105] Use the fully connected layer of the model to be trained to process the low-dimensional feature data to obtain the predicted labels corresponding to the low-dimensional feature data, and determine the label loss according to the difference between the predicted labels and the true labels corresponding to the calibrated samples;
[0106] Combine the contrast loss, reconstruction loss and label loss to obtain the model loss of the model to be trained.
[0107] The calculation methods of the above losses are described below.
[0108] Contrast loss. For a selected pair of calibrated samples, after using the encoder to extract the two low-dimensional feature data of this pair of calibrated samples, calculate the similarity between the two low-dimensional feature data (which can be expressed as a percentage), and then determine whether this pair of calibrated samples belong to the same category, that is, determine whether the corresponding true labels of the two are the same, or determine whether the similarity between the true labels of the two is greater than a certain threshold. If the true labels are the same or the similarity of the true labels is greater than a certain threshold, it is determined that the two in this pair of calibrated samples belong to the same category of samples. If the true labels are different and the similarity is not greater than a certain threshold, it is determined that the two in this pair of calibrated samples belong to different category of samples.
[0109] If they are samples of the same category, it is hoped that the low-dimensional feature data output by the encoder is as similar as possible. At this time, 100% can be subtracted from the similarity between the two low-dimensional feature data, and the obtained difference is used as the contrast loss; if they are samples of different categories, it is hoped that the low-dimensional feature data output by the encoder is as different as possible. At this time, the similarity between the two low-dimensional feature data can be directly used as the contrast loss.
[0110] Among them, the similarity between the two low-dimensional feature data can specifically be the cosine similarity. The specific calculation method of the cosine similarity can refer to relevant existing technologies and will not be elaborated here.
[0111] Reconstruction loss. When calculating the reconstruction loss, first, the low-dimensional feature data of a calibration sample can be input into the decoder, and the decoder processes the low-dimensional feature data of the calibration sample to obtain the simulated hyperspectral data of the calibration sample.
[0112] The decoder can also include two neural networks that share weights and parameters. When inputting the low-dimensional feature data into the decoder, it can specifically be input into any one of the neural networks. As Figure 4 shown, one neural network in the decoder can include four layers with the number of neurons being 32, 64, 128, and 423 in sequence. Among them, the first two layers (i.e., the fully connected layers of 32 and 64) can be set with Tanh activation functions and Dropout parameters respectively. Dropout can be set to 0.2 or other values. The layer with 423 neurons serves as the output layer of the neural network.
[0113] After obtaining the simulated hyperspectral data, the reconstruction loss can be calculated based on the simulated hyperspectral data and the first hyperspectral data of the same calibration sample.
[0114] Still taking Figure 4 as an example, the low-dimensional feature data of calibration sample b is input into the decoder to obtain the corresponding simulated hyperspectral data b, and then the reconstruction loss is calculated using the first high-quality spectral data b and the simulated hyperspectral data b of the same calibration sample.
[0115] When calculating the reconstruction loss, the mean absolute error value between the first high-quality spectral data b and the simulated hyperspectral data b can be calculated, and the calculation result is used as the reconstruction loss. Therefore, the reconstruction loss can also be called the mean absolute error loss. The calculation method of the mean absolute error value can refer to relevant existing technologies.
[0116] Label loss. When calculating the label loss, as Figure 4 shown, first, the low-dimensional feature data of a calibration sample can be input into the fully connected layer of the model to be trained through a layer normalization process. The number of neurons in this fully connected layer can be the same as the number of variables of the low-dimensional feature data, for example, both are 32.
[0117] This fully connected layer can process the input low-dimensional feature data and output the predicted label corresponding to the low-dimensional feature data.
[0118] When the method of this embodiment is used in a qualitative analysis scenario, the predicted label can reflect the category of the sample corresponding to the input low-dimensional feature data; when used in a quantitative analysis scenario, the predicted label can reflect the components and proportions of the sample corresponding to the input low-dimensional feature data.
[0119] After obtaining the predicted label, the label loss is calculated based on the predicted label and the true label of the same calibration sample.
[0120] TakingFigure 4 For example, the low-dimensional feature data a of the calibration sample a is input into the fully connected layer to obtain the predicted label of the calibration sample a. Then, according to the predicted label and the true label of the calibration sample a, the label loss is calculated.
[0121] When calculating the label loss, the cross-entropy between the predicted label and the true label can be calculated, and the calculation result is used as the label loss. Therefore, the label loss can also be called the cross-entropy loss. The specific calculation method of the cross-entropy can refer to the relevant existing technologies.
[0122] When combining the above three losses, they can be directly added together, and the obtained result is used as the model loss. Or weights can be set for each loss in advance, and the three are weighted and averaged according to the weights, and the obtained result is used as the model loss. As an example, the weight of the alignment loss can be set to 1, the weight of the reconstruction loss can be set to 6, and the weight of the label loss can be set to 9.
[0123] Optionally, the label loss can be calculated using one of a pair of calibration samples, and the reconstruction loss can be calculated using the other of the pair of calibration samples. Figure 4 For example, the low-dimensional feature data a of the calibration sample a is used to calculate the label loss, and the low-dimensional feature data b of the calibration sample b is used to calculate the reconstruction loss.
[0124] S304, Process the second low-quality spectral data of the sample to be measured based on the spectral reconstruction model to obtain the second high-quality spectral data.
[0125] The second high-quality spectral data serves as the basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
[0126] The process of the spectral reconstruction model processing the second low-quality spectral data to obtain the second high-quality spectral data is the same as the process of processing the first low-quality spectral data to obtain the simulated hyperspectral data described above, and will not be elaborated here.
[0127] In some alternative embodiments, before using the spectral data of the calibration sample and the spectral data of the sample to be measured, the spectral data can be preprocessed.
[0128] The preprocessing process includes, but is not limited to, removing the baseline, standardizing, selecting spectral lines, calculating spectral peak areas, etc. The specific implementation process can refer to the relevant existing technologies and will not be elaborated here.
[0129] The method for obtaining the second low-quality spectral data can be the same as that for the first low-quality spectral data.
[0130] In some alternative embodiments, the spectral reconstruction model can also be used for auxiliary analysis. Specifically, after inputting the second low-quality spectral data of the sample to be measured into the spectral reconstruction model, first, the encoder of the spectral reconstruction model can process the second low-quality spectral data into low-dimensional feature data of the sample to be measured. At this time, the low-dimensional feature data can be input into Figure 4 the fully connected layer shown to obtain the predicted label of the sample to be measured, and the predicted label can be used as the result of qualitative analysis and / or as the result of quantitative analysis.
[0131] In some alternative embodiments, the encoder can also process the low-quality spectral data into low-dimensional feature data through techniques such as partial least squares regression and convolutional neural networks, not limited to the above-mentioned siamese neural network.
[0132] The advantage of using a siamese neural network is that for the case of a small number of calibration samples, the sample number can be increased by combining different calibration samples, achieving the effect of data augmentation.
[0133] Correspondingly, the decoder can use a general fully connected neural network, not limited to the above-mentioned siamese neural network structure.
[0134] The method for improving the quality of spectral data based on spectral reconstruction provided in this embodiment has the beneficial effect that:
[0135] This solution constructs a spectral reconstruction model using the low-quality spectral data and high-quality spectral data of calibration samples. Then, when obtaining the low-quality spectral data of the sample to be measured, the spectral reconstruction model can be used to reconstruct the low-quality spectral data of the sample to be measured into high-quality spectral data, improving the analysis performance of miniaturized spectrometers by enhancing the quality of spectral data.
[0136] The following describes the execution process of the method in this embodiment with an example.
[0137] In this example, first, 385 coating samples of 7 brands are randomly sorted and numbered. The samples numbered 50, 100, 150, 200, 250, and 300 are used as calibration samples for constructing the model in sequence, and the samples numbered 85 at the end are used as test samples (equivalent to the aforementioned samples to be measured). The Raman spectra data of the calibration samples are subjected to baseline removal and normalization processing, and the Raman spectra data with 260 variables after preprocessing are used as the aforementioned first low-quality spectral data. The Laser-induced breakdown spectroscopy (LIBS) data of the calibration samples are subjected to baseline removal, spectral line selection, and spectral peak area calculation processing, and the Laser-induced breakdown spectroscopy data with 386 variables after preprocessing are used as the aforementioned first high-quality spectral data.
[0138] Among them, for the Raman spectroscopy data part, a portable Raman system was used for collection. Some parameters of this system are as follows: the working wavelength is 457 nm, the laser power is 15 mW, the wavelength range is 303.9 - 1199.4 cm⁻¹, and the spectral resolution is 3.5 cm⁻¹. For the laser-induced breakdown spectroscopy data part, a bench-top laser-induced breakdown spectroscopy system was used for collection. Some parameters of this system are as follows: Nd:YAG laser, the working wavelength is 1064 nm, the frequency is 2 Hz, the pulse energy is 90 mJ, the wavelength range is 186.9 - 979.2 nm, and the spectral resolution is 0.09 nm.
[0139] After using the first low-quality spectral data and the first high-quality spectral data of the above 300 calibration samples to construct a spectral reconstruction model according to the aforementioned method, 85 test samples were used to test the accuracy of this spectral reconstruction model, and the accuracy curves shown in (a) and (b) of Figure 5 were obtained. As a comparison, this curve shows the accuracy of classifying the second high-quality spectral data reconstructed by the method of the aforementioned embodiment using different algorithms. Among them, PLS-DA represents a classification method implemented based on partial least squares discriminant analysis, SVM represents a classification method implemented based on support vector machines, and TSR-Net represents the classification method proposed in this embodiment. L-RS represents low-quality Raman spectra.
[0140] It can be seen from the accuracy curve that the second high-quality spectral data reconstructed by the method of this embodiment can obtain relatively accurate classification results when used for classification.
[0141] Among them, the accuracy of the classification result can be represented by the similarity between the predicted label on the test sample in the classification result and the known true label of the test sample. The higher the similarity, the higher the accuracy.
[0142] At the same time, for the laser-induced breakdown spectroscopy data reconstructed by the model constructed in this embodiment, when used to establish a conventional model, the accuracy of the conventional model has been significantly improved.
[0143] During the training process of the model, the model loss gradually converges as the number of iterations increases, as shown in (c) of Figure 5 , indicating that the method of this embodiment can successfully construct a spectral reconstruction model without the situation of non-converging loss and unable to complete the training.
[0144] Among the 85 test samples, only two test samples were misclassified by the trained spectral reconstruction model, that is, there are only two test samples where the predicted label is inconsistent with the true label, as shown in (d) of Figure 5 .
[0145] In addition, please refer to Figure 5 In (e) of Figure 5 , the Raman spectral data of the test sample (equivalent to the second low-quality spectral data of the aforementioned sample to be measured) is processed using the trained spectral reconstruction model to obtain the laser-induced breakdown spectral data reconstructed by the model (equivalent to the second high-quality spectral data of the sample to be measured in S304). By comparing the reconstructed laser-induced breakdown spectral data with the true laser-induced breakdown spectral data measured from the test sample, it is found that the reconstructed high-quality spectral data and the measured high-quality spectral data are highly similar, and the mean absolute error (MAE) is only 0.068. This shows that the spectral reconstruction model constructed by the method of this embodiment can use low-quality spectral data to reconstruct simulated high-quality spectral data (i.e., the second high-quality spectral data) that is highly close to the true high-quality spectral data, thereby effectively improving the analysis performance of the miniaturized spectrometer.
[0146] Figure 5 In (e) of Figure 5 , curve R represents the measured high-quality spectral data, and curve H represents the reconstructed high-quality spectral data.
[0147] The embodiment of the present application also provides a device for improving the quality of spectral data based on spectral reconstruction. Please refer to Figure 6 This device may include the following units.
[0148] An acquisition unit 601, configured to acquire the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample;
[0149] A dimensionality reduction unit 602, configured to perform data dimensionality reduction processing on the first low-quality spectral data to obtain low-dimensional feature data of the first low-quality spectral data;
[0150] A construction unit 603, configured to construct a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data;
[0151] A processing unit 604, configured to process the second low-quality spectral data of the sample to be measured based on the spectral reconstruction model to obtain the second high-quality spectral data, and the second high-quality spectral data is used as the basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
[0152] Optionally, when the construction unit 603 constructs a spectral reconstruction model according to the low-dimensional feature data and the first high-quality spectral data, it is specifically configured to:
[0153] Obtain the model to be trained;
[0154] Determine the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data;
[0155] In the case where the model loss does not meet the convergence condition, update the model to be trained according to the model loss until the model loss meets the convergence condition;
[0156] When the model loss meets the convergence condition, the model to be trained is determined as the spectral reconstruction model.
[0157] Optionally, when the construction unit 603 determines the model loss of the model to be trained based on the low-dimensional feature data and the first high-quality spectral data, it is specifically used for:
[0158] Compare the differences in the low-dimensional feature data of a pair of calibration samples output by the siamese neural network of the model to be trained to obtain a comparison loss;
[0159] Use the decoder of the model to be trained to process the low-dimensional feature data to obtain the simulated hyperspectral data corresponding to the low-dimensional feature data, and determine the reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data;
[0160] Use the fully connected layer of the model to be trained to process the low-dimensional feature data to obtain the predicted label corresponding to the low-dimensional feature data, and determine the label loss according to the difference between the predicted label and the true label corresponding to the calibration sample;
[0161] Combine the comparison loss, the reconstruction loss and the label loss to obtain the model loss of the model to be trained.
[0162] Optionally, when the acquisition unit 601 acquires the first low-quality spectral data, it is specifically used for:
[0163] Detect the calibration sample through at least one of a portable spectrometer, a handheld spectrometer and a micro spectrometer to obtain the first low-quality spectral data, and the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data.
[0164] Obtain the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample.
[0165] Optionally, when the acquisition unit 601 acquires the first high-quality spectral data, it is specifically used for:
[0166] Detect the calibration sample through a bench spectrometer to obtain the first high-quality spectral data, and the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, X-ray fluorescence spectral data.
[0167] For the device for improving the quality of spectral data based on spectral reconstruction provided in this embodiment, the specific working principle and beneficial effects can refer to the relevant steps and beneficial effects of the method for improving the quality of spectral data based on spectral reconstruction provided in any embodiment of this application, and will not be elaborated here.
[0168] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0169] It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0170] Those of ordinary skill in the art can implement or use the present application. Various modifications to these embodiments will be readily apparent to those of ordinary skill in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for improving the quality of spectral data based on spectral reconstruction, characterized in that Including: Obtaining first low-quality spectral data and first high-quality spectral data corresponding to a calibration sample; Performing data dimensionality reduction processing on the first low-quality spectral data to obtain low-dimensional feature data of the first low-quality spectral data; Constructing a spectral reconstruction model based on the low-dimensional feature data and the first high-quality spectral data; Processing second low-quality spectral data of a sample to be measured based on the spectral reconstruction model to obtain second high-quality spectral data, and using the second high-quality spectral data as a basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
2. The method according to claim 1, wherein The constructing a spectral reconstruction model based on the low-dimensional feature data and the first high-quality spectral data includes: Obtaining a model to be trained; Determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data; In the case where the model loss does not meet the convergence condition, updating the model to be trained according to the model loss until the model loss meets the convergence condition; In the case where the model loss meets the convergence condition, determining the model to be trained as the spectral reconstruction model.
3. The method according to claim 2, wherein The determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data includes: Comparing the differences in the low-dimensional feature data of a pair of calibration samples output by the siamese neural network of the model to be trained to obtain a contrast loss; Processing the low-dimensional feature data by a decoder of the model to be trained to obtain simulated hyperspectral data corresponding to the low-dimensional feature data, and determining a reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data; Processing the low-dimensional feature data by a fully connected layer of the model to be trained to obtain a predicted label corresponding to the low-dimensional feature data, and determining a label loss according to the difference between the predicted label and the true label corresponding to the calibration sample; Combining the contrast loss, the reconstruction loss, and the label loss to obtain the model loss of the model to be trained.
4. The method according to claim 1, wherein Obtaining the first low-quality spectral data includes: Detecting the calibration sample by at least one of a portable spectrometer, a handheld spectrometer, and a micro spectrometer to obtain first low-quality spectral data, where the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, and laser-induced breakdown spectral data; Obtaining first low-quality spectral data and first high-quality spectral data corresponding to a calibration sample.
5. The method according to claim 1, wherein Obtaining the first high-quality spectral data includes: Detecting the calibration sample by a bench-top spectrometer to obtain the first high-quality spectral data, where the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, and X-ray fluorescence spectral data.
6. An apparatus for improving the quality of spectral data based on spectral reconstruction, characterized in that, Including: An obtaining unit for obtaining first low-quality spectral data and first high-quality spectral data corresponding to a calibration sample; A dimensionality reduction unit for performing data dimensionality reduction processing on the first low-quality spectral data to obtain low-dimensional feature data of the first low-quality spectral data; A construction unit, configured to construct a spectral reconstruction model based on the low-dimensional feature data and the first high-quality spectral data; A processing unit, configured to process the second low-quality spectral data of a sample to be measured based on the spectral reconstruction model to obtain second high-quality spectral data, and the second high-quality spectral data is used as a basis for quantitative analysis and / or qualitative analysis of the sample to be measured.
7. The device according to claim 6, characterized in that, When constructing the spectral reconstruction model based on the low-dimensional feature data and the first high-quality spectral data, the construction unit is specifically configured to: Obtain a model to be trained; Determine the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data; In the case that the model loss does not meet the convergence condition, update the model to be trained according to the model loss until the model loss meets the convergence condition; In the case that the model loss meets the convergence condition, determine the model to be trained as the spectral reconstruction model.
8. The device according to claim 7, characterized in that When determining the model loss of the model to be trained according to the low-dimensional feature data and the first high-quality spectral data, the construction unit is specifically configured to: Compare the differences in the low-dimensional feature data of a pair of calibration samples output by the siamese neural network of the model to be trained to obtain a comparison loss; Process the low-dimensional feature data by using the decoder of the model to be trained to obtain simulated hyperspectral data corresponding to the low-dimensional feature data, and determine a reconstruction loss according to the difference between the simulated hyperspectral data and the first high-quality spectral data; Process the low-dimensional feature data by using the fully connected layer of the model to be trained to obtain a predicted label corresponding to the low-dimensional feature data, and determine a label loss according to the difference between the predicted label and the true label corresponding to the calibration sample; Combine the comparison loss, the reconstruction loss, and the label loss to obtain the model loss of the model to be trained.
9. The device according to claim 6, wherein, When obtaining the first low-quality spectral data, the obtaining unit is specifically configured to: Detect the calibration sample by using at least one of a portable spectrometer, a handheld spectrometer, and a micro spectrometer to obtain first low-quality spectral data, where the first low-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, and laser-induced breakdown spectral data; Obtain the first low-quality spectral data and the first high-quality spectral data corresponding to the calibration sample.
10. The device according to claim 6, characterized in that, When obtaining the first high-quality spectral data, the obtaining unit is specifically configured to: Detect the calibration sample by using a bench spectrometer to obtain the first high-quality spectral data, where the first high-quality spectral data includes at least one of Raman spectral data, visible spectral data, near-infrared spectral data, laser-induced breakdown spectral data, inductively coupled plasma mass spectrometry data, and X-ray fluorescence spectral data.