Method for generating synthetic spectral data
By synthesizing spectral data based on a modeled distribution, the method addresses the scarcity of training data in spectral analysis, enabling more accurate and reliable deep learning models for spectral data tasks.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- COMMISSARIAT A LENERGIE ATOMIQUE ET AUX ENERGIES ALTERNATIVES
- Filing Date
- 2022-06-21
- Publication Date
- 2026-05-22
AI Technical Summary
Existing deep learning algorithms for spectral data analysis, particularly in techniques like LIBS, face challenges due to the limited availability of training data, leading to issues such as overfitting and poor generalization performance, as they require a large number of spectral realizations that are often not readily available.
A method is proposed to synthesize spectral data by modeling the distribution of existing data and generating new spectra based on this distribution, allowing for the creation of a statistically representative dataset that can be used to train machine learning models, thereby increasing the number of available training samples without altering the original data's distribution.
This approach enables the use of deep learning algorithms with a larger, representative dataset, improving prediction accuracy and reducing uncertainties by ensuring the generated spectra maintain the same distribution as the original data, thus enhancing model performance.
Smart Images

Figure 00000019_0000 
Figure 00000019_0001 
Figure 00000020_0000
Abstract
Description
Title of the invention: Method for generating synthetic spectral data
[0001] The invention relates to the field of spectral data analysis, that is, data exhibiting a plurality of intensity values in different wavelength channels or spectral bands. The data can be both multi- or hyperspectral, where the number of spectral bands ranges from a few dozen to hundreds, and data from emission or absorption spectra of a chemical species, containing thousands of wavelength channels. The invention is applicable to any type of spectral analysis where a large number of replicates of the input data are required and these are not readily available in large quantities. The invention is particularly, but not exclusively, applicable to quantitative analysis (e.g., concentration determination) or to the classification of samples for which spectral data are measured.
[0002] More specifically, the invention relates to a method of synthesizing synthetic spectral data to provide training data to a machine learning engine for the analysis of species associated with spectral data, in particular, but not exclusively, for the quantitative or qualitative analysis of chemical species.
[0003] A possible application of the invention relates to the determination of the concentration of chemical elements or the classification of samples from spectral data, for example, acquired using laser-induced breakdown spectroscopy (LIBS). The invention is not limited to this particular technique; it can be applied to any type of spectroscopy technique that produces multi- or hyperspectral data or spectral data of the emission or absorption of chemical species.
[0004] The invention applies to all types of spectral analysis. In fact, the invention can be used in quantitative analysis, such as predicting a quantity characterizing samples to be analyzed. It also applies to qualitative analysis, such as segmentation or scene or map identification by a technique that produces multi- or hyperspectral images or spectra of chemical species obtained by a spectroscopic technique such as LIBS or others. Furthermore, it can also be applied to sample generation for super-resolution and other unsupervised learning techniques. the difference being simply the nature of the variables to be predicted or dealt with, which are, for example, continuous in quantification (e.g., the concentration of a species), discrete in classification (e.g., a class or category label), or of the same type as the input data for unsupervised analysis (e.g., the intensity values of spectral bands of a pixel in super-resolution images).
[0005] In the context of spectral data, different processing methods are used for different types of analysis. In particular, multivariate deep learning methods, based primarily on artificial neural networks, have been explored and used, for example, for quantitative analysis (calibration, regression) or for sample classification. Examples of such methods are described in references [1]-[3]. However, these algorithms are generally characterized by their ability to learn from a very large number of realizations (spectra), which limits their use when the available datasets contain a limited number of realizations.
[0006] In contrast to the most widely used approaches based on fully connected neural networks as presented in [4], recent developments in spectral signature analysis have led to the introduction of architectures inspired by object detection and image classification algorithms based on convolutional neural networks (see, for example, [5], [6]). Although the same problem arises for all neural network models, this particular type of architecture aims at learning models from training data, which requires a large number of realizations to correctly learn how to associate input data with output data, for example, in supervised learning.For example, standard image processing datasets contain approximately 10⁴ to 10⁶ training samples (see
[20] ), whereas typical LIBS datasets contain tens or hundreds of spectra (see [7]), or several thousand to tens of thousands for LIBS mapping (see [8]). This observation also holds true for other types of spectroscopy.
[0007] Obtaining a large number of spectral data is a problem to be solved. For example, in LIBS spectroscopy, the collection of a large number of spectra may be prevented by the destruction of the sample surface, or by an available surface that is too small, or even by a simple question of time (for example, the impossibility of probing a given area quickly enough).
[0008] Beyond LIBS spectroscopy, the deficit of spectral training data can also be attributed to the high cost of obtaining a sufficient amount of labeled data for training.
[0009] There is therefore a need to realistically increase the amount of data learning available for spectral data.
[0010] The problem of a lack of realizations in spectral analysis is rarely addressed in the literature. There are a few works, discussed below, aimed at enriching the information given to architectures (e.g., neural networks) or at focusing only on an arbitrarily relevant part of the information, but, from the point of view of deep learning techniques, the absence of a large number of different realizations (i.e., spectra) can still lead to problems of overfitting or poor generalization performance.
[0011] In general, data augmentation and synthesis are methods used in deep learning, for example, in computer vision. The basic idea is to create an oversampling of the input data in a non-trivial way. Classically, data augmentation enriches the training data by using transformations (rotations, broadenings, reflections, etc.) of the training data to produce new realizations (see, for example, [9],
[10] ,
[12] ,
[18] ) in most deep learning applications, such as image classification, time series analysis, natural language processing, etc. This procedure makes it possible to produce an arbitrary number (except for constraints related to the size or shape of the data) of examples generated directly from the distribution of the training data.The effect is a regularization and stabilization of learning, which generates a model that generalizes better, either for classification or for regression tasks. The synthesis of new data is commonly used for image processing (e.g., super-resolution
[11] ). Furthermore, the development of deep learning models on smaller datasets, particularly spectroscopic datasets or in the context of one-shot learning in computer vision, is a current topic of research.
[0012] For example, reference [2] relates to a "data augmentation" method for the LIBS technique using time-resolved spectra of chemical elements for multivariate analysis with shallow neural networks. That is, for each crater on the surface, instead of a single spectral signature, several spectra are recorded at different times after the laser firing. The concatenation of these spectra is then used, for each crater, as a representative of the measurement, which now has an additional temporal direction, hence the name "time-resolved spectra." The dataset used for the neural network analysis thus consists of a collection of time-resolved spectra. Here, the term "data augmentation" is not used correctly. Indeed, the number of realizations is not The quantity of information for a given realization has not actually increased, but it has. One could argue that the quality of the data has certainly increased, even if no new data has been produced. The analysis proposed in reference [3] uses the same type of time-resolved data, without explicitly mentioning "data increase."
[0013] The methods described in references
[13] ,
[14] use deep learning methods for LIBS data analysis based on convolutional neural networks. However, the problem of data augmentation is not addressed there. More recently, the authors in
[15] introduced a data augmentation technique derived directly from the standard deep learning image processing methodology. Their analysis is, again, based on convolutional neural networks and focuses on two-dimensional elemental maps with a spatial resolution of 150 pm between craters. Starting from the maps obtained from the intensity of preselected lines, they use cross-sections, recombinations, image filters (e.g., the addition of Gaussian noise and a median filter), and reflections to produce additional training data for sample classification.Note that, in this case, the authors do not directly use the spectral information contained in the original data, but rather extract it from maps to exploit its spatial information. The augmentation is then performed directly on the maps. Within the context of image classification, and for the purposes illustrated by the authors, the techniques used in the article can improve the generalization capabilities of the classifier network. However, for more general purposes, using slices and recombinations to generate new images does not directly modify the data associated with each pixel (i.e., each crater), but rather reorganizes it across the map: such a data augmentation technique leads to oversampling of the collected data at the intensity map level, rather than the production of spectra.For example, other types of analysis, such as multivariate regression for quantitative analysis, may not benefit significantly from this treatment, as it can be considered a simple replication of the input data to the regression network (even if it may lead to slight performance improvements). Furthermore, very small elemental maps, in which only a small number of laser shots are performed, may benefit only marginally, as the number of relevant transformations is considerably reduced.
[0014] Review article
[16] presents the concept of data augmentation by proposing the generation of an arbitrary number of spectra by adding random noise to each experimental spectrum. However, no implementation of this technique is shown. in the article and no definition of random noise is offered.
[0015] Other analyses described in reference
[17] use different types of LIBS spectroscopy data, for example, by considering only specific wavelength channels for the analysis, with the aim of reducing the size of the training data relative to the size of the neural network model. This approach allows the use of a reduced version of the input data, where the information deemed relevant has been pre-extracted to improve the analysis. However, this can still lead to overfitting problems and poor generalization due to the limited amount of available data, as well as a possible reduction in performance due to information loss resulting from the pre-selection of the input data.
[0016] In the context of multi- or hyperspectral image analysis, one can also mention traditional data augmentation methods, generally defined for tasks such as object detection or semantic segmentation (for example, reference [9] gives examples and a complete bibliography of the state of the art). However, in this context, the purpose of the analysis is different and generally limited to the classification or characterization of scenes (similarly, these techniques have also been applied in the context of LIBS spectroscopy in
[15] as discussed above).
[0017] The invention aims to overcome the limitations of the prior art by providing a method for synthesizing spectral data, which allows for better use of deep learning algorithms and, more generally, any algorithm that requires a large amount of spectral input data. This contribution makes it possible to implement more efficient algorithms, capable of reducing prediction uncertainties and building reliable models, but which require a large amount of training data.
[0018] The invention proposes a method for synthesizing spectral data, usable for training purposes such as regularization and oversampling of training data, or directly as training data. The synthesis method according to the invention is based on experimental data to model the signal distribution.
[0019] This distribution can then be used to generate an arbitrary number of spectra, which statistically represent the real data. This new dataset can be used to train deep learning algorithms, which require a large amount of data: since this data models a real distribution, the algorithms maintain their predictive capacity and accuracy on new data acquired experimentally by a spectroscopic method.
[0020] The invention, unlike certain state-of-the-art techniques, relates to the generation of an arbitrary number of truly different spectral training data, statistically representing the experimental data set, without constraint on the number of wavelength channels or spectral bands contained in the spectra.
[0021] The invention proposes a technique different from the prior art for synthesizing an arbitrary number of spectra. Since directly adding random noise to a limited number of spectra can modify the training distribution (i.e., it can change the nature of the distribution, given that the number of realizations is relatively small), the spectra are first modeled based on a known or estimated statistical distribution (for example, using a kernel density estimation method), and then generated according to their statistical distribution to broaden the feature space of the input data, i.e., covering a larger part of the distribution's domain of definition. In this way, the generated dataset is always a statistical representation of the original data with an arbitrarily large number of replicas.Random noise (e.g., Gaussian or uniform) can then be added separately to each synthesized replica to improve the algorithm's generalization capabilities. Using synthesized data provides a sufficiently large input dataset that the addition of noise is negligible on average, with no overall impact on the data distribution. Conversely, adding noise to a limited number of data points can significantly alter the nature of the data and disrupt the algorithm's learning. Generating from a statistical distribution ensures that each replica is a different representation of the training data, enabling the algorithm to learn a greater number of features, and that the number of replicas is high enough to guarantee that, statistically, the training distribution is representative of the analyzed samples.
[0022] Unlike the prior art, the invention proposes an augmentation method directly related to the nature of spectral signatures to solve the problem of the number of spectra available for training. Since no prior knowledge of the type of spectral data is required (for example, it can be estimated), the same principle presented here can be extended to any type of multi- or hyperspectral data, not necessarily related to the LIBS technique.
[0023] The invention relates to a method for modeling the distribution of spectra for the realistic synthesis of data, compared to experimental data. The invention also provides for a step of adding random noise from the synthesized data, as opposed to adding the noise directly to the original data. This technique makes it possible to generate an arbitrary number of data that are actually representative of the samples and then to modify the spectral intensities, without altering on average the original distribution of the experimental data (which, in applications, consists of only a few realizations, and is not representative of the true distribution of the data).
[0024] Unlike conventional computer vision data augmentation techniques, any transformation (shift, translation, reflection, dilation) applied to spectral data will inevitably alter the physical meaning of the spectra: for example, shifting an emission line attributed to one element in wavelength may lead to its being attributed to another element. The invention proposes to generate new training spectra, that is, to synthesize training data using a theoretical model of the distribution of real data. In other words, the spectral profile obtained experimentally by a spectroscopic method is used to generate spectra having, on average, the same distribution for each wavelength channel. This approach makes it possible to solve the problem of the number of realizations (spectral signatures) without distorting the physical content of the spectra.Spectra are generated using random extractions from this distribution: the method also allows covering a larger part of the space in which the original data are defined (for example, in the case of spectroscopic data, the wavelength space).
[0025] The invention relates to a computer-implemented method for synthesizing spectral data comprising the steps of: - Acquire a set of spectral data, each associating a spectrum with a sample having a given chemical composition, using a spectroscopic method. - Determine a theoretical model of the intensity distribution of the spectrum for each wavelength channel of the spectrum, - Generate a set of synthetic spectral data by generating, for each wavelength channel of the spectrum, an intensity drawn randomly according to the probability distribution of the theoretical model.
[0026] According to a particular aspect of the invention, the theoretical model is based on a probability distribution according to a Poisson law parameterized by the intensity measured on the acquired spectrum.
[0027] According to a particular aspect of the invention, the spectral data set includes several spectral measurements for the same sample and the method includes a step of determining the average spectrum over the set of measurements.
[0028] According to a particular aspect of the invention, the synthetic spectral data are generated by adding to the randomly drawn intensity a noise value drawn according to a uniform distribution in an interval centered on the intensity and of configurable width.
[0029] According to a particular aspect of the invention, the synthetic spectral data are generated by adding to the randomly drawn intensity a noise value drawn according to a normal distribution centered on the intensity, the standard deviation of which is a modifiable parameter.
[0030] According to a particular aspect of the invention, spectral data are acquired by means of a laser-induced plasma atomic emission spectroscopy method.
[0031] According to a particular aspect of the invention, the spectral data are derived from emission or absorption spectra of chemical species.
[0032] The invention also relates to a method for quantitative or qualitative analysis of spectral data comprising the steps of: - Generate a set of synthetic spectral data by executing the spectral data synthesis method according to the invention, - Train a machine learning model from the generated synthetic spectral data. - Use the trained model to perform a quantitative or qualitative analysis of spectral data.
[0033] The invention further relates to a computer program comprising instructions for the execution of a method according to the invention, when the program is executed by a processor, and to a processor-readable recording medium on which is recorded a program comprising instructions for the execution of a method according to the invention, when the program is executed by a processor.
[0034] Other features and advantages of the present invention will become more apparent from the following description in relation to the following accompanying drawings.
[0035] [Fig. 1] represents an example of spectral data characterizing a sample containing different chemical species,
[0036] [Fig.2] represents a diagram of the steps for implementing a method of ge generation of synthetic spectral data according to the invention,
[0037] [Fig.3] represents a flowchart of the steps in implementing a method machine learning of a spectral data analysis model according to the invention,
[0038] [Fig.4] represents a quantile-quantile diagram of the real and syn distributions Therapeutic solution for a cement sample (type I) with the addition of NaCl,
[0039] [Fig. 5a] represents an example of a mean spectrum
[0040] [Fig. 5b] represents an illustration of the results obtained by the invention with a mo- Gaussian-type delimitation,
[0041] [Fig.5c] represents an illustration of the results obtained by the invention with a model based on a "tophat" kernel
[0042] LIBS technology enables material analysis by laser ablation and spectroscopy. The data acquired via this technique are spectral data which correspond, for each point in a zone, to an emission spectrum comprising atomic lines characteristic of the elemental chemical composition of the sample.
[0043] LIBS spectral data are obtained by focusing a laser beam onto a point on a surface to be analyzed. The plasma emission resulting from this focusing is collected and processed spectroscopically to obtain a spectrum of atomic lines. The process is iterated for each point in the area to be analyzed.
[0044] Figure 1 shows, by way of illustration, an example of an atomic line spectrum obtained for a sample having a certain chemical composition. In Figure 1, the spectral signatures of certain chemical elements (Ca, Al) have been identified, which correspond to atomic lines in channels of given wavelengths.
[0045] As explained in the preamble, the invention aims to generate synthetic spectral data from one or more spectral data measurements of the type described in [Fig.1].
[0046] The method according to the invention is described in [Fig.2].
[0047] The first step 110 consists of acquiring spectral data by means of a The appropriate acquisition device depends on the intended application. If the application involves a qualitative or quantitative analysis of samples, for example, of a material, the data are spectral data and are acquired, for example, using a spectrometry device, such as laser-induced plasma atomic emission spectroscopy, or a device based on a mass spectrometry technique coupled with laser ablation, an ion beam, an X-ray beam, synchrotron radiation-induced spectrometry, charged particle beam spectrometry, Raman spectrometry, or infrared spectroscopy. If the application involves a method for mapping a geographic area, the multi- or hyperspectral data are acquired, for example, using a multi- or hyperspectral imaging sensor onboard a satellite payload.The invention applies more generally to any other multi- or hyperspectral data acquisition device capable of generating, for a given sample, a spectrum within a given wavelength range.
[0048] The first step 110 may consist of measuring a single spectrum per sample or several spectra per sample.
[0049] In an optional step 121, the measured spectral data are pre-processed In order to estimate and correct any potential acquisition-related offset, the different measured spectra are normalized to ensure homogeneity, and any blind spots are removed. In other words, each measured spectrum can be normalized in various ways, for example, by a known emission / absorption line or band, by the maximum intensity, or by other methods. If several spectra are used that are assumed to be representative of the measurement, one can also focus on a specific wavelength channel, consider the average intensity, and discard spectra containing outliers for that channel from the overall data. This preprocessing allows the use of only the most representative spectra of the sample, without necessarily modeling defects simultaneously.
[0050] If several spectral measurements are performed on the same sample, the spectra are averaged in step 122. In other words, several spectra representing the same sample can be used to model the distribution (for example, following several laser shots on the same sample in the LIBS technique). The spectra used for generating synthetic data are averaged to obtain a more accurate representation of the analyzed sample. Put another way, instead of using a single spectrum as representative of a sample, the spectroscopic measurement can be replicated several times, and the average spectrum obtained from a sample can be used for synthesis. This approach allows for a more accurate representation of the sample, taking into account possible differences on average across the surface.However, it should be noted that this embodiment of the invention is more specifically applicable to spectral data without an image concept, that is, for data for which the spectroscopic measurement can be repeated without changes in the physical meaning of the data (each spectrum must be representative of the same distribution). Applying this embodiment to multi- or hyperspectral maps implies the presence of several realizations of the same image in order to average the contribution of a single pixel. This application is not possible with the LIBS technique since the destructive nature of the laser's interaction with the surface does not allow the measurement to be reproduced at the same location. On the other hand, acquiring multi- or hyperspectral images using an orbital mapping method, for example, allows the same image to be replicated multiple times.
[0051] In all cases, an experimental measurement of a spectrum is obtained.
[0052] Next, a model is determined (step 130) of the distribution of intensity values of the spectral lines from the experimental measurement.
[0053] In the case of spectral data obtained by a LIBS acquisition method, the main source of noise at low intensities and of signal at high intensities is constituted by the photons that impacted the detector. We can therefore estimate the actual distribution of spectral data using a distribution that models the photon count.
[0054] The distribution model used is therefore based on a Poisson probability distribution expressed by the formula p^X = k) = — ' oa is 'a variable of the distribution which is here the intensity of the spectral lines and 1 is the parameter of the Poisson law.
[0055] If we denote ln the parameter of the Poisson distribution for the channel of wavelength n, this parameter also corresponds to the expected mean of the distribution for the channel n. Consequently, within the framework of the invention, for each channel of wavelength n, we impose ln = In, that is to say the peak of the probability distribution of the synthetic spectra in a channel n is equal to the intensity In recorded for the channel in the experimental spectrum which we consider to model the synthetic spectra (that provided as input to step 130, possibly averaged in step 122).
[0056] Next, in step 140, new synthetic spectral data are generated from the model obtained in step 130 for each wavelength channel n. A new synthetic spectrum is obtained by determining each intensity of the spectrum for each wavelength 11 by means of a random sampling following the intensity distribution model obtained in step 130. The random sampling is calculated by inverting the cumulative distribution function and using it to represent a random variable, uniformly distributed in the interval [0, 1], in the probability space. It is thus possible to generate an arbitrary number of spectra having statistically the same properties as the experimental spectra 110.
[0057] By way of illustration, [Fig. 4] shows the quantile-quantile diagram of the actual and synthetic distributions for a cement sample (type I) with the addition of NaCl. The data were synthesized by modeling the intensity using a Poisson distribution. The diagram shows points aligned on the bisector of the first quadrant: the observed quantiles effectively overlap the quantiles of the experimental distribution.
[0058] A synthetic spectral dataset 150 is then obtained, in greater numbers than could be obtained experimentally. The synthetic dataset 150 can then be used as a training set comprising spectra that simultaneously represent the same distribution of the input data and different realizations of the experimental measurements (i.e., new data, independent of the experimental data).
[0059] In one embodiment of the invention, instead of modeling the intensity of Since each wavelength channel follows a Poisson distribution, the intensity distribution of the spectrum can be modeled using, for example, a nonparametric kernel density estimation (KDE) method, as described, for instance, in M. Rosenblatt, “Remarks on Some Nonparametric Estimates of a Density Function,” Ann. Math. Statist. 27 (3) 832–837, September 1956. In this variant, a kernel function K(z, h) is used to estimate the density f(x) of a random variable x (the intensity, in the case of spectra), using a number of realizations (experimental spectra). The form of f(x) is estimated by a function 1 1 ' !=1, . 'f' (x) — V h) For each value of x. The parameter h represents a bandwidth, which can be adapted to improve the estimation of f(x) by
[0060] The function can be estimated by different choices of the kernel K. In variants that can be used for spectral analysis, one can choose K(x, h) ≤ (a "Gaussian" kernel), or, for example, K(x, h) ≤ 9(hx) (a so-called "tophat" kernel), where 6 is the Heaviside function. The choice of h normally depends on the type of data to be modeled: a smaller bandwidth allows the kernel profile to be better adapted to the data, at the risk of generating oversampling effects. To choose h, one can, for example, use quantile-quantile diagrams to compare the distribution of the real data and the distribution of the synthesized data using the estimator j of the spectral intensity density.
[0061] Figures 5a, 5b, 5c show the comparison of the modeling by a Gaussian kernel and a "tophat" kernel of a cement sample (type I) with the addition of NaCl analyzed by a LIBS technique. The average spectrum 500 is shown in [Fig. 5a]
[0062] Different spectra 501, 502, 503, 504 obtained for a Gaussian nucleus are shown in [Fig. 5b]. Different spectra 510, 520, 530, 540 obtained for a "tophat" nucleus are shown in [Fig. 5c].
[0063] For each spectrum, an associated quantile-quantile diagram is also represented.
[0064] Normally, the data are better reproduced using low values of the bandwidth, since the quantiles are aligned on the bisector of the diagram. Higher values of h show a deviation of the quantiles at both low and high intensities. The comparison also shows a better fit to the "top-hat" kernel data for high values of h. Conversely, at low values of / 1, a Gaussian kernel fits the data better.
[0065] In one embodiment, the synthetic data distribution can be made even more realistic by adding, during the generation of the synthetic data, an additional random noise source for each wavelength channel. Such a source is modeled as a difference in the number of photons reaching the detector.
[0066] The intensity of a spectrum for the wavelength is then given by In = ( 1 + Um ) In, where, for each channel of wavelength n, In follows a Poisson distribution P with parameter InX (i.e., In ~ P(In), where In is the intensity recorded experimentally for the channel (possibly averaged at step 122) and corresponds to the expected mean of the distribution of In), m is a noise parameter chosen such that Um is a number uniformly distributed in the interval [ - m , ni],
[0067] In an alternative embodiment, one can define jn — fl . jn, where m is a noise parameter chosen such that Nm is a number distributed according to a normal distribution centered at 1 and with a standard deviation m, i.e. Nm ~ N(l, ni).
[0068] In one embodiment, the generated synthetic spectral data 150 can be added (step 160) to the measured input data 110 to construct a training dataset.
[0069] Alternatively, it is also possible to use only the synthetic spectra 150 as a training set because, in general, the number of spectra generated is much greater than the number of experimental data, to the point that the latter become statistically negligible.
[0070] The dataset obtained by the method according to the invention can be used to train a machine learning engine as illustrated in an example in [Fig.3].
[0071] Synthetic spectral data are generated in step 301 from initial training spectral data measured in step 300, and then used as training data to train an analysis model in step 302. The analysis model can target a quantitative analysis, for example, an estimation of the concentration of a chemical species in a sample from the analysis of its spectrum, or a qualitative analysis, for example, a classification of spectra according to the type of sample.
[0072] The machine learning model is, for example, based on one or more convolutional neural networks or any other equivalent machine learning algorithm. The training data can be used to perform oversampling and / or regularization of deep learning methods. References [9]-
[10] -
[12] provide, by way of illustration, various methods learning methods adapted to the qualitative or quantitative analysis of spectral data.
[0073] Once the model has been trained, it can be used in step 303 to perform a qualitative or quantitative analysis of new spectral data measured in step 304.
[0074] The steps of the invention can be implemented as a computer program comprising instructions for its execution. The computer program can be stored on a storage medium readable by a processor.
[0075] The reference to a computer program that, when executed, performs any of the functions described above, is not limited to an application program running on a single host computer. Rather, the terms computer program and software are used here in a general sense to refer to any type of computer code (for example, application software, firmware, microcode, or any other form of computer instruction) that can be used to program one or more processors to implement aspects of the techniques described herein. The computing means or resources may, in particular, be distributed ("cloud computing"), possibly using peer-to-peer technologies.The software code can be executed on any suitable processor (e.g., a microprocessor) or processor core, or a set of processors, whether located in a single computing device or distributed across multiple computing devices (e.g., as potentially accessible within the device's environment). The executable code of each program, enabling the programmable device to implement the processes according to the invention, can be stored, for example, on the hard drive or in read-only memory. Generally, the program(s) can be loaded into one of the device's storage means before being executed. The central processing unit can command and direct the execution of the instructions or portions of software code of the program(s) according to the invention, instructions which are stored on the hard drive or in read-only memory, or in the other aforementioned storage elements. References
[0076] [1] M. H. Mozaffari and L.-L. Tay, “A Review of 1D Convolutional Neural Networks toward Unknown Substance Identification in Portable Raman Spec-trometer,” ArXiv200610575 Cs Eess, 2020, Accessed: Oct. 29, 2021. [Online]. Available: http: / / arxiv.org / abs / 2006.10575
[0077] [2] L. Narlagiri and V. R. Soma, “Simultaneous quantification of Au and Ag com position from Au-Ag bi-metallic LIBS spectra combined with shallow neural network model for multi-output régression,” Appl. Phys. B, vol. 127, no. 9, p. 135, 2021, doi: 10.1007 / s00340-021-07681-y.
[0078] [3] C. Lu, B. Wang, X. Jiang, J. Zhang, K. Niu, and Y. Yuan, “Détection of K in soil using time-resolved laser-induced breakdown spectroscopy based on convolutional neural networks,” Plasma Sci. Technol., vol. 21, no. 3, p. 34014, 2019, doi: 10.1088 / 2058-6272 / aaef6e.
[0079] [4] F. Rosenblatt, The perceptron, a perceiving and recognizing automaton Project Para. Cornell Aeronautical Laboratory, 1957.
[0080] [5] Y. LeCun et al., “Backpropagation Applied to Handwritten Zip Code Ré cognition,” Neural Comput., vol. 1, no. 4, pp. 541-551, 1989, doi: 10.1162 / neco.1989.1.4.541.
[0081] [6] Y. LeCun et al., “Handwritten digit récognition with a back-propagation network,” Adv. Neural Inf. Process. Syst., vol. 2, 1989.
[0082] [7] D. W. Hahn and N. Omenetto, “Laser-induced Breakdown Spectroscopy (LIBS), Part II: Review of Instrumental and Méthodologie al Approaches to Material Analysis and Applications to Different Fields,” Appl. Spectrosc., vol. 66, no. 4, pp. 347-419, 2012, doi: 10.1366 / 11-06574.
[0083] [8] L. Jolivet, M. Leprince, S. Moncayo, L. Sorbier, C.-P. Lienemann, and V. Motto- Ros, “Review of the recent advances and applications of LIBS-based imaging,” vol. 151, pp. 41-53, 2019, doi: 10.1016 / j.sab.2018.11.008.
[0084] [9] C. Shorten and T. M. Khoshgoftaar, “A survey on Image Data Augmentation for Deep Learning,” J. Big Data, vol. 6, no. 1, p. 60, 2019, doi: 10.1186 / s40537-019-0197-0.
[0085]
[10] A. Mikolajczyk and M. Grochowski, “Data augmentation for improving deep learning in image classification problem,” in 2018 International Interdisciplinary PhD Workshop (IIPhDW), Swinoujscie, 2018, pp. 117-122. doi: 10.1109 / IIPHDW.2018.8388338.
[0086]
[11] K. Li, D. Dai, E. Konukoglu, and L. Van Gool, “Hyperspectral Image Super- Resolution with Spectral Mixup and Heterogeneous Datasets,” ArXiv210107589 Cs, 2021, Accessed: Jan. 12, 2022. [Online]. Available: http: / / arxiv.org / abs / 2101.07589
[0087]
[12] Q. Wen et al., “Time Sériés Data Augmentation for Deep Learning: A Survey,” in Proceedings of the Thirtieth International Joint Conférence on Artificial Intelligence, Montreal, Canada, 2021, pp. 4653-4660. doi: 10.24963 / ijcai.2021 / 631.
[0088]
[13] J. Chen, J. Pisonero, S. Chen, X. Wang, Q. Fan, and Y. Duan, “Convolutional neural network as a novel classification approach for laser-induced breakdown spectroscopy applications in lithological récognition,” Spectrochim. Acta Part B At. Spectrosc., vol. 166, p. 105801, 2020, doi: 10.1016 / j.sab.2020.105801.
[0089]
[14] L. Zou et al., “Online simultaneous détermination of H2O and KC1 in potash with LIBS coupled to convolutional and back-propagation neural networks,” J. Anal. At. Spectrom., vol. 36, no. 2, pp. 303-313, 2021, doi: 10.1039 / D0JA00431F.
[0090]
[15] T. Chen et al., “Deep learning with laser-induced breakdown spectroscopy (LIBS) for the classification of rocks based on elemental imaging,” Appl. Geochem., vol. 136, p. 105135, 2022, doi: 10.1016 / j.apgeochem.2021.105135.
[0091]
[16] L.-N. Li, X.-F. Lin, F. Yang, W.-M. Xu, J.-Y. Wang, and R. Shu, “A review of artificial neural network based chemometrics applied in laser-induced breakdown spectroscopy analysis,” Spectrochim. Acta Part B At. Spectrosc., vol. 180, p. 106183, Jun. 2021, doi: 10.1016 / j.sab.2021.106183.
[0092]
[17] J. El Haddad et al., “Artificial neural network for on-site quantitative analysis of soils using laser induced breakdown spectroscopy,” Spectrochim. Acta Part B At. Spectrosc., vol. 79-80, pp. 51-57, 2013, doi: 10.1016 / j.sab.2012.11.007.
[0093]
[18] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. MIT Press, 2016.
[0094]
[19] J. J. Bird, D. R. Faria, C. Premebida, A. Ekart, and P. P. S. Ayrosa, “Overcoming Data Scarcity in Speaker Identification: Dataset Augmentation with Synthetic MFCCs via Character-level RNN,” in 2020 IEEE International Conférence on Autonomous Robot Systems and Compétitions (ICARSC), Ponta Delgada, Portugal, 2020, pp. 146-151. doi: 10.1109 / ICARSC49921.2020.9096166.
[0095]
[20] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li and L. Fei-Fei, ImageNet: A Large- Scale Hierarchical Image Database. IEEE Computer Vision and Pattern Récognition (CVPR), 2009.
Claims
Demands
1. A computer-implemented method for synthesizing spectral data comprising the steps of: - Acquiring (110) a set of spectral data associating each a spectrum with a sample having a given chemical composition, by a spectroscopic method, each spectrum exhibiting a plurality of intensities as a function of wavelength channels - Determining (130) a theoretical model of the distribution of spectral intensities for each wavelength channel of the spectrum, - Generating (140) a set of synthetic spectral data (150) by generating for each wavelength channel of the spectrum, an intensity drawn randomly according to the probability distribution of the theoretical model.
2. Method of spectral data synthesis according to claim 1 wherein the theoretical model is based on a probability distribution according to a Poisson law parameterized by the intensity measured on the acquired spectrum.
3. Method of spectral data synthesis according to any one of the preceding claims wherein the spectral data set comprises several spectral measurements for the same sample and the method comprises a step (122) of determining the average spectrum over the set of measurements.
4. Method of synthesizing spectral data according to any one of the preceding claims wherein the synthetic spectral data are generated (140) by adding to the randomly drawn intensity a noise value drawn according to a uniform distribution in an interval centered on the intensity and of parameterizable width.
5. Method of synthesizing spectral data according to any one of claims 1 to 3 wherein the synthetic spectral data are generated (140) by adding to the randomly drawn intensity a noise value drawn according to a normal distribution centered on the intensity, the standard deviation of which is a modifiable parameter.
6. Method for synthesizing spectral data according to any one of the previous claims wherein the spectral data are acquired (110) by means of a laser-induced plasma atomic emission spectroscopy method.
7. Method of synthesizing spectral data according to any one of the preceding claims wherein the spectral data are derived from emission or absorption spectra of chemical species.
8. A method for quantitative or qualitative analysis of spectral data comprising the steps of: - Generating (301) a set of synthetic spectral data by executing the spectral data synthesis method according to any one of the preceding claims, - Training (302) a machine learning model from the generated synthetic spectral data, - Using (303) the trained model to perform a quantitative or qualitative analysis of spectral data (304).
9. A computer program comprising instructions for carrying out a method according to any one of claims 1 to 7, when the program is executed by a processor.
10. Processor-readable recording medium on which is recorded a program containing instructions for executing a method according to any one of claims 1 to 7, when the program is executed by a processor.