Method and device for carrying out a spectral analysis for determining a spectrum of a sample
Patent Information
- Application Number
- EP2023825471
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-12-06
- Filing Date
- 2023-12-06
- Publication Date
- 2025-09-10
AI Technical Summary
Current spectral analysis methods for determining element concentration in samples are time-consuming and costly due to iterative evaluation procedures, requiring known parameters and lengthy calibration processes.
The method employs neural networks trained on large datasets of simulation and actual spectra to perform quantitative and qualitative analysis, using inverse functions and feature reduction techniques to accelerate and enhance the precision of spectral analysis, allowing for rapid and accurate determination of element concentrations and presence/absence of chemical elements.
This approach significantly reduces evaluation time and computational resources while maintaining high accuracy, enabling quick and precise quantitative and qualitative analysis of sample spectra, and facilitates calibration of unknown measuring devices.
Smart Images

Figure 1.1
Abstract
Description
[0001] Method and device for performing a spectral analysis to determine a spectrum of a sample
[0002] The invention relates to a method and a device for spectral analysis for determining a spectrum of a sample.
[0003] In spectral analysis, spectra of a sample are determined using a spectrometer or measuring device. These spectra contain information about the sample's physical properties. These spectra are evaluated to determine, for example, the element concentration in a layer on or within the sample, as well as the presence of chemical elements. These processes must be performed quickly and with high repeatability to enable reliable conclusions from the measurement. Therefore, a fundamental goal in such spectral analysis is to improve the quantitative and / or qualitative analysis and to reduce the time required for this spectral analysis.
[0004] Until now, the element concentration in the sample was determined, for example, by an iterative evaluation procedure.
[0005] This involved selecting a number of parameters from a physical model and using them as a basis. The measured spectrum was then compared with a theoretical spectrum derived from the physical model. After iterative optimization of the parameters, a result was output representing a parameter set that best matches the measured and theoretical spectrum. This iterative procedure is time-consuming and costly.
[0006] The invention is based on the object of proposing a method and a device for carrying out a spectral analysis to determine a spectrum of a sample in order to enable at least a fast and precise quantitative analysis.
[0007] This problem is solved by a method for performing a spectral analysis to determine a spectrum from a sample, in which at least one acquired spectrum is fed to and evaluated by at least one first network architecture with an analyzing neural network that is trained for quantitative analysis of the spectrum. The first network architecture is trained using a plurality of simulation spectra, which are generated using a simulation method based on the existing physical model S = P(ct>).
[0008] Furthermore, at least one second network architecture with a second neural network is trained for a qualitative analysis of the spectrum. Thus, both the quantitative and qualitative analysis can be performed in a single process step, and a result for the at least one element concentration and / or the at least one element identification of the sample can be output with high accuracy / certainty.
[0009] The first and second network architectures may be structured differently.
[0010] The analyzing neural network of the at least one network architecture is preferably trained to have an inverse function P -1(S) of a physical model S = P(T, ,K) is approximated, so that in the quantitative analysis at least one element concentration or layer thickness of the sample from which the spectrum was recorded is output and / or in the qualitative analysis the presence or absence of at least one chemical element in the sample is determined and output. To carry out the quantitative analysis it was previously necessary to analyze the spectrum using the physical model, whereby all relevant parameters had to be known. This model corresponds abstractly to the equation S = P(c|o), where the spectrum S is a non-linear function P of and stands for all variables that are to be determined, such as the concentration of all relevant elements and, where applicable, the layer thickness.The goal of quantitative analysis is to determine the amplitude of a measured spectrum under given measurement conditions of a known measuring instrument or device. It has been recognized that an inverse function P is required for this purpose. -1 (S) is required to determine the variables without iterative procedures. By training the analyzed neural network on this inverse function, a fast and precise quantitative analysis for at least one element concentration in the sample and / or a qualitative analysis for the absence or presence of at least one element in the sample becomes possible.
[0011] Furthermore, it is preferably provided that the first network architecture is trained using a plurality of simulation spectra, which are generated starting from the existing physical model S=P(0, T, , K) using the simulation method, in particular a Monte Carlo simulation, and / or that the first network architecture is trained using a plurality of actually acquired spectra. Training the neural network of the first network architecture with a plurality of simulation spectra requires a one-time increase in time and possibly increased computing power. However, the evaluation time for each acquired spectrum can be significantly reduced after training the first network architecture.
[0012] The simulation spectra are preferably determined using at least relevant and / or predetermined parameters, such as the concentration of the chemical element or the concentration of at least one element of an alloy, element-specific physical constants, the layer thickness of the at least one layer, different measurement conditions of the measuring devices, and / or characteristic properties of different measuring devices and / or device-specific data from one or more different measuring devices. Furthermore, it is preferably provided that the second network architecture for the qualitative analysis is trained using a plurality of simulation spectra, which represent a plurality of, in particular, differing intensity distributions of energy spectra of the chemical elements.Preferably, the network architecture is not trained for all existing chemical elements of the Periodic Table of Elements (PSE), but rather for a specific selection of elements required for spectrum analysis. In particular, specific chemical elements can be selected for X-ray fluorescence analysis, for example, those with an atomic number greater than 9. For other spectral analyses, appropriately adapted chemical elements can be selected.
[0013] The first network architecture for quantitative analysis is preferably constructed by scaling the acquired spectrum with a scaling network using a scaling factor to compensate for various possible excitation conditions. The scaled spectrum is then fed to a prediction network, and this prediction network outputs a result from the acquired spectrum. This, in turn, enables a more precise quantitative determination of the network.
[0014] A neural network is used for the prediction network and / or the scaling network, with a dense convolutional network (DenseNet) being preferred. This enables a shortened simulation time.
[0015] For the quantitative and / or qualitative analysis of the acquired spectrum, a feature reduction method is preferably used to generate pre-processed simulation spectra. This feature reduction method, in turn, serves to accelerate the training process. Furthermore, it is envisaged that, in the second network structure, a multilayer network (MLP), a convolutional neural network (CNN), or a dense convolutional neural network (DenseNet), for example, is trained based on a large number of spectra that are pre-processed using one of the feature reduction methods. In particular, it has proven advantageous for constructing the second network structure to perform feature selection as a feature reduction method, followed by selecting and training a DenseNet.
[0016] One of the feature reduction methods can provide for the number of features to be evaluated from the already generated spectra, each of which comprises, for example, 1024 features, to be reduced to a number of preferably 512, 256, 128, 64, 32, or 16 features. The features of a spectrum also include the frequencies, or the energies correlated with them, which are output as so-called channels, particularly when using A / D converters for the detector. The channels are equivalent to the features.
[0017] Another feature reduction method that can be used is mean compression, in which the number of features of the acquired spectrum is reduced by averaging to a target spectrum with a reduced number of features.
[0018] Furthermore, the feature reduction method can be implemented as feature selection, in which the number of features or channels of the respective spectrum is reduced based on their original size in the generated spectrum. The generated spectrum is one that was created using the simulation method described above.
[0019] Furthermore, it can alternatively be provided that a feature transformation is carried out, in particular with an autoencoder network, in which a weighting of the features or channels of the respective spectrum is carried out according to relevant and non-relevant features and the non-relevant features are eliminated.
[0020] By means of the aforementioned feature reduction methods, approximately similar performance results in the accuracy and precision of the results can be achieved within a range in which the features of the spectrum are reduced, for example, from 1024 to up to 32 features, which are advantageously between 0.92 and 0.98 in an evaluation range from 0 to 1.
[0021] When using feature selection as a feature reduction method, a watermelon model is preferably selected, in which the selection of features for selection is preferably carried out by a Bayesian error rate estimation.
[0022] Furthermore, it is preferably provided that the analyzing neural network is trained with a model for network reduction. Preferably, the neural network is reduced by components such as neurons, filters, and / or parameters, in particular to eliminate redundant components. This can also enable a reduction in the required resources.
[0023] Furthermore, it is preferably provided that the analyzing neural network is trained with a quantization model in which the data size, in particular of a node of the neural network, is reduced. Preferably, a bit size of 32 bits (32-bit floats) or less is selected.
[0024] The aforementioned feature reduction methods, which are preferably performed after training the neural network of the first and / or second network architecture, can be applied individually, in any combination, or even cumulatively. Using the aforementioned feature reduction methods, the size of the respective neural network used can be reduced, while the evaluation results are the same or even better. Using at least one feature reduction method, improved precision and shortened evaluation of sample measurements can be enabled, significantly reducing data and computing costs.
[0025] Furthermore, it is preferably provided that a neural network of a meta-network is trained using simulation data from a number of known devices to carry out a meta-learning process. Known devices are understood to be those that are already calibrated and preferably have at least one characteristic property of the device already determined. This allows for minimizing calibration costs for the devices, particularly when manufacturing a large number of devices or measuring instruments.
[0026] Furthermore, it is preferably provided that a meta-learning method with the meta-network is used to calibrate unknown measuring devices for spectral analysis or devices. In this method, the properties and / or the changing measurement conditions and / or the measurement tasks of the unknown measuring device are recorded and the neural network of the meta-network is additionally trained using these simulation data and / or measured data from the unknown measuring device. The duration of training the meta-network using simulation data from the unknown measuring device is significantly reduced compared to the duration of training for at least one neural network for the unknown measuring device. After training the meta-network to calibrate the unknown network, the device is ready for quantitative and / or qualitative analysis.Unknown devices are preferably understood to be devices that are completed in production but not yet calibrated. For the present method, a multilayer perceptron (MLP), a convolutional neural network (CNN), or a dense convolutional neural network (DNN), in particular a dense convolutional network (DenseNet), can be used to form the analyzing neural network.
[0027] Furthermore, to carry out the spectral analysis, it is preferably provided that in a first step the qualitative analysis of the acquired spectrum is carried out using the second analyzing neural network and the result is output. This makes it possible to determine within a short process time whether the element(s) to be examined are present in the sample. In some cases, such an analysis may be sufficient. In this case, the method can be terminated. In most cases, however, a statement about the concentration of the element(s) is required. In this case, the quantitative analysis is carried out using the first neural network following the qualitative analysis. By training the neural network with a large number of simulation spectra, a very precise determination of the concentration of the element(s) in the sample can be carried out within a very short evaluation time.Alternatively, both the quantitative and qualitative analyses can be performed simultaneously. It is understood that only the quantitative analysis can be performed with the first neural network.
[0028] The object underlying the invention is further achieved by a device for carrying out a spectral analysis to determine the spectrum of a sample, in particular for carrying out the method according to one of the previously described embodiments, which device comprises a source for emitting primary radiation onto the sample, as well as a detector for detecting secondary radiation which is emitted after the sample has been excited with the primary radiation, wherein a computer-assisted evaluation device evaluates the at least one detected spectrum by the detector, wherein at least one analyzing neural network with at least a first network architecture is provided for the quantitative analysis of the spectrum and at least one second network architecture with a second analyzing neural network is provided for the qualitative analysis of the spectrum, as well as an output device,which outputs a result of the at least one spectrum analyzed by the at least one neural network for the at least one element concentration and / or the at least one element identification of the sample.
[0029] The object underlying the invention is further achieved by a computer program for carrying out a spectral analysis for determining a spectrum of a sample, which is provided in particular for the evaluation device of the aforementioned device, wherein the computer program is stored on at least one computer-readable storage medium which can be executed in the computer-assisted evaluation device and causes the evaluation device to carry out a method according to one of the preceding embodiments.
[0030] The invention, as well as further advantageous embodiments and developments thereof, are described and explained in more detail below with reference to the examples shown in the drawings. The features shown in the description and the drawings can be used individually or in any combination according to the invention. They show:
[0031] Figure 1 is a schematic view of a device for performing a spectral analysis,
[0032] Figure 2 is a diagram of a spectrum of an alloy determined by spectral analysis,
[0033] Figure 3 is a schematic view of training a neural network with simulation spectra, Figure 4 is a schematic representation of a first network architecture of the neural network,
[0034] Figure 5 is a schematic view of a network structure for the network architecture according to Figure 4,
[0035] Figure 6 is a diagram showing a comparison between test losses before and after training the neural network,
[0036] Figure 7 is a diagram showing the evaluation times of different network structures,
[0037] Figure 8 shows a table comparing evaluation times between conventional measuring instruments and those supported by a neural network,
[0038] Figure 9 is a schematic view of a second network architecture for qualitative analysis,
[0039] Figure 10 is a schematic view of a network structure for the second network architecture,
[0040] Figure 11 is a diagram illustrating the performance of feature reduction methods applied to a neural network,
[0041] Figure 12 is a schematic view of a spectrum with a full set of features,
[0042] Figure 13 is a schematic view of a spectrum with a reduced number of features,
[0043] Figure 14 is a schematic view of a data size of a neural network, Figure 15 is a schematic view of a reduced data size of the neural network in Figure 14,
[0044] Figure 16 is a table showing the time reduction of the evaluation time of different feature reduction methods when using different neural networks,
[0045] Figure 17 is a view of steps for training a neural network for spectral analysis with high performance, and
[0046] Figure 18 shows a schematic diagram of a meta-learning procedure for calibrating measuring instruments.
[0047] Figure 1 schematically shows a device 11 for performing spectral analysis to determine a spectrum of a sample 12. This device 11 comprises a source 14 for generating primary radiation 15. This primary radiation 15 is directed onto the sample 12. A focusing element 16, for example an optical lens or a collimator or the like, can be provided between the source 14 and the sample 12. The primary radiation 15 emits secondary radiation 17 in the sample 12 or in at least one layer 13 on the sample 12, which is detected by a detector 18. This device 11 comprises a controller 19, by which at least the source 14 is controlled. The detector 18 converts the detected secondary radiation 17 into a spectrum and forwards it to an evaluation device 21. This evaluation device 21 is, for example, in a spectrometer.This evaluation device 21 comprises a neural network 22, which evaluates the data acquired by the evaluation unit 21. A display device 23 outputs the determined result based on the trained neural network 22. The device 11 can be, for example, an X-ray fluorescence measuring device in which the source 14 is designed as an X-ray tube and the detector 18 comprises an A / D converter to detect the energies of the secondary radiation 17 and convert them into so-called channels and output them. Indices of these channels are proportional to the detected energy. The detected energy or the output channel is referred to below as a feature.
[0048] Alternatively, the device 11 can be configured, for example, to perform laser-induced breakdown spectroscopy (LIBS), and the source 14 is configured as a laser source. The detector 18 is configured as a spectrometer for detecting emitted light.
[0049] Figure 2 shows a diagram of a spectrum of an alloy determined by device 11. The elements contained in the alloy are represented in the spectrum by fluorescence lines of different intensities I, which are plotted along the Y-axis, as well as corresponding fluorescence line energies that correspond to the indices of the channels, each of which is plotted in a feature M along the X-axis. The concentration / layer thickness of the chemical element in the layer 13 or the sample 12 can be determined via the intensity. The chemical element can be determined by assigning the intensity to the respective feature.
[0050] The basis for the output of such a spectrum by the evaluation device 21 is an abstract physical model S = P (0, T, , K), where 0 denotes the concentration of the chemical element, T the layer thickness on an object, X the measurement conditions during the measurement by the device 11 and K the characteristic properties of the device 11. In the previous classic method for determining, for example, an element concentration on the basis of the physical model, parameters were initially fixed and after the measurement with the device 11 the recorded values of the element concentration were optimized in an iterative process until the theoretical spectra show the highest possible agreement with the measured spectrum in order to then output a result.
[0051] This is accompanied by the problem that the evaluation time is long, and the measurement conditions and properties, as well as the calibration if necessary, must be disclosed. Based on this, the goal is to enable fast and precise evaluation of the measurement results through the use of at least one neural network. Furthermore, the application of at least one neural network should be applicable to various end devices.
[0052] The use of an analyzing neural network, which is trained and constructed as follows, enables complex nonlinear functions to be captured very precisely. It was also found that at least one neural network generates an inverse function P -1 (S) can be approximated. This is not possible with the current spectral analysis.
[0053] Against this background, it is proposed to propose a neural network for spectral analysis of the device 11, which at least accelerates the quantitative analysis and enables it with a high precision.
[0054] For the quantitative analysis of element concentrations, it is necessary that the neural network is trained with a large number of spectra, especially simulated spectra. This training can be based on the physical model S = P( <t>), where 0, T, , K, are trained by at least one simulation method on the basis of specific parameters of the samples 12 and / or specific parameters of the device 11. Randomly selected or set parameters can also be used for training. Training can also be carried out using actually recorded features of the spectra. A combination can also be provided. A large number of spectra can be generated using the simulation method. The neural network is advantageously trained with 80,000 to 150,000 spectra.
[0055] Figure 3 shows that starting from the physical model (S= P(c|3)) 24, a plurality of spectra are generated according to the image 25, by which the neural network 22 of the first network architecture 26 is trained.
[0056] The neural network 22 is based on neurons and comprises an input layer, one or more hidden layers, and an output layer. For example, a single-layer network (MLP - multilayer perceptrons) can be provided. Likewise, a convolutional neural network (CNN) or a dense neural network (DNN) or other architectures, such as a dense convolutional neural network (DenseNet), can be used.
[0057] Figure 4 schematically illustrates the structure of the first network architecture 26. Starting from a recorded spectrum 27, this recorded spectrum 27 is evaluated by means of a scaling network 28 and modified with a scaling factor 29, for example, to compensate for different excitation conditions for generating the secondary radiation 17. The spectrum 31 scaled by the scaling network 28 is evaluated with a prediction network 32, and a result 33 of the recorded spectrum 27 is output.
[0058] Figure 5 shows a schematic view of a possible structure of the DenseNet (Densely Connected Convolutional Network) 35. This DenseNet 35 is preferably used for the scaling network 28 and / or the prediction network 32 of the first network architecture 26. The DenseNet 35 can, for example, comprise an input layer 36, a convolutional layer 37, and then a so-called density block (DenseBlock) 38. A further, also selection-specific, sequence of such layers occurs up to the output layer 39.
[0059] The dense block 38 is shown schematically in an enlarged scale in Figure 5. It comprises, for example, four consecutive folded layers 37, each layer 37 being in contact with the adjacent one.
[0060] Figure 6 shows a diagram showing which number of spectra are advantageous for training the neural network 22 in order to achieve a low error rate in the output of results. The error rate of failed tests is plotted along the Y-axis, and the number of spectra is plotted along the X-axis. A comparison is shown, in which the fictitious line with the round dots represents spectra after training the first network architecture 26. The fictitious line marked with a cross shows a comparison of failed tests and failed training. This shows that too small a size of the training data of spectra leads to an increased error rate. Furthermore, it is clear that in a range greater than 80 k (80,000 spectra), preferably 100 k (100,000 spectra) to 160 k (160,000 spectra), a lower loss and a lower error rate are achieved, and this range orthe number of spectra for training the neural network 22, in particular the first network architecture 26, is to be selected in order to achieve a satisfactory result for implementation in real operation.
[0061] Figure 7 shows a schematic diagram comparing various networks for application in the first network architecture 26. The measured mean error (MAE) is plotted along the Y-axis, and the response time in seconds for outputting the result is plotted along the X-axis. It is evident that the application of the CNN architecture, and in particular the DenseNet architecture, results in the shortest evaluation times compared to the MLP architecture when applied in the first network architecture 26.
[0062] Figure 8 shows an overview of the results of evaluating spectra measured on various samples 12 using various devices 11. The previously known evaluation is compared in the "Reference" column with the application of the trained first network architecture 26 for quantitative analysis in the "New" column. For the comparison, the mean absolute error rate MAE in percent and the evaluation time in seconds are shown. The contrasting mean error rates differ negligibly from each other. By applying the first network architecture 26, a significant reduction in evaluation time is achieved, with an average duration of 0.049 + / - 0.002 seconds. This corresponds to a factor of at least 20 times compared to the reference.This clearly shows that the trained first network structure 26 with a large number of simulation spectra, particularly between 80,000 and 160,000 spectra, achieved a significant reduction in evaluation time. The mean absolute error (MEA) remains approximately the same when compared to the quantitative analysis supported by the first network architecture 26.
[0063] A qualitative analysis of the spectrum to identify the chemical element of the sample 12, whether it is present or absent, can be carried out analogously to the quantitative analysis by training the neural network with several actually recorded and / or simulated spectra 27 of different pure elements.
[0064] For the qualitative analysis relating to element identification, a second or further network structure 41 is preferably provided for outputting a result from the acquired spectrum 27. Such a second network structure 41 is shown in Figure 9. Based on the simulated spectra 27, one or more feature reduction methods can be selected for evaluation. One of the feature reduction methods involves mean compression 42. Another feature reduction method can be feature selection 43. In addition, a feature transformation 44 can be performed as a feature reduction method. A preprocessed spectrum 45 is determined using one of these methods 42, 43, 44.This pre-processed spectrum 45 is fed to a second neural network 46, for example with a network architecture MLP 47, CNN 48 or DenseNet 35, and subsequently a result 33 is output in which the identified chemical element is specified.
[0065] In the mean value compression 42, it is preferably provided that the number of features of the determined spectra of the features to be evaluated in the spectrum 27 are reduced to a target spectrum by averaging.
[0066] During feature selection 43, the number of features can be reduced starting from a complete spectrum. For example, the so-called watermelon model can be applied. The watermelon model, also called a watermelon model, is based on a Bayesian error rate estimation. This model first uses a kernel density estimation to approximate the true distribution of the data. Subsequently, a Bayesian error rate estimation is calculated, and the features are individually evaluated for their independence or redundancy. The redundant features are then identified.
[0067] During the feature transformation 44, the features of the respective spectrum 27 are preferably weighted according to relevant and non-relevant features, and the non-relevant features in the spectrum are eliminated.
[0068] The feature transformation 44 can be implemented, for example, using a network architecture according to Figure 10. This is a so-called autoencoder 51. Starting from an input layer 36, the spectrum 27 is fed to an encoder 55, which, for example, has at least a first and second dense layer 56, 57, as well as at least one further dense layer 58 between the encoder 55 and a subsequent decoder 59. At least one dense layer 61, 62 can be provided in the decoder 59. The output layer 39, in turn, follows the decoder 59.
[0069] These previously described procedures serve to train the neural network 22, for example with the first network architecture 24 and / or the second network architecture 41. After this training, the neural network 22 can be further optimized to reduce the evaluation time and / or the computing power.
[0070] Figure 11 shows a diagram in which the feature reduction methods are plotted as a function of the number of features in the spectra 27, as well as as a function of various reduction methods. A factor is plotted along the Y-axis, which indicates the accuracy output of the result with the FI factor from the acquired spectrum in a range between 0 and 1, where 1 corresponds to 100% accuracy. The number of features per spectrum 27 is plotted along the X-axis. Line 61 shows a performance curve when using the autoencoder 55 according to Figure 11. Line 62 shows the performance curve for the compression method 42, and line 63 shows the performance curve for feature selection 43, with the Watermelon model being selected in particular. For all three feature reduction methods 42, 43, and 44, a DenseNet 35 was used as the neural network 22.This shows that, when reducing the number of features per spectrum from, for example, 1024 to 32, a high accuracy of >92% can be achieved in all three cases. The two reduction methods according to lines 62 and 63 enable significantly improved accuracy of the determined spectrum than is the case with the reduction method according to line 61. Thus, feature selection, and in particular the Watermelon model, is the preferred feature reduction method, especially for the first and / or second network architectures 26, 41, to reduce the evaluation time.
[0071] A further feature reduction can consist of a targeted reduction in the number of features of the spectra 27. Figure 12 shows several spectra analogous to the spectrum in Figure 2, which is plotted along the x-axis over, for example, 1024 features. In the spectra shown in Figure 13, the number of features of the spectra 27 has been reduced, for example, from 1024 features to 32 features. For this feature reduction, the previously described feature selection can also be performed, for example.
[0072] Furthermore, a network reduction of the neural network 22 can be applied to accelerate the evaluation time. With such a network reduction, the network architecture can be checked in a first step to determine which neurons, filters, or parameters impair performance. Then, redundant components in particular are eliminated. In a second step, the parameter size and the floating-point operations (FLOPs) can be reduced within the network reduction method. For example, a neural network 22 is symbolically represented in Figure 14 in order to verify the number of features of the spectrum 27 according to Figure 12. Figure 15 symbolically represents a reduced neural network 22 compared to that in Figure 14, which can be trained using a spectrum 27 according to Figure 14.The data size of the neural network 22 according to Figure 14 is, for example, 5.2 MB, and the computing time is, for example, 540 ms. In Figure 15, the reduced neural network 22 has a data size of, for example, 0.1 MB and a computing time of, for example, 0.9 ms. This clearly demonstrates the advantages of network reduction.
[0073] Furthermore, the evaluation time of the neural network 22 can be shortened by network quantization. Typically, a bit size of 32 bits is used for the floating-point operations (floats). Network quantization reduces the bit size for the floating-point operations to a bit size lower than 32 bits and / or an integer representation (int). Quantization can preferably result in a reduction to 16-bit floats or 8-bit ints, or, for example, to 8-bit ints with 16-bit activations. For example, by quantizing a 32-bit float to a 16-bit float, the data size can be reduced by approximately 50%. The same applies to further bit reduction.
[0074] Figure 16 shows another table showing a performance comparison of the various reduction methods. The data sizes (MB) and the evaluation time (ms), as well as the accuracy with the factor (FI), are compared. Furthermore, the application of a CNN architecture 48 on the one hand and a DenseNet architecture 35 on the other hand are compared. The first row of the table, under "Base," shows the neural network 22 of the first and / or second network architecture 26, 41 without performing feature reduction and / or network reduction and / or network quantization.
[0075] Through feature selection 43, a significant reduction in data size by a factor of 5 can be achieved with the CNN architecture 48 compared to the baseline. The data size remains the same with the DenseNet architecture. The evaluation time can be reduced by 13.7 times for the CNN architecture 48 and by 16 times for the DenseNet architecture 35. The accuracy of the result remains virtually the same when applying feature selection compared to the baseline.
[0076] The network reduction method and / or the network quantization method can enable a further reduction in data size and analysis time, as shown in the table. However, the accuracy of the results remains unchanged compared to the baseline. An optional combination of the three reduction methods listed—feature reduction, network reduction, and / or network quantification—can also be applied. If all three reduction methods are used in conjunction to optimize the neural network 22, the data size can be reduced by approximately 29 times and analysis time by 65 times for the CNN architecture 48. With the DenseNet architecture 35, the file size can be reduced by 52 times and the analysis time by 600 times.It is therefore obvious that by reducing the features, in particular by selecting the features, of the spectra for the quantitative and / or qualitative analysis on the one hand and by training the neural network 22 with these spectra r and preferably by additional network reduction and / or network quantization of the neural network 22, not only can the computational effort be reduced and thus the costs reduced, but also the evaluation times can be shortened considerably.
[0077] Figure 17 shows a preferred embodiment for training the neural network 22. This embodiment is particularly suitable for producing a large number of devices, in particular mass production, wherein the evaluation time is significantly reduced while maintaining high accuracy of the results. In a first step 71, the neural network 22 is trained using the actually acquired spectra and / or spectra simulated using a simulation method for qualitative and quantitative analysis. Preferably, the number of simulated spectra is many times greater, in particular at least by a factor of 100, than the number of actually acquired spectra. The training is based on a second network structure 41, which is described in more detail in Figure 9, in order to train the first network architecture 26 on the chemical elements.In a further step 72, an analogous procedure is followed as in step 71 for the quantitative analysis, as shown in Figures 4 and 5. This provides, in particular, that pre-processed spectra are acquired from the determined spectrum 27 using the feature reduction method of feature selection 42, which are then trained using the DenseNet.
[0078] Based on this, in a further step 73 for feature reduction, the feature selection 42 is selected, which is described in Figures 12 and 13. Subsequently, in step 74, the preprocessed spectra of the first and / or second network architecture 26, 41 are used as a basis. Subsequently, in step 75, network reduction is performed, followed by network quantization in step 76. This allows for the creation of a process-optimized method for performing spectral analysis, particularly for large-scale production, which achieves a reduction in data size by up to 52 times and a reduction in evaluation time by up to 600 times compared to conventional methods. At the same time, it is possible to use conventional measuring devices with reduced computing power.
[0079] Figure 18 shows a schematic view of a meta-learning method 64 with a meta-network 65. Such a meta-learning method 64 is used in particular for the calibration of devices 11 or measuring instruments manufactured in a factory. Furthermore, after a certain period of operation, it is often necessary to recalibrate the devices 11 in order to correct any drift that may occur during operation of the measuring instrument.
[0080] The meta-learning method 64 provides for information to be learned using various tasks so that a meta-network can quickly adapt to a new, unknown task, in particular to a device 11. In a first step 66, spectra 27 are acquired from a sample 12 using a known measuring device 11 under various measurement conditions. These spectra 27 are evaluated by a first network architecture 26 of the neural network 22. Spectra 27 are then determined from further samples 12 using further known measuring devices 11, and / or simulation spectra are generated, taking into account not only the various measurement conditions but also various characteristic properties of the measuring device 11. These spectra from the at least one known device 11 are used in step 66 to train a meta-network 65. An agnostic meta-learning model (MAML - Model-Agnostic Meta-Learning) can be used for this purpose.The meta-network 65 is thus trained for the calibration of the devices 11.
[0081] To calibrate an unknown measuring device 11, the meta-network 65 is first trained in step 67 with the measurement conditions and / or measurement properties of the unknown measuring device 11. Subsequently, the unknown measuring device 11 is calibrated by the meta-network 65. Advantageously, the calibration of the unknown measuring devices 11 can be performed without measuring a sample 12. Alternatively, at least one measurement can be performed on the unknown measuring device 11. In this process, the meta-network 65 is further trained, thereby performing an improved calibration on the unknown measuring device 11.
[0082] After a predetermined or longer operating period of the measuring device 11, recalibration may be required, as per step 68. Measurement data from the measuring device 11 to be recalibrated is acquired, so that the meta-network 65 is in turn trained using this measurement data. Due to the further training of the meta-network 65 by the measuring device 11 to be recalibrated, the meta-network 65 can be further trained and perform a faster and improved recalibration.
[0083] The meta-learning method 64 is preferably divided into two processes, namely a training before calibrating the unknown measuring device 11 and a training after calibrating the then known measuring device 11. The meta-learning method 64 can enable significant cost savings for the industry and the customer.< / t>
Claims
Claims Method for carrying out a spectral analysis for evaluating a spectrum of a sample (12), - with a measuring device (11), - in which a primary radiation (15) from a source (14) is directed onto the sample (12), - in which secondary radiation is emitted by the sample (12) or at least one layer (13) of the sample (12) by means of primary radiation (15), - in which a spectrum of the secondary radiation (17) is detected by a detector (18), - in which the at least one detected spectrum is fed by the detector (18) for evaluation to a computer-assisted evaluation device (21), in which the at least one detected spectrum is evaluated in at least one analyzing neural network (22) of at least one first network architecture (26) which is trained for quantitative analysis of the spectrum, - in which the neural network (22) of the first network architecture (26) is trained by a plurality of simulation spectra which are based on a physical model (S = P( <t>)) is generated using a simulation method, - in which at least one second network architecture (41) is trained with a second analyzing neural network (46) for the qualitative analysis of the detected spectrum, - in which a concentration and / or an identification of the at least one element of the sample (12) is output as a result from the spectrum analyzed by the neural networks (22, 46). Method according to claim 1, characterized in that the neural network (22, 46) of the at least one network architecture (26, 41) is trained to generate an inverse function P _ X (S) is approximated, so that in the quantitative analysis the at least one element concentration of the acquired spectrum of the sample (12) is determined and output and / or in the qualitative analysis the absence or presence of at least one chemical element of the acquired spectrum of the sample (12) is determined and output. Method according to one of the preceding claims, characterized in that the neural network (22) of the first network architecture (26) is trained by a plurality of simulation spectra which are generated starting from the physical model (S=P(0, T, , K)) using a simulation method, in particular a Monte Carlo simulation, and / or that the first network architecture (26) is trained with a plurality of actually acquired spectra.Method according to claim 3, characterized in that the simulation spectra are determined with predetermined parameters, in particular with the concentration of the chemical element or concentrations of the at least one element of an alloy, element-specific physical constants, the layer thickness of at least one layer, different measuring conditions of the measuring devices (11) and / or characteristic properties and / or device-specific data from the one or different measuring devices (11). Method according to claim 1, characterized in that the neural network (46) of the second network architecture (41) is trained with a plurality of spectra of at least one chemical element, in particular differing intensity distributions of energy spectra of the at least one chemical element, which are determined by an X-ray fluorescence analysis, and preferably the at least one chemical element has an atomic number greater than 9.Method according to one of the preceding claims, characterized in that for the quantitative analysis of the acquired spectrum in the first network architecture (26), a scaling network (28) is used, in which the at least one acquired spectrum of the sample (12) is adjusted with a scaling factor (29), preferably to compensate for one or more, in particular different, excitation conditions, and that the scaled spectrum is subsequently fed to a prediction network (32), and a result of the spectrum is output by this prediction network (32). Method according to claim 6, characterized in that the prediction network (32) and / or the scaling network (28) is constructed with a neural network, in particular a dense convolutional network (DenseNet), which comprises at least one dense block with convolutional layers (DenseBlock), in which each layer within the block is preferably in contact with the other layers.Method according to one of claims 3 or 5, characterized in that for the quantitative analysis and / or for the qualitative analysis of the acquired spectrum, at least one feature reduction method is applied to the generated simulation spectra. Method according to claim 8, characterized in that, starting from a plurality of simulation spectra and / or acquired spectra, preprocessed spectra are generated using the at least one feature reduction method, and the neural network is trained using these preprocessed spectra. Method according to claim 8 or 9, characterized in that, by the one feature reduction method, the number of features of the spectrum to be evaluated is reduced from 1024 to 512, 256, 128, 64, 32, or 16 features. Method according to one of claims 8 to 10, characterized in that - that the feature reduction method is carried out as a mean compression, in which the number of features of the acquired spectrum and / or simulation spectra is reduced to a target spectrum by averaging, and / or - that the feature reduction method is carried out as a feature selection, in which a reduction of the number of features of the respective spectrum is carried out starting from their original size and reduced to a pre-processed spectrum, and / or - that the feature reduction method is carried out as a feature transformation, preferably with an autoencoder network, in which the features of the respective spectrum are weighted according to relevant and irrelevant features, and the irrelevant features are eliminated. Method according to claim 11, characterized in that the feature selection is carried out with a watermelon model. in which the selection of the features is preferably carried out by a Bayes error rate estimation. Method according to one of the preceding claims, characterized in that the analyzing neural network (26, 46) is trained with at least one further model for reducing the evaluation time after the simulation of the spectra and / or reduction of the spectra for the quantitative analysis and / or the selection of the chemical elements for the qualitative analysis. Method according to claim 13, characterized in that the analyzing neural network (22, 46) is trained with a model for network reduction after the simulation of the spectra and / or reduction of the spectra for the quantitative analysis and / or the selection of the chemical elements for the qualitative analysis, in which model the neural network (22, 46) is reduced by components, such as neurons, filters and / or parameters, in particular is reduced by the redundant components.Method according to claim 14, characterized in that a filter reduction is performed on the neural network (22, 46) comprising convolved layers, and / or a neuron reduction is performed on the neural network (22, 46) comprising dense layers. Method according to claim 13, characterized in that, after simulating the spectra and / or reducing the spectra for the quantitative analysis and / or selecting the chemical elements for the qualitative analysis, the analyzing neural network (22, 46) is trained with a quantization model in which the data size, in particular of a node of the neural network (22, 46), is reduced, preferably to a bit size of less than 32 bits (32-bit floats). Method according to one of the preceding claims, characterized in that the measuring device (11) is trained by a neural network of a meta-network (65) using simulation data from a number of known measuring devices to carry out a meta-learning method (64) for unknown measuring devices (11). Method according to claim 17, characterized in that for calibrating unknown measuring devices (11) for spectral analysis, the meta-learning method (64) with the meta-network (65) is used, in which the properties and / or the changing measurement conditions and / or the measurement tasks of the measuring device (11) to be calibrated are recorded, and the meta-network (65) is additionally trained using simulation data and / or measured data from the unknown measuring device (11). Method according to claim 17 or 18, characterized in that the meta-network (65) becomes operational at least for the quantitative analysis of the unknown measuring device (11) through the training.Method according to one of the preceding claims, characterized in that the analyzing neural network (22, 46, 65) is constructed at least as a multi-layer network (MLP), a convolutional neural network (CNN), or a dense neural network (DNN), in particular a dense convolutional network (DenseNet). Method according to one of the preceding claims, characterized in that in a first step for examining the sample (12), the quantitative analysis is carried out, and for further analysis of the sample (12), the quantitative analysis of the sample (12) is carried out. Method according to one of the preceding claims, characterized in that the source (14) is designed as an X-ray tube and X-ray radiation is emitted. Measuring device for carrying out a spectral analysis to determine a spectrum of a sample, with a source (14) for emitting a primary radiation (15) onto the sample (12), with a detector (18) for detecting a secondary radiation (17) which is emitted after the excitation of the sample (12) by means of primary radiation (15), with a computer-assisted evaluation device (21) which evaluates the at least one detected spectrum of the detector (18), wherein at least one analyzing neural network (22) with at least a first network architecture (26) is provided for the quantitative analysis of the spectrum, wherein at least one second network architecture (41) with a second analyzing neural network (46) is provided for the qualitative analysis of the spectrum, and with an output device which outputs a result of at least one function generated by the at least one neural network (22,46) for a concentration and / or an identification of the at least one element of the sample (12). A computer program for performing a spectral analysis to determine a spectrum of a sample (12), wherein the computer program is stored on at least one computer-readable storage medium that is executable on a computer system and causes the computer system to execute a method according to one of claims 1 to 22.< / t>